1 comment

tudorizer8 days ago
That's to be expected from specs alone, no? The memory bandwidth is the culprit. The comparison is almost unfair for most inference, especially on a dense model.

I wrote more details about what's appropriate for a DGX Spark here: https://spark.enverge.ai/blog/dgx-spark-prefill-vs-decode

[disclosure: co-founder of the company linked]

Read the full thread on Hacker News →

Related stories