TensorFold's new inference engine delivers decode speeds over 62 tokens per second on a single stream and 119 across five, with prefill at 2,500 tokens per second and a 256k context window. It outperforms prior vLLM…
1 comment
spottedmarley1 day ago
Been waiting for MTP in llamacpp for a about a week now so I can speed up 3.8:flash-next but this might be just what I'm looking for. :)
Update: Looks good so far
Flash-Next on llama.cpp with no MTP: 28.8t/s decode, ~620t/s prefill
on llama.cpp with MTP (froze twice): 44.1t/s decode, ~593t/s prefill
on TensorFold + MTP: 61.0t/s decode, 2,455t/s prefill
That's about 2.1 to 2.4x on decode and about 4x on prefill (!!)
Need to do some long context testing but Im stoked.
Read the full thread on Hacker News →
Related stories
- Hacker News · 1 points · about 7 hours ago
- Hacker News · 1 points · 9 days ago
- DEV Community · 1 points · 3 days ago
- Hacker News · 1 points · 3 days ago
- Turning GLM-5.3-Flash into a Jev-like decision modelprivatemode.aiHacker News · 113 points · 4 days ago
- Ant Group releases finance-focused Ling-3.0-flash-Finartificialanalysis.aiHacker News · 1 points · 8 days ago