TensorFold's new inference engine delivers decode speeds over 62 tokens per second on a single stream and 119 across five, with prefill at 2,500 tokens per second and a 256k context window. It outperforms prior vLLM…

2 points•soltanov•1 day ago•1 comment•

1 comment

spottedmarley1 day ago
Been waiting for MTP in llamacpp for a about a week now so I can speed up 3.8:flash-next but this might be just what I'm looking for. :)

Update: Looks good so far

Flash-Next on llama.cpp with no MTP: 28.8t/s decode, ~620t/s prefill

on llama.cpp with MTP (froze twice): 44.1t/s decode, ~593t/s prefill

on TensorFold + MTP: 61.0t/s decode, 2,455t/s prefill

That's about 2.1 to 2.4x on decode and about 4x on prefill (!!)

Need to do some long context testing but Im stoked.

Read the full thread on Hacker News →

Related stories