How our Kimi K3 megakernel on TPU v7 reaches over 700 tokens/s with speculative decoding and nearly 2× GB200's batch-one decode throughput.

5 points•xutingl•7 days ago•1 comment•

1 comment

xutingl7 days ago
Inferact's first TPU megakernel for Kimi K3 reaches 709 tokens/s on low-concurrency decode, against 450 tokens/s for our GB200 baseline, both with DSpark speculative decoding. Without speculative decoding, the megakernels for K3 and Qwen 3.8 27B deliver roughly 1.4 to 2× the decode throughput of the GB200 baseline at batch sizes 1 through 8.

Read the full thread on Hacker News →

Related stories