How our Kimi K3 megakernel on TPU v7 reaches over 700 tokens/s with speculative decoding and nearly 2× GB200's batch-one decode throughput.
1 comment
xutingl7 days ago
Inferact's first TPU megakernel for Kimi K3 reaches 709 tokens/s on low-concurrency decode, against 450 tokens/s for our GB200 baseline, both with DSpark speculative decoding. Without speculative decoding, the megakernels for K3 and Qwen 3.8 27B deliver roughly 1.4 to 2× the decode throughput of the GB200 baseline at batch sizes 1 through 8.
Read the full thread on Hacker News →
Related stories
- Hacker News · 124 points · about 8 hours ago
- Show HN: A Simple GPU Job Schedulergithub.comHacker News · 2 points · 4 days ago
- DEV Community · 15 points · 14 days ago
- DEV Community · 1 points · 9 days ago
- DEV Community · 10 points · 7 days ago
- Sparse resources helped my GPU-driven renderer memory usagezino2201.substack.comHacker News · 1 points · 10 days ago