tpu

3 stories and discussions about tpu, aggregated from every source we track.

1.

Gemma 4 E2B, E4B, 12B and a 26B-A4B fp8 build read by their label probabilities with vLLM on one TPU v6e chip, checked against the same read on an NVIDIA L4 and against Jev 1.13.0's published results. What fits one chip, how to read labels when vLLM on TPU returns only the top 32 log-probabilities, speed, cost, and why no 31B loads today.

10 points•xbill•6 days ago•0 comments
2.

How our Kimi K3 megakernel on TPU v7 reaches over 700 tokens/s with speculative decoding and nearly 2× GB200's batch-one decode throughput.

5 points•xutingl•7 days ago•1 comment•
3.
2 points•simonpure•10 days ago•0 comments•

Related topics