Running a Jev-Style Decision Model on One TPU v6e: What Fits, What It Costs, and What Changes From a GPU

DEV Community·10 points·xbill·6 days ago·dev.to

Gemma 4 E2B, E4B, 12B and a 26B-A4B fp8 build read by their label probabilities with vLLM on one TPU v6e chip, checked against the same read on an NVIDIA L4 and against Jev 1.13.0's published results. What fits one chip, how to read labels when vLLM on TPU returns only the top 32 log-probabilities, speed, cost, and why no 31B loads today.

Read the full article at dev.to →

Related stories

Related topics