Running a Jev-Style Decision Model on One TPU v6e: What Fits, What It Costs, and What Changes From a GPU
DEV Community·10 points·xbill·6 days ago·dev.to
Gemma 4 E2B, E4B, 12B and a 26B-A4B fp8 build read by their label probabilities with vLLM on one TPU v6e chip, checked against the same read on an NVIDIA L4 and against Jev 1.13.0's published results. What fits one chip, how to read labels when vLLM on TPU returns only the top 32 log-probabilities, speed, cost, and why no 31B loads today.
Read the full article at dev.to →
Related stories
- Hacker News · 1 points · 9 days ago
- Hacker News · 3 points · 3 days ago
- Hacker News · 1 points · 5 days ago
- DEV Community · 1 points · 7 days ago
- DEV Community · 8 points · 7 days ago
- Hacker News · 7 points · 11 days ago