Gemma 4 QAT on One TPU v5e: What Runs and What Doesn't
DEV Community·4 points·xbill·about 9 hours ago·dev.to
Google's quantization-aware-trained Gemma 4 weights, repacked into int4 and int8 formats vLLM serves on TPU, on one v5e chip. What runs: every size from E2B to 26B, level with bf16 through 12B, up to 2.4 points above Google's own 4-bit exports at the same speed, and 12B at 675 output tokens per second. What doesn't: bf16 above E2B, 31B in any build, and 26B past a 2,176-token context.
Read the full article at dev.to →
Related stories
- DEV Community · 8 points · about 8 hours ago
- DEV Community · 10 points · 8 days ago
- DEV Community · 7 points · 7 days ago
- DEV Community · 7 points · 14 days ago
- DEV Community · 10 points · 22 days ago
- DEV Community · 0 points · 7 days ago