Gemma 4 QAT on One TPU v5e: What Runs and What Doesn't

DEV Community·4 points·xbill·about 9 hours ago·dev.to

Google's quantization-aware-trained Gemma 4 weights, repacked into int4 and int8 formats vLLM serves on TPU, on one v5e chip. What runs: every size from E2B to 26B, level with bf16 through 12B, up to 2.4 points above Google's own 4-bit exports at the same speed, and 12B at 675 output tokens per second. What doesn't: bf16 above E2B, 31B in any build, and 26B past a 2,176-token context.

Read the full article at dev.to →

Related stories

Related topics