Repacked QAT Gemma 4 on One TPU v5e: 12B Serves at 675 Tokens per Second

DEV Community·8 points·xbill·about 8 hours ago·dev.to

Google's quantization-aware-trained Gemma 4 weights, repacked into int4 and int8 formats vLLM serves on TPU. On one v5e chip the repacks serve every size from E2B to 26B, read the suite level with bf16 through 12B, score up to 2.4 points above Google's own 4-bit exports at the same speed, and put 12B on the chip at 11.31 GiB and 675 output tokens per second.

Read the full article at dev.to →

Related stories

Related topics