Post-training quantization (PTQ) is widely used to deploy large language models efficiently, but its effect on reasoning models is not well understood. Across math, coding, and science QA, we find that aggressive PTQ…
1 comment
Highly quantized models, especially with highly quantized KV caches, will, effectively, attend to the wrong tokens and be unable to easily discern highly similar tokens. The bastardized way of explaining this is gradient descent techniques get stuck in localized minimum and global maximums, so what happens when you turn the slopes into hard stair steps?
We need to move to smaller models and smaller caches and better samplers, not new quant methods (although I'm willing to also take those too).
Read the full thread on Hacker News →
Related stories
- Reasoning Yield: share of tokens spent resolving uncertaintyjeffauriemma.leaflet.pubHacker News · 1 points · 8 days ago
- Hacker News · 4 points · 7 days ago
- Hacker News · 1 points · 7 days ago
- Hacker News · 2 points · 8 days ago
- Microsoft Publisher will no longer be supported after October 2026support.microsoft.comHacker News · 3 points · 7 days ago
- Initial-scale no longer neededquirksmode.orgHacker News · 1 points · 5 days ago