VRAM for local LLMs: why memory bandwidth sets your tokens per second
DEV Community·2 points·axrisi·about 20 hours ago·dev.to
VRAM for LLMs is a bandwidth problem: every token streams the whole model from memory. Bandwidth per tier, the 20x offload cliff, and what fits in 16, 24 or 48 GB.
Read the full article at dev.to →
Related stories
- Hacker News · 1 points · about 14 hours ago
- Hacker News · 1 points · 8 days ago
- Hacker News · 2 points · 7 days ago
- Hacker News · 1 points · 6 days ago
- Hacker News · 3 points · 4 days ago
- Hacker News · 6 points · 10 days ago