When we train large ML models, we have to split (or "shard") their parameters or inputs across many accelerators. Since LLMs are mostly made up of matrix multiplications, understanding this boils down to understanding…
1 comment
gryfft7 days ago
[2025]
Read the full thread on Hacker News →
Related stories
- Ars Technica · 0 points · 8 days ago
- Transformer: A Novel Neural Network Architecture for Language Understandingresearch.googleblog.comLobsters · 4 points · about 9 years ago
- MongoDB Sharding: Manage a Sharded Cluster Visuallyvisualeaf.comHacker News · 3 points · 9 days ago
- Hacker News · 44 points · 9 days ago
- Hacker News · 2 points · 5 days ago
- Hacker News · 54 points · 1 day ago