When we train large ML models, we have to split (or "shard") their parameters or inputs across many accelerators. Since LLMs are mostly made up of matrix multiplications, understanding this boils down to understanding…

1 points•lawrenceyan•7 days ago•1 comment•

1 comment

gryfft7 days ago
[2025]

Read the full thread on Hacker News →

Related stories