The Next 3× in Inference Won't Come From Faster Kernels
1 comment
floathub1 day ago
One unique idea for solving the chip shortage: pack more models into each existing GPU by having them share the resources. :-)
Read the full thread on Hacker News →
Related stories
- Hacker News · 134 points · about 11 hours ago
- Writing Rust code that's faster than state-of-the-art libraries by asking agents to make the code fasterminimaxir.comLobsters · 5 points · 8 days ago
- Hacker News · 2 points · 8 days ago
- Hacker News · 2 points · 9 days ago
- Hacker News · 1 points · 9 days ago
- The Basics of Transformer Inferencejax-ml.github.ioHacker News · 2 points · 9 days ago