Looped transformers and Mixture-of-Experts (MoE) offer complementary routes to efficient scaling: recurrence increases computational depth at fixed parameters, while MoE sparsity expands total capacity at fixed active…
0 comments
No comments yet.
Related stories
- Hacker News · 2 points · 7 days ago
- Field Notes on Scaling MoE Expert Parallelism with DeepEPnousresearch.comLobsters · 2 points · 8 months ago
- Runtime Dynamic Compression of Mixture of Experts [pdf]timdettmers.comHacker News · 1 points · 8 days ago
- Ars Technica · 0 points · 4 days ago
- Lobsters · 6 points · about 14 years ago
- Scaling to 100k Usersalexpareto.comLobsters · 18 points · over 6 years ago