Notes on MiMo-V2.6 Pro’s GQA and sliding-window attention, agent training tasks, reward signals, and large RL batches.
1 comment
nchmy8 days ago
This is silly. We can put aside whether artificial analysis' benchmarks are useful. But they're comparing mimo 2.6 to deepseek 4.0 Pro, despite deepseek essentially deprecating 4.0 pro with the release of 4.1 Flash (they literally planned to stop offering it altogether, but ended up keeping it)
Read the full thread on Hacker News →
Related stories
- Hacker News · 1 points · 1 day ago
- Hacker News · 1 points · 5 days ago
- Ars Technica · 0 points · 5 days ago
- The Verge · 0 points · 8 days ago
- The Verge · 0 points · 5 days ago
- Xiaomi MiMo Desktopmimo-ai.xiaomimimo.comHacker News · 1 points · 9 days ago