Notes on MiMo-V2.6 Pro’s GQA and sliding-window attention, agent training tasks, reward signals, and large RL batches.

3 points•ModelForge•9 days ago•1 comment•

1 comment

nchmy8 days ago
This is silly. We can put aside whether artificial analysis' benchmarks are useful. But they're comparing mimo 2.6 to deepseek 4.0 Pro, despite deepseek essentially deprecating 4.0 pro with the release of 4.1 Flash (they literally planned to stop offering it altogether, but ended up keeping it)

Read the full thread on Hacker News →

Related stories