Introducing the MiMo-V2.6 series: frontier intelligence, all the modalities, built in public.

1130 points•volf_•9 days ago•482 comments•

482 comments

rao-v9 days ago
I know we have strong views on what a truly open model is (open weights, open training data, open training code etc.) but I really like how transparent they’ve been about the training of this model.

The realtime dashboard they shared during training (https://mimo.xiaomi.com/rl/) was an incredible learning and teaching tool for me, and they’ve been unusually comprehensive in sharing details about their methodology (check out that tech report - it's got lots of clever behind the scene tricks like Google or Deepseek writeups) and benchmark scores (even the stuff they didn’t do well on).

If you’re releasing an open model going forward, please consider offering the community more of this transparency!

bicepjai9 days ago
I was absolutely mind blown when I saw how they were publishing that training dashboard while US models publish 100s of pages of reports (just provide a "copy as MD" button, folks, in the future). I was thinking about doing something similar but did not know how to show it, and this is a perfect example for someone who wants to show whatever they are training, for me it was local training on a consumer GPU.

My dream is to see this like a dashboard for a model trained across distributed machines, like Bitcoin mining, where minted coins are given to people whose machines were used for training. I don't know if they are worth it, but bragging rights alone, like a tag they can put on a website or social media, will be good enough for me.

BodyCulture9 days ago
The distributed training collective is a good idea, let’s discuss details:

A) How to prevent malicious injection of bad training data?

B) How to handle copyright violations, will participants be responsible and will they have to pay the creators?

C) Can this create an income stream for content creators and how to avoid abuse, eg feeding with AI content?

So many more, but let’s focus on these before we break things fast because we didn’t think about them.

rao-v9 days ago
This dashboard is almost certainly built on verl (https://github.com/verl-project/verl), which comes with a bunch of dashboarding capabilities built in (that doesn't look too dissimiliar to these dashboards).
sally_glance9 days ago
Not an expert on this but I think the RL runs need to work sequentially? I wonder what the opportunities for distributed execution would be... Maybe parallelizing the benchmark task or inference
ignoramous9 days ago
> got lots of clever behind the scene tricks like Google or Deepseek writeups) and benchmark scores

Xiaomi MiMo is led by Luo Fuli, a former Alibaba & DeepSeek employee. Perhaps it is due to Luo just how similar Xiaomi's tech & GTM approach is to DeepSeek's.

- How Luo Fuli Keeps an Earthy Touch as she Soars Through the AI World, https://newsen.pku.edu.cn/news_events/news/people/15385.html (https://archive.vn/I8Pmu).

- Luo Fuli, the 30-year-old ‘AI genius girl’ behind DeepSeek’s success?, https://e.vnexpress.net/news/tech/personalities/who-is-luo-f... (https://archive.vn/sb3B6).

pimeys9 days ago
Open tech is cool. Speeds up all progress...
earthnail9 days ago
Thanks so much for sharing this. As someone who mostly watches from the sideline, can you share what you can see in this dashboard that someone like me can't see? Is it the metrics themselves that they measure (the metrics tab is absurdly detailed), something in the notices, or something else I missed?
rao-v9 days ago
I might turn this into a blogpost if folks are interested, but my god there is so much clever info in that dashboard.

Here is one really neat bit:

A cutting edge training idea (for agents, it's been used elsewhere for ages) is on-policy RL, basically, it's not enough to say "here is an end to end agentic sequence (including tool calls etc.) that is perfect" you want to say "here is a sequence you might actually have generated that turns out to be correct".

Basically, it's more training efficient to improve models with small tweaks to do more of the right thing they are already doing sometimes than from some perfect oracular "this is the way" answer.

(if you've ever tried to teach humans new skills, you’ve probably noticed this too!)

When you do that, you care about how far the model you are updating (improving) has deviated from the one being used to generate rollouts (agentic rollouts for hard problems can take hours with lots of tool calls, so you can't keep redeploying every slight improvement).

Lo and behold, the dashboard literally has:

partial/avg_staleness (likely the measure of how many micro iterations the "generate answers" model is behind the "improving based on the occasional right answer" model)

train_infer_diff/new_infer/kl (a more direct KL divergence based way of measuring how differently the two models generate tokens)

How cool is that?!

And don't get me started on the clever ideas hiding behind dynsam/avg@n ...

tancop9 days ago
The best thing they did is being open about all the setbacks they had to deal with. They logged every restart with a reason, talked about dropping a cyber dataset after it degraded coding benchmarks. Also published real time training loss, benchmark scores after every checkpoint and running cost estimates.

Really the only thing missing was dataset descriptions, the dashboard only had random IDs like "dataset-zrso". I guess it's their lawyers fault.

verdverm9 days ago
the existence, who else has a live dashboard for the RL late-training?
dang9 days ago
> The realtime dashboard

'twas discussed a few days ago:

Xiaomi Mimo 2.6 live post-training dashboard - https://news.ycombinator.com/item?id=49732270 - Sept 2026 (155 comments)

figassis9 days ago
Because the world is conditioned to distrust chinese models (pick your reason here), I believe this is critical for them in order to kill any arguments outside the actual merits. They probably spent a lot of time making this call and might pay off on the long run.
embedding-shape9 days ago
Whatever well-founded/or not distrust people have in Chinese models, this dashboard proves/shows nothing that can make them trust it more or less. It's like providing the journalctl logs of your HTTP server on your website and claim this proves NSA isn't listening or something.
dgellow9 days ago
> Because the world is conditioned to distrust chinese models

I think you mean mostly the US

margorczynski9 days ago
China will most probably win the AI race in the long run because of one major bottleneck the US has - energy. The electric energy and grid buildout in China has been massive since a long time and there is simply no way for the US to quickly catch up.

No matter how much cash you throw you can't just materialize a 100 nuclear reactors to power the data centers.

xynelius9 days ago
I was curious how much energy is actually needed to power these datacenters, so I did a little bit of math.

Looking at Nvidia revenues in the past few years, there's maybe $300 billion worth of GPUs currently deployed in the U.S. The B200 costs ~$40k, so we have 7.5 million B200-equivalents, which draw 1000W. Running these at full capacity requires 66 TWh a year, or ~1.5% of total current U.S. electricity consumption. Maybe a bit more to account for inefficiencies, cooling, and other components, but not more than ~2.5% total I would guess.

So it's not that much in reality, but will definitely grow fast.

traceroute668 days ago
> Running these at full capacity requires 66 TWh a year,

I think your numbers are off.

For a start you are effectively calculating a GPU only number.

I think 100Twh would be the minimum level to think about "all-in". And even that is probably being generous.

Remember, afterall that Google have just bought half the capacity (4.1Twh) of a nuclear power plant in Finland, on top of 630 MW of wind and 94MW of battery.

This is to cater for three new sites at Kajaani, Muhos, and Vaala and expansion at Hamina. So basically 3.5 datacentres.

But Finland is quite a small place. The US has more sites and bigger sites, so the numbers probably grow exponentially very quickly.

podgorniy9 days ago
It's even less if adjusted for non 100% (more like 0.35% of total). Yet factual impact on the grid and other industries will not be as small as numbers appear. For example most probably there will appear transformer (electric) and witchgear deficites. And this is not the only supply chain bottleneck in this subject
boguscoder9 days ago
For 1000w draw it also generates almost as much heat, is your math including all the cooling required?
beachy9 days ago
It's worth looking at similar industries with enormous electricity requirements such as aluminium smelting, where the plant can be located in a friendly country but the product is owned and controlled back in the US.

Aluminium is often described as "congealed electricity". Ship bauxite to wherever power is cheap and stranded, turn it into metal, and ship the metal out. Here in NZ, Tiwai Point is the textbook case, with London-based Rio Tinto running a smelter on the other side of the world that exists mainly because Manapōuri hydro had nowhere else to go.

AI data centres can be just the same - even more so, since the plant's assets (its chips) are virtually perishables, so there is less concern about assets becoming stranded if the host goes rogue. All the US needs is friendly and stable allied countries with cheap power.

nicolasjungers9 days ago
All the US needs is friendly and stable allied countries...
busssard9 days ago
datacenters will quickly be used by the country themselves... A country can only use Aluminum with the required processing industry existing.

Datacenters enable anyone with a computer to use it.

gpt59 days ago
the bottleneck right now is compute, not energy, and it's not even close. That is why RAM, SSD, CPU, and GPU prices are increasing exponentially, while solar panels are dropping.

Also, unlike China, US companies are building data centers all over the world, which gives them higher distribution and ability to colocate with the energy production sources.

Lastly, energy production costs have been decreasing over the last couple of decades. If they will increase, the market will react, as it always does. Looking backwards does not predict the future in this case.

ilaksh9 days ago
I think we are talking about the medium term for China to pull ahead. They just released a new domestic AI chip which seems in important ways to have caught up to Nvidias previous generation. They have a LOT more manufacturing capacity. They are collaborating via open source by default. They have better materials access and much more energy capacity. They have many more people overall and more researchers.

There is a strong chance most of the researchers are pulled out of the US and Europe if WWIII really kicks off or even if there is just more global crisis and concern.

One other thing about the power needs. Within a few years, the power efficiency of AI chips is likely to improve by a factor of 20, 50 or more times by switching to true compute-in-memory architecture with new materials that have made rapid progress lately.

aenis9 days ago
Its also pretty telling that the release of the new Chinese chip wasn't met with an intense discussion here on HN. And it's arguably way, way more important than a version bump on some benchmaxxed model or two. The advances in Chinese chipmaking are super exciting.
ShinyLeftPad9 days ago
> There is a strong chance most of the researchers are pulled out of the US and Europe if WWIII really kicks off

for wwiii it would likely imply a war in asia too, so it's not as if PRC will be a safe place for those researches to run away to.

glub9 days ago
> They are collaborating via open source by default.

It's almost as if collaboration is the foundation of scientific progress. Too bad US has lost the notes.

jesterson9 days ago
> China will most probably win the AI race in the long run because of one major bottleneck the US has - energy

Plus another bottleneck - China produces engineers, the US produces lawyers.

utopiah9 days ago
I'm curious about that, do you have relevasnt metrics? I imagine stats out of universities, engineer schools, etc and maybe number of patents could be used but fearing those could be gamed.
stymaar9 days ago
Flash[1]: 309B total / 15B activated parameters

Pro [2]:, 1.02T total / 42B activated parameters

[1]: https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL

[2]: https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL

verdverm9 days ago
gandreani9 days ago
Those this mean they've fine-tuned this Qwen 3.5 9B on output from the V2.6 model?
embedding-shape9 days ago
Been playing around with both of these since last night. So far, when enabling reasoning (which is binary on/off), it seems to me like Flash either is less "token efficient" or just likes to think more, or I'm doing something else wrong, because most prompts I send to both, Flash reasons more and for longer than Pro, which is the opposite of my expectations.

Are others seeing the same thing?

the_duke9 days ago
Most of the Chinese models get into long complicated thinking loops before they accomplish something more complicated.

This is also true for Deepseek 4(.1) .

alfiedotwtf9 days ago
Maybe it’s the quant you’re using?
verdverm9 days ago
curious why the HF pill (on the right) always has inaccurate values
wren69919 days ago
Packed 4-bit weights are often identified as u8 byte arrays in the safetensors metadata, so the HF UI counts them as half a weight each.
bopbop98769 days ago
I believe it's because this model is natively fp8 (for the most part), and that display struggles native quants.
stymaar9 days ago
I noticed the same, and I wonder as well.
segmondy9 days ago
more like 500B in FP8
user439289 days ago
I don't trust any of the benchmarks where Opus 5 surpasses Astra or Fable 5.1.

Maybe Terminal Bench 4.0 and ExploitGym are reasonable.

Terminal Bench 4.0

  GPT 6 Astra             59.6
  Claude Fable 5.1        55.1
  Claude Opus 5           49.0
  MiMo-V2.6-Pro           34.9
  MiMo-V2.6-Flash         28.8
  DeepSeek V4.1 Flash     26.8
  MiMo-V2.5-Pro            1.5
ExploitGym

  GPT 6 Astra             42.4
  Claude Fable 5.1        30.4
  Claude Opus 5           22.1
  MiMo-V2.6-Pro           17.8
  MiMo-V2.6-Flash          6.0
  MiMo-V2.5-Pro            0.1
DeepSWE v1.1

  DeepSeek V4.1 Flash     74.2
  Claude Opus 5           74.0
  GPT 6 Astra             74.0
  MiMo-V2.6-Pro           71.9
  Claude Fable 5          70.0
  MiMo-V2.6-Flash         67.9
  MiMo-V2.5-Pro           19.0
dom969 days ago
Why not? In my own benchmark Opus 5 does in fact come out on top[1]

1 - https://bench.killswitch-lang.org/

user439289 days ago
Good question, maybe I am underestimating it based on its absolutely horrible writing style.
HighGoldstein9 days ago
I think we are still far from nailing down good LLM benchmarks, because the more general-purpose your software the harder the question of what makes it good becomes. Is Python a good programming language? Is Java? Is C? I think it's a similar class of problem. You can benchmark rudimentary things like execution speed similar to how you can benchmark tokens/second, but these metrics don't tell the whole story.
3abiton9 days ago
We're past the one model fits them all kind of LLM. Most of the recent release actually regress on world knowledge for example, but optimize for something different: tool usage, thinking process, and agentic approach. And yes, in my own usage, some usecases Opus beats Fable.
mokre9 days ago
Maybe you should not trust any of the benchmarks!
novaleaf9 days ago
can you recommend any benchmark websites that show up-to-date details like this?

TerminaBench, DeepSwe sites are out of date.

UnfitFootprint9 days ago
Yeah it’s a shame a lot of these benchmarks are behind. My favourite was ‘SlopCodeBench’ [1] as I’m most interested in ai reinforcing its own bad decisions, but it’s not even up to current gen oai

1: https://www.scbench.ai/

simonw9 days ago
phainopepla29 days ago
I think we can say pretty confidently they aren't pelican-bench-maxxing
Kurtz798 days ago
A sentence that I would not have expected to read on HN as recently as last year, but that makes perfect sense today.

Jokes asides, @simonw any plan to include a 3D model version (make a 3D model of a Pelican riding a bicycle in Blender)?

Given that Astra seems to have improved a lot in 3D modeling capabilities (and that matches my experience) I'm actually quite interested to see if/when other models catch up and how they stack against it.

Or if anyone knows what would be a useful existing benchmark for that skill.

written-beyond9 days ago
can you update this website, I just wish the entire layout wouldn't shift when the page gets loaded and the timestamps in the title look very ugly and take up a lot of space.
simonw9 days ago
What operating system and device?
tesnorindian9 days ago
A different pelican on a different bicycle direction(R2L) finally. Wondering why MiMo V2.6 Pro choose R2L and Flash choose L2R for the bicycle direction.
ciefa9 days ago
The one with the fish hahaha that's awesome
brcmthrowaway9 days ago
Just me, or do these look bad?

Qwen3.8-27b pelican was amazing on Mac.

https://www.nudgehost.com/dpjn3uwe

knicholes9 days ago
Two legs on one side is a little sus.
idiotsecant9 days ago
Looking terrible isn't nessesarily a bad thing. The pelican is heavily pre trained now. Having a crappy pelican means you didn't try to juke the stats.

Read the full thread on Hacker News →

Related stories