Xiaomi MiMo-V2.6: Training in Public, and Why the Model Race Is Still Open

A2Agent Team · 2026-09-23T00:00:00Z

Xiaomi has released MiMo-V2.6, a native multimodal agent family in two sizes, Pro and Flash, under the MIT License. The headline number is easy to repeat: MiMo-V2.6-Pro scores 46 on the Artificial Analysis Intelligence Index, the highest among open-source models. The more interesting story is how Xiaomi got there, how fast it got there, and what that says about where the model market is heading.

The quick facts

  • Architecture: 1.02T-parameter MoE with 42B active parameters; text, image, video and audio input; 1M-token context.
  • Training: one mixed reinforcement learning run trains coding, general, visual and cybersecurity agents together. Pro and Flash each completed 30 RL steps in under six days, generating around 750K trajectories.
  • Headline scores (Pro): 46 on the AA Intelligence Index, 71.9 on DeepSWE v1.1, 89.9 on Terminal-Bench 2.1, 82.0 on OSWorld-Verified.
  • What's open: model weights, technical report, RL environments and training code.

1. Training in public is the real innovation

Most labs publish weights and a report after the fact. Xiaomi is also opening the RL environments and training code, and framing the whole release as "built in public." That changes what the community gets. Weights let you use a model. Environments and training code let you study how agentic behavior is produced, reproduce parts of it, and build on it.

This has a demonstration effect. Once one serious player shows that large-scale agentic RL can be opened up without losing its edge, it gets harder for other open-source labs to justify releasing only weights.

It is also very on-brand for Xiaomi. Founder Lei Jun has long built the company's image on transparency and closeness to the user community: letting fans follow along, making the process part of the product story. Opening the training process applies the same playbook to AI. It turns a technical release into a narrative people can follow and share, which buys influence far beyond what a leaderboard score alone would.

2. A fast-rising open-source contender

The generational jump is what makes MiMo worth taking seriously. Comparing MiMo-V2.6-Pro with its predecessor MiMo-V2.5-Pro on the same charts:

Benchmark MiMo-V2.5-Pro MiMo-V2.6-Pro
AA Intelligence Index 26 46
DeepSWE v1.1 19.0 71.9
GDPVal 2.1 (AA) 1107 1673
Toolathlon-verified 49.1 76.9
Terminal Bench 4.0 1.5 34.9
CyberGym 40.0 94.0

That is not incremental tuning. In one generation MiMo went from a mid-table open model to one that sits next to Claude Opus 5, GPT 6 Astra and DeepSeek V4.1 Flash on many agent benchmarks. Flash tracks Pro closely on most of them, which matters for anyone who cares about cost.

Artificial Analysis Intelligence Index v4.3, with MiMo-V2.6-Pro at 46

The gaps are real too, and worth stating plainly. On Terminal Bench 4.0, Pro's 34.9 trails GPT 6 Astra's 59.6. On ProgramBench it scores 26.5 against Claude Opus 5's 37.0. On ExploitGym it posts 17.8 against 42.4. Two of the charts (MiMo Code Bench and MiMo Visual Coding) are in-house benchmarks. On the AA index, 46 still sits below the top closed models at 51 to 53. "On par across most agent benchmarks" is fair; "caught up everywhere" would not be.

MiMo-V2.6 Pro and Flash compared with Claude, GPT and DeepSeek models across coding, general, visual and cyber benchmarks

3. The model landscape is less settled than it looks

It is tempting to read the leaderboard as a done deal: a few frontier labs at the top, everyone else competing for what's left. MiMo-V2.6 is a useful counterexample. A company better known for phones and EVs went from a score of 26 to 46 in one generation and now leads the open-source field on the AA index.

A few things make this kind of jump possible. Reinforcement learning on agent tasks rewards whoever builds the best environments and training loops, not only whoever has the most pretraining compute. Open releases let new players start from a much higher baseline. And companies with deep product ecosystems, whether phones, cars, or home devices, have both the motivation and the real-world use cases to push agentic models hard.

None of this means the current leaders are about to be overtaken. It does mean the ranking is a snapshot, not a verdict. New entrants can still show up and move the line, and anyone building on top of these models should plan for that.

4. Uncertainty is where the opportunity is

If the model layer is still unsettled, everything built on top of it is too. AI applications, agent products, and the whole chain of development that depends on foundation models will keep shifting as new models arrive and old rankings change. The best model for coding agents, computer use or multimodal workflows today may not be the best one six months from now.

That uncertainty is easy to read as risk. We read it as opportunity. A settled market leaves little room for new ideas; an open one keeps rewarding teams that move fast, try new models early, and build products around capabilities that did not exist a generation ago. Each release like MiMo-V2.6 expands what can be built, and each expansion creates room for new applications on top of it. As long as that cycle continues, the opportunity in this industry keeps compounding, and we believe it will keep growing at an exponential pace.

We are very optimistic about where this goes.

Why this matters to A2Agent

This is one of the core reasons we started A2Agent. If no single model is going to win everything, builders should not have to bet their product on one provider. They need a simple way to reach the frontier closed models and the fast-rising open ones through one place, compare them, and switch as the landscape moves.

Releases like MiMo-V2.6 are exactly why that flexibility matters. The next strong model may come from a lab nobody expected, and we want the teams building on A2Agent to be able to use it the day it becomes worth using.

Links