An Unexpected Open Weights Leader: It’s Xiaomi’s MiMo-V2.6-Pro RL and MiMo-V2.6-Flash

Xiaomi, best known for its smartphones and electric vehicles, released the highest-scoring open weights model on Artificial Analysis’ Intelligence Index.

Share
Chart ranks Artificial Analysis models by score; Xiaomi's MiMo-V2.6-Pro-RL leads open weights category.

Xiaomi, best known for its smartphones and electric vehicles, released the highest-scoring open weights model on Artificial Analysis’ Intelligence Index. It performs similarly to GPT-6 Sol, but costs less per task.

What’s new: Along with MiMo-V2.6-Pro-RL and its smaller sibling MiMo-V2.6-Flash, the company also released MiMo-V2.6-Distill-Qwen-9B (Alibaba’s Qwen3.5-9B fine-tuned on MiMo outputs). In addition to the models, Xiaomi also published more than 7,000 reinforcement learning (RL) task environments (software workspaces in which a model attempts tasks and receives feedback) and the code to train on them.

  • Input/output: Text, images, video, and audio in (up to 1 million tokens), text out (up to 128,000 tokens, 44.4 tokens per second)
  • Architecture: Mixture-of-experts transformer; 1.02 trillion parameters, 42 billion active per token for MiMo-V2.6-Pro-RL; 309 billion parameters, 15 billion active per token for MiMo-V2.6-Flash
  • Features: Optional reasoning (on by default), tool calling, optional UltraSpeed mode that generates output up to 20 times faster at 10 times the cost per token
  • Performance: MiMo-V2.6-Pro-RL ranks first among open weights models on Artificial Analysis’ Intelligence Index (46), and MiMo-V2.6-Flash (59.58 percent) ranks first among open weights models on Vals Index, with MiMo-V2.6-Pro-RL (59.47 percent) just behind it, within the margin of error
  • Availability/price: Monthly subscriptions from $6 to $100; MiMo-V2.6-Pro API $0.435/$0.0036/$0.87 per million input/cached/output tokens, UltraSpeed mode $4.35/$0.036/$8.70 per million input/cached/output tokens; MiMo-V2.6-Flash API $0.14/$0.0028/$0.28 per million input/cached/output tokens; batch processing at half the standard price of each model
  • Weights/license: All three models free to download for commercial and noncommercial use under the MIT license
  • Undisclosed: Knowledge cutoff, specific training datasets

How it works: MiMo-V2.6-Pro processes images, video, and audio through separate encoders that feed its language model. According to Xiaomi’s technical report, the training recipe combines methods from earlier work, much of it Xiaomi’s own. The main innovation rewards coding attempts for quality rather than only for passing tasks.

  • Xiaomi pretrained the model first on 27 trillion tokens of text from web pages, books, academic papers, code, and STEM material, then on 3 trillion tokens of text, images, video, and audio for multi-modality. In a mid-training stage, it continued training on records of agents performing coding, visual, and research tasks and extended the input context to 1 million tokens.
  • After a brief round of supervised fine-tuning, the company fine-tuned the model in one RL run that mixed all task types rather than training separate specialists. Coding accounted for 68 percent of the tasks, while design of websites and other visual artifacts (games, 3D scenes, slides, SVGs) accounted for 13 percent; tool use accounted for 12 percent, cybersecurity 4 percent, and context following 3 percent. The RL run used Group Relative Policy Optimization, which generates several attempts at each prompt and scores each one against the others. Each training step drew 1,568 prompts and generated 16 attempts per prompt, for around 25,000 total attempts. The model attempted tasks in variations of one minimal agent harness, configured differently for general, coding, visual, and cybersecurity tasks, so success wouldn’t depend on using a particular harness. The RL run took 30 steps and just over 123 hours.
  • Automated tests that scored coding attempts during training reveal whether code works, not whether it’s well made, so Xiaomi additionally used AI models to grade passing attempts by quality. For coding tasks the model often passed, an agent studied a set of earlier attempts and wrote task-specific checklists for the quality of its approach. During training, each attempt’s reward equaled its test result (1 or 0) multiplied by its two checklist scores. For coding tasks the model did not often pass, a grader model reviewed all 16 attempts at a task together, inspected the repository, and ran tests as needed. It ranked passing attempts by criteria such as the approach’s suitability, how few changes it made, and consistency with the codebase’s conventions. Training then reinforced higher-ranked attempts more strongly and lower-ranked ones less.
  • To counter reward hacking, such as downloading a published solution instead of writing a fresh one, the team scrubbed leftover answers from task environments, blocked network access, and ran an agent that searched for exploits until it found none. Attempts confirmed to use a leaked answer received zero reward, the same as a failed attempt, and they stayed below 2 percent of attempts throughout training.
  • The company trained separate teacher models on tasks whose results are hard to check automatically, such as game development, scientific research, and embodied intelligence (controlling robots and other physical systems). Then it trained MiMo-V2.6-Pro to imitate them through on-policy distillation, in which the student learns to align its output with the teacher’s prediction of which token should come next.

Performance: MiMo-V2.6-Pro topped open weights models on two independent composite evaluations, making a large gain over its predecessor. It matched some proprietary models at a small fraction of their cost per task, but it trailed the leaders and generated output slowly.

  • On Artificial Analysis’ Intelligence Index v4.3.2, a weighted average of 10 evaluations across math, science, coding, and reasoning, MiMo-V2.6-Pro set to reasoning (46, $0.13 and 19.5 minutes per task) led open weights models, ahead of GLM-5.3 set to max reasoning (45, $2.01 and 10.7 minutes per task) and Kimi K3 set to max reasoning (44 and $2.00 per task). Its predecessor, MiMo-V2.5-Pro, scored 26. Among proprietary models, it tied Grok 4.7 set to high reasoning (46, $2.73 and 14.2 minutes per task) and trailed several models from Anthropic, Meta, and OpenAI.
  • On the Vals Index, which weights seven agentic benchmarks of finance, coding, and legal work by each sector’s share of U.S. GDP, the smaller MiMo-V2.6-Flash set to reasoning (59.58 percent accuracy, $0.20 per test, 45.03 minutes) and MiMo-V2.6-Pro set to reasoning (59.47 percent accuracy, $0.39 per test, 52.33 minutes) outperformed all other open weights models tested. Both trailed 15 proprietary models, led by Claude Opus 5.5 set to max reasoning (69.69 percent accuracy, $22.30 per test, 72 minutes).
  • On Vals’ CyberBench v1.1, which tests whether agents can reproduce and patch vulnerabilities in open-source software, MiMo-V2.6-Flash set to reasoning (75.36 percent accuracy, $0.05 per test, 32.48 minutes) ranked first of nine models and MiMo-V2.6-Pro set to reasoning (72.86 percent accuracy, $0.09 per test, 31.73 minutes) ranked third, ahead of Muse Spark 1.3 Max (72.74 percent accuracy, $3.34 per test, 19.75 minutes) and Claude Fable 5.1 (70.42 percent accuracy, $4.03 per test, 22.12 minutes).

Behind the news: Recently, Anthropic alleged that, over 20 days in March and April 2026, Xiaomi routed MiMo users’ conversations and coding sessions to Claude through the coding tools OpenClaw and OpenCode, producing more than 400,000 exchanges to use as training data. To date, Xiaomi has not publicly responded to Anthropic’s accusation.

Why it matters: Code that passes tests isn’t necessarily good code. Xiaomi found that a version of its model that was trained without the code quality grader picked up habits that make software harder to maintain: The model added code the task didn’t call for, let errors pass silently, and were lax in checking on incoming data until tests passed. Trained with the code quality grader, the model made smaller, more focused changes. Rewarding code that a human reviewer would accept is as important as rewarding code that runs.

We’re thinking: Open weights are great, but open recipes are even better. By releasing its reinforcement learning tasks, the software that runs them, and its training code, Xiaomi lets others repeat the steps that produced much of its models’ improvements. It also disclosed what those steps cost: $2.6 million for MiMo-V2.6-Pro and $0.9 million for MiMo-V2.6-Flash. Teams planning to train agents get a set of tasks they don’t have to build and a rough price for reinforcement learning at this scale. We hope other labs can reproduce and extend on what Xiaomi learned training these models.