Xiaomi's MiMo-V2.6-Pro Tops 114 Open Models With a 1T Parameter Beast

Xiaomi's new omnimodal MoE debuts as the top open-weights model on Artificial Analysis, trained with a single mixed RL run for around $2.6M.

·
·
Xiaomi's MiMo-V2.6-Pro Tops 114 Open Models With a 1T Parameter BeastPRO
  • Xiaomi released MiMo-V2.6-Pro-RL, a 1.02T total / 42B active omnimodal MoE under MIT license.
  • Debuts as top open-weights model on Artificial Analysis Intelligence Index at 46.
  • Trained with one unified RL run across coding, agents, visual, and cybersecurity domains.
  • Groupwise agentic grading replaces binary pass/fail rewards with self-referential ranking.
  • Full RL pipeline cost around $2.62M and finished in under six days.
  • Pricing at $0.435/M input, $0.87/M output; supports 1M-token context with FP8 weights.

Xiaomi’s MiMo team released MiMo-V2.6-Pro-RL, a sparse Mixture-of-Experts model with 1.02 trillion total parameters and 42 billion activated for each token. It accepts text, images, video, and audio within a 1 million-token context window. Artificial Analysis gives it an Intelligence Index score of 46, the highest among the 114 open-weight models it tracks, and lists hosted API prices of $0.435 per million input tokens and $0.87 per million output tokens.

The release includes the weights, technical report, training framework, and reinforcement-learning environments under the MIT license. Its technical focus is a unified RL program that Xiaomi completed in under six days while publishing metrics from the final training runs.

One RL run, many skills

Xiaomi calls the method You Only RL Once. The training mix combines coding, visual tasks, computer use, cybersecurity, and several agent harnesses in one run. Harness diversity forms part of the training distribution, with the goal of transferring strategies across domains and into harnesses absent from training.

A fully asynchronous architecture allowed rollout workers to generate experience while training workers updated the model. Each update used 1,568 samples, supported contexts up to 1 million tokens, and processed a reported 3.5 billion to 3.7 billion tokens per step. Xiaomi used Group Relative Policy Optimization, or GRPO, which updates the model by comparing multiple responses to the same task instead of training a separate value model.

Rewards beyond pass or fail

Xiaomi’s agentic grader supplements executable test results with relative judgments about successful trajectories. Two mechanisms provide that additional signal:

  • Groupwise Reward Synthesis (GRS): constructs task-specific rubrics offline from contrasting rollouts, then combines rubric scores with test outcomes.
  • Groupwise Advantage Redistribution (GAR): ranks successful trajectories during training and assigns more update weight to higher-quality responses, favoring shorter paths and lower token use.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads