On September 21, 2026, Xiaomi released and open-sourced the MiMo-V2.6 series — three checkpoint families headed by two natively omnimodal sparse-MoE models, published as ungated, MIT-licensed Hugging Face repositories:
- MiMo-V2.6-Pro-RL — flagship: 1.02T total / 42B active parameters, 1M-token context, text/image/video/audio input, text output. (FACT, HF card + AA agree)
- MiMo-V2.6-Flash-RL — efficiency tier: 309B total / 15B active, same 1M-token context and modalities. (FACT)
- MiMo-V2.6-Distill-Qwen-9B — a 9B agentic-RL starting point distilled on MiMo-generated data. (FACT, HF; company describes it as a research entry point)
- Plus 7,000+ high-quality RL task environments (software engineering, vulnerability reproduction, knowledge-intensive work, web design & development), an end-to-end RL training framework (built on verl, uni-agent, mini-swe-agent) and composable mini-harnesses — i.e., Xiaomi opened the training machinery, not just the weights. (COMPANY CLAIM as to coverage/quality; release contents CONFIRMED from official docs)
- The official announcement (dated Sep 22) is titled "MiMo-V2.6: Scaling Up Reinforcement Learning for Self-Improvement," with the blog page meta: "Introducing the MiMo-V2.6 series: frontier intelligence, all the modalities, built in public."
Independent measurement (INDEPENDENTLY VERIFIED): Artificial Analysis scores MiMo-V2.6-Pro at 46 on the Intelligence Index v4.3.2 — the highest score of any open-weights model on its leaderboard, #1 of 114 in its class (large open-weights reasoning models), tied overall with the proprietary Grok 4.7 (xhigh). AA lists the release date as September 21, 2026; API pricing $0.435/M input, $0.87/M output with an unusually deep 99% cache discount, $0.13 per Intelligence Index task (vs ~$3.74 for Grok 4.7), 54.2 t/s, TTFT 2.67 s, 1M context, and a single API provider (Xiaomi's own).
Company-published performance claims (COMPANY CLAIM until replicated): Pro reaches 71.9 on DeepSWE v1.1 (vs 67.9 Flash, 19.0 MiMo-V2.5-Pro, 74.0 Claude Opus 5 / GPT-5.6 Sol, 70.0 Claude Fable 5); on most agent benchmarks Xiaomi says Pro is on par with Claude Opus 5 and GPT-5.6 Sol; AA index "46.32" (per Unite.AI's quote of the announcement); the series "surpasses Kimi K3 and Qwen3.8 Max"; multimodal capabilities across 3D game generation, Blender modeling, Computer Use, embodied control of a Franka Panda arm, materials research (MOF design for PFAS adsorption), and a Lean 4 kernel-verified formalization of Li & Yorke's "Period Three Implies Chaos" (>6,000 lines).
Production RL transparency — the unusual part: Xiaomi livestreams its production RL runs. In under six days the Flash and Pro checkpoints each completed 30 RL steps over ~750,000 trajectories total, at reported costs of ~$0.85M (Flash) and ~$2.62M (Pro); average pass rate on training tasks rose 25% (Flash) and 12% (Pro) and DeepSWE v1.1 (a held-out long-horizon benchmark) improved from 48.8 → 65.7 (Flash) and 58.4 → 72.6 (Pro). (COMPANY CLAIM — figures published in the official announcement; costs not auditable from outside.)
Availability (FACT): API model names mimo-v2.6-pro, mimo-v2.6-flash, mimo-v2.6-pro-ultraspeed (all-lowercase); API prices unchanged from V2.5; an UltraSpeed mode (up to 20× inference speed) at 10× the token price; Xiaomi MiMo Desktop client official release with membership; also in AI Studio, MiMo Code, and OpenRouter, plus ModelScope mirrors.