News Weekly
LV 10 XP
0% read
Your progress · 0/5 chapters
About 6 min total
ModelsISSUE #2 · STORY 9 OF 20Sep 21, 2026CONFIRMED

Xiaomi's new open AI model tops the open-weight rankings

Xiaomi's MiMo-V2.6-Pro scored 46 on Artificial Analysis' Intelligence Index, the highest ever measured for an open-weights model there. Xiaomi also released its training environments, framework and harnesses under the permissive MIT license.

Illustration: an enormous translucent iceberg-vault floats over a calm abstract sea, filled with countless tiny glowing modules — an artistic impression of a trillion-parameter open model released under a permis…

Read it your way

CHAPTER 1 · THE 60-SECOND VERSIONPicked for Explorers

The open ceiling moved up

On Sept 21, 2026, Xiaomi released MiMo-V2.6 Pro and Flash as ungated, MIT-licensed models on Hugging Face. Artificial Analysis scored Pro at 46 on its Intelligence Index, the top open-weights result on that leaderboard. The closed leaders, Claude Fable 5.1 and GPT-6 Astra, still lead the overall index at 53.

The measured scorePro scores 46 on Artificial Analysis, ahead of GLM-5.3 at 45 and Kimi K3 at 44 among open weights.
Not the overall leaderXiaomi's own docs concede the 53 lead for closed models, so 46 is the open-field top, not all of AI.
The full toolkitThe release includes weights plus 7,000 training environments, an RL framework and mini-harnesses, per official docs.
Flash is practicalThe 309B Flash variant with 15B active parameters is the realistic self-hosted pick; Pro needs multi-node clusters.
Finish this chapter for +15 XP
Flip the switch

How the open ceiling moved

YOU GETA new ceiling: 46MiMo-V2.6-Pro tops open weights and ties the proprietary Grok 4.7 on the same index.
YOU GETA reproduction kitEnvironments, framework, harnesses and a technical report invite independent rebuilding.
YOU GETA reset cost lineAA measures $0.13 per index task, one to two orders below top closed models.
Play with the numbers · +10 XP

What an open-weight agent workload costs you

DRAG THE SLIDER
MiMo-V2.6 Pro
$260
$0.13 per task, per Artificial Analysis
Grok 4.7
$7,480
$3.74 per task, per Artificial Analysis
You'd save with MiMo
$7,220
every month

Per-task costs measured by Artificial Analysis. Xiaomi's own benchmark wins are company claims awaiting independent reruns.

Your next move · as a Explorer

Test the open frontier

1Check the Hugging Face repos and confirm they are ungated and MIT.
2Rebuild the AA open-weights rows and the per-point cost math yourself.
3Benchmark Flash on three of your own coding tasks if you have a GPU.

Switch your reading mode at the top to see a different next move.

Tap to open

Things to keep an eye on

Pop quiz · unlock the MIT Maven badge

Did it stick?

0/3
What does MiMo-V2.6-Pro's score of 46 represent on Artificial Analysis?+20 XP
What did Xiaomi open-source besides the weights?+20 XP
Which license covers the MiMo-V2.6 models?+20 XP
Your call · +5 XP

With open weights at 46 and $0.13 per task, does a closed model at 53 still deserve its premium?

Deep dive

The full research, labeled and sourced

CONFIRMED11 sources · 78 min
Story identity
FieldValue
Story IDS09
TitleXiaomi open-sources MiMo-V2.6 Pro & Flash (MIT): 1T-param flagship tops open-weight rankings at 46
OrganizationXiaomi (Xiaomi MiMo team; official brand "Xiaomi MiMo")
CategoryOpen-Weight Model
Event date2026-09-21 (weights live on Hugging Face per HF repository creation timestamps 2026-09-21T15:39 UTC; Artificial Analysis lists release date as September 21, 2026)
Announcement date2026-09-22 (official docs news page, "Update Time September 22, 2026"; official blog page mimo.xiaomi.com/mimo-v2-6)
Article dates2026-09-21 (Unite.AI, CCLeaks); 2026-09-22 (official announcement/docs)
Window check2026-09-21 ∈ [2026-09-18, 2026-09-22] — eligible
Evidence statusCONFIRMED (official announcement + docs fetched in full; both HF model cards fetched in full; HF API + README YAML checked; Artificial Analysis model page and open-source leaderboard page fetched in full; two independent outlets fetched)

Corrections / refinements to the discovery record (important):

  • "Tops open-weight rankings at 46" — INDEPENDENTLY VERIFIED. Artificial Analysis' open-source model comparison page (fetched 2026-09-23) lists MiMo-V2.6-Pro at 46 on the Intelligence Index v4.3.2, ahead of GLM-5.3 (max) at 45 and Kimi K3 (max) at 44 — first place among open-weights models. The prose on that page states: "MiMo-V2.6-Pro and GLM-5.3 (max) are the highest intelligence open source models." The caveat: 46 is the top of the open-weights field, not the overall index — closed models Claude Fable 5.1 and GPT-6 Astra lead at 53 (per Xiaomi's own docs and S03 research). Grok 4.7 (xhigh) also scores 46 but is proprietary, not open-weights.
  • "Docs 'Update Time September 21'" — actual docs page reads "Update Time September 22, 2026." The event date is still 2026-09-21: Hugging Face repository createdAt timestamps (2026-09-21T15:39:33Z for Pro, 15:39:51Z for Flash, 18:18:40Z for Distill-9B) and Artificial Analysis both fix the weights going live on Sep 21; the official announcement text is dated Sep 22. Both dates are inside the window. No eligibility impact.
  • Speed figures are time-sensitive and inconsistent across the public record. AA's live model page (researched 2026-09-23) shows 54.2 output tokens/s (below the class median of 67.1) and TTFT 2.67 s; Unite.AI's Sep 21 report cited AA at 129.7 t/s / 2.17 s TTFT; a secondary router post (OrcaRouter, cited in S03 sources) cited 134.3 t/s. AA's page figures reflect the latest measurement window (possibly post-UltraSpeed deployments). This research uses the live AA page figures as the current record and flags the spread.
  • Flash's HF page label says "311B"; the model card's Model Summary table and AA say 309B total / 15B active. Both the page label field and the card table appear in the same repo; the card table (309B/15B) is treated as canonical, with the 311B being a display-rounding artifact. Similarly Pro is labeled "1T" on the page and 1.02T in the card table. Trivial, but worth pinning before citing.
  • "Surpasses Kimi K3 and Qwen3.8 Max" is a company claim that is only partially independently corroborated. AA independently shows MiMo-V2.6-Pro (46) ahead of Kimi K3 (max) (44) — corroborated. AA's open-source page lists "Qwen3.8 27B (xhigh)" at 34, which MiMo beats, but the exact "Qwen3.8 Max" variant named by Xiaomi does not appear on AA's open-source leaderboard — the surpass claim against that specific naming is COMPANY CLAIM.
  • API license metadata quirk: the HF REST API api/models/... payload returns license: None for all three repos even though the rendered model-card pages and the README YAML front-matter both declare license: mit, and AA independently records the MIT license. MIT is CONFIRMED (rendered card + README YAML + AA); the API-field inconsistency is a Hugging Face metadata-storage quirk to be aware of when scripting license audits.

✓

What happened?

🎓 For Explorer

On September 21, 2026, Xiaomi released and open-sourced the MiMo-V2.6 series — three checkpoint families headed by two natively omnimodal sparse-MoE models, published as ungated, MIT-licensed Hugging Face repositories:

  • MiMo-V2.6-Pro-RL — flagship: 1.02T total / 42B active parameters, 1M-token context, text/image/video/audio input, text output. (FACT, HF card + AA agree)
  • MiMo-V2.6-Flash-RL — efficiency tier: 309B total / 15B active, same 1M-token context and modalities. (FACT)
  • MiMo-V2.6-Distill-Qwen-9B — a 9B agentic-RL starting point distilled on MiMo-generated data. (FACT, HF; company describes it as a research entry point)
  • Plus 7,000+ high-quality RL task environments (software engineering, vulnerability reproduction, knowledge-intensive work, web design & development), an end-to-end RL training framework (built on verl, uni-agent, mini-swe-agent) and composable mini-harnesses — i.e., Xiaomi opened the training machinery, not just the weights. (COMPANY CLAIM as to coverage/quality; release contents CONFIRMED from official docs)
  • The official announcement (dated Sep 22) is titled "MiMo-V2.6: Scaling Up Reinforcement Learning for Self-Improvement," with the blog page meta: "Introducing the MiMo-V2.6 series: frontier intelligence, all the modalities, built in public."

Independent measurement (INDEPENDENTLY VERIFIED): Artificial Analysis scores MiMo-V2.6-Pro at 46 on the Intelligence Index v4.3.2 — the highest score of any open-weights model on its leaderboard, #1 of 114 in its class (large open-weights reasoning models), tied overall with the proprietary Grok 4.7 (xhigh). AA lists the release date as September 21, 2026; API pricing $0.435/M input, $0.87/M output with an unusually deep 99% cache discount, $0.13 per Intelligence Index task (vs ~$3.74 for Grok 4.7), 54.2 t/s, TTFT 2.67 s, 1M context, and a single API provider (Xiaomi's own).

Company-published performance claims (COMPANY CLAIM until replicated): Pro reaches 71.9 on DeepSWE v1.1 (vs 67.9 Flash, 19.0 MiMo-V2.5-Pro, 74.0 Claude Opus 5 / GPT-5.6 Sol, 70.0 Claude Fable 5); on most agent benchmarks Xiaomi says Pro is on par with Claude Opus 5 and GPT-5.6 Sol; AA index "46.32" (per Unite.AI's quote of the announcement); the series "surpasses Kimi K3 and Qwen3.8 Max"; multimodal capabilities across 3D game generation, Blender modeling, Computer Use, embodied control of a Franka Panda arm, materials research (MOF design for PFAS adsorption), and a Lean 4 kernel-verified formalization of Li & Yorke's "Period Three Implies Chaos" (>6,000 lines).

Production RL transparency — the unusual part: Xiaomi livestreams its production RL runs. In under six days the Flash and Pro checkpoints each completed 30 RL steps over ~750,000 trajectories total, at reported costs of ~$0.85M (Flash) and ~$2.62M (Pro); average pass rate on training tasks rose 25% (Flash) and 12% (Pro) and DeepSWE v1.1 (a held-out long-horizon benchmark) improved from 48.8 → 65.7 (Flash) and 58.4 → 72.6 (Pro). (COMPANY CLAIM — figures published in the official announcement; costs not auditable from outside.)

Availability (FACT): API model names mimo-v2.6-pro, mimo-v2.6-flash, mimo-v2.6-pro-ultraspeed (all-lowercase); API prices unchanged from V2.5; an UltraSpeed mode (up to 20× inference speed) at 10× the token price; Xiaomi MiMo Desktop client official release with membership; also in AI Studio, MiMo Code, and OpenRouter, plus ModelScope mirrors.


Δ

What changed?

  • The strongest open-weights release ever measured on AA's Intelligence Index landed — with the training stack attached. MiMo-V2.6-Pro's 46 beats the previous open-weights frontrunners (GLM-5.3 max 45, Kimi K3 max 44) on an independent leaderboard, and Xiaomi open-sourced not only MIT weights but 7,000+ RL environments, the RL framework, and mini-harnesses. (FACT on the score; scope of the open-sourced machinery per official docs, CONFIRMED)
  • The open/closed intelligence gap on AA's composite index closed to one point for the top open model vs the top proprietary field ordering: open-weights 46 vs Claude Fable 5.1 / GPT-6 Astra at 53 — and MiMo ties the proprietary Grok 4.7 at 46. The "open weights are one generation behind" assumption no longer holds on this composite. (INTERPRETATION grounded in AA data)
  • Licenses shifted: the earlier MiMo-V2-Flash repo was Apache-2.0; the V2.6 cards are MIT (FACT: card YAML license: mit), removing even attribution friction for commercial derivatives.
  • RL training economics went public and quasi-reproducible: ~$2.62M for a 6-day, 30-step production RL run on the flagship, with the environment/framework code released so others can reproduce the loop — a step-change in transparency for frontier-scale RL. (COMPANY CLAIM on the costs; reproducible via the released framework — EARLY RESEARCH)
  • A new price-performance reference point for open weights: AA measures $0.13 per Intelligence Index task for Pro at $0.435/$0.87 per M with a 99% cache discount — far below the ~$3.7/task of Grok 4.7 and ~$3.3/task of GPT-6 Astra Max, i.e., per-intelligence-point cost collapsed by ~1–2 orders of magnitude. (INDEPENDENTLY VERIFIED on the AA numbers; the cross-vendor comparison is arithmetic on AA data)
  • Hardware-adjacent pressure: a 1T-parameter MoE with 1M-token training at 3.5–3.7B tokens/step (company claim) raises the baseline for what "open-source AI" requires to run — and makes Flash (15B active) the practical on-prem entry point.

↔

Before → Change → After

🎓 For Explorer

Before (as of mid-September 2026):

  • Open-weights leadership on AA's Intelligence Index: GLM-5.3 (max) at 45, Kimi K3 (max) at 44 — no open model had crossed 46; MiMo-V2.5-Pro sat at 19.0 on DeepSWE v1.1 with no competitive agentic profile (FACT: Xiaomi's own before/after table).
  • Open-weights releases typically shipped weights + a card + maybe a small eval harness; training environments and RL code stayed proprietary. (INTERPRETATION, widely observable)
  • Xiaomi MiMo's earlier open releases (MiMo-V2-Flash) were Apache-2.0, and V2.5 was open-sourced with an "Orbit 100 trillion token plan" (from the docs news index).
  • Closed frontier: Claude Fable 5.1 / GPT-6 Astra at 53 holding a ≥7-point edge over any open model; pricing war in full swing (Grok 4.7 same week at $2/$6; GPT-6 Sol/Luna the next day at half price).

Change (Sep 21–22, 2026):

  • MiMo-V2.6-Pro-RL (1.02T/42B, MIT, ungated) hits 46 on AA — #1 open-weights score, tied overall with proprietary Grok 4.7; Flash (309B/15B) ships the same 1M context and modalities at 10× lower token price than Pro.
  • Xiaomi releases weights plus 7,000+ RL environments, an end-to-end RL framework, mini-harnesses, a technical report, and a Distill-9B starting point, with a livestreamed 6-day production RL run documented end-to-end.
  • API pricing unchanged from V2.5; UltraSpeed (≤20×) launched; Desktop client leaves beta; Distill-9B demonstrates 61.1 → 66.2 on SWE-bench Verified after RL (company claim).

After:

  • For the first time, the top open-weights model and a top proprietary model sit at the identical composite index score (46/46) out of 673 tracked models, while the overall lead (53) remains closed — so the open-vs-closed narrative splits into "parity at the top of open" vs "overall lead still closed."
  • Every lab's next release must now answer "open at 46 with the RL stack attached" — either on score, on price-per-task ($0.13), on license (MIT), or on transparency (livestreamed RL runs). (PREDICTION)
  • Developer default deployment choice for cost-sensitive agentic work shifts toward open weights: MiMo-V2.6-Flash (15B active, $0.14/$0.28) becomes a credible self-hosted default, with Pro as the moat-level option. (INTERPRETATION)

⚙

How it works

  • Architecture (FACT, HF cards): both checkpoints are sparse MoE on a hybrid sliding-window/global-attention backbone with "You Only RL Once" mixed training. Pro: 70 layers (60 SWA + 10 GA), hidden 6144, 384 routed experts with 8 active per token, SWA window 128, 1M max context. Flash: 48 layers (39 SWA + 9 GA), hidden 4096, 256 experts / 8 active. First block uses global attention with a dense FFN; remaining blocks interleave local SWA and GA with sparse MoE FFNs and no shared experts.
  • Omnimodal I/O (FACT): a 681M-param MiMo ViT (28 layers, 24 SWA + 4 full; patch 2×16×16) for vision; audio via a 308M-param AudioTokenizer (20 RVQ codebooks) plus a 127M-param audio patch encoder (25 Hz → 6.25 Hz); text/image/video/audio input, text output; 1M-token context (about 1,500 A4 pages per AA).
  • Speculative decoding (FACT): a 5-layer SWA multi-token-prediction drafter predicts 7 subsequent tokens per forward pass for parallel verification (EAGLE-style recipes in the cards).
  • RL recipe (COMPANY CLAIM as to efficacy; mechanics from cards/docs): fully asynchronous GRPO on huge batches — 1,568 prompts × 16 rollouts per step, billions of tokens per update, 1M-context training with 3.5–3.7B tokens per step. Grading is itself scaled: Groupwise Reward Synthesis (GRS) builds task-specific rubrics offline by contrasting rollouts; Groupwise Advantage Redistribution (GAR) ranks passing trajectories online — together forming a self-improvement loop that steers toward shorter paths and fewer tokens.
  • Stability/anti-hacking measures (COMPANY CLAIM): the MoE router is frozen during RL to suppress expert load drift; a reward-hacking defense spans reward design, adversarial evaluation, anomaly detection, and verifier cross-checks.
  • Multi-harness design (COMPANY CLAIM): tasks and harnesses (code, general agents, visual, cyber) are mixed in the same batch so capabilities reinforce each other; lightweight composable "mini-harnesses" decouple system prompts, tools, and context management, and Xiaomi claims generalization to harnesses never seen in training.
  • Distillation (COMPANY CLAIM): MOPD2 (Multi-Prefix Multi-Teacher On-Policy Distillation) follows the mixed RL run, reusing teacher/SFT prefixes so decision points train without regenerating preceding turns — extending RL gains to hard-to-verify tasks; the Distill-Qwen-9B artifact applies this to a small startable checkpoint.
  • Serving (FACT): SGLang and vLLM recipes published on the model cards (Docker images, tensor-parallel/EP layouts); also served via Xiaomi's API (OpenAI-compatible), OpenRouter, AI Studio, MiMo Code, and Desktop; UltraSpeed mode at up to 20× token speed at 10× price.

!

Why it matters

🎓 For Explorer
  1. The open-weights ceiling moved — and the gap to closed models narrowed dramatically on a respected composite. AA 46 (#1 open, tied with proprietary Grok 4.7; closed leaders at 53) is the strongest independently measured open-weights result to date. The "open weights are always a generation behind" claim is no longer true at the top of the open field. (FACT on scores; the claim about what it means is INTERPRETATION)
  2. Xiaomi open-sourced the reproduction kit, not just the model. 7,000+ RL environments, the verl-based end-to-end framework, mini-harnesses, a technical report, and a 9B starting point mean other labs and enterprises can rebuild the training loop — this is a public-good move that could accelerate open agentic-RL research globally. (FACT on the release scope; direction of effect is PREDICTION)
  3. Per-task cost collapse for open agentic AI: $0.13 per Intelligence Index task (vs $3.74 Grok 4.7, ~$3.3 GPT-6 Astra Max) resets the economics of open-weights agentic workloads — with a 99% cache discount that rewards long-context/RAG patterns. (INDEPENDENTLY VERIFIED numbers; the reset is INTERPRETATION)
  4. The pricing war now has an open-weights supply side: same week as Grok 4.7 ($2/$6) and GPT-6 Sol/Luna (half price), Xiaomi shipped MIT weights at $0.435/$0.87 (Pro) and $0.14/$0.28 (Flash) with unlimited commercial use — "frontier for less" is now a three-company reality plus a free-for-use open layer. (FACT)
  5. Transparency as a competitive weapon: livestreamed production RL, published run costs, and full training-code release is a genuinely different posture from closed labs — it converts Xiaomi's AI research into a trust signal and a recruiting/ecosystem asset. (INTERPRETATION)

✦

What became possible?

🎓 For Explorer
  • Self-hosting a top-of-open-weights reasoning model at 1M context with native text/image/video/audio input — no API dependency, no per-token bill, MIT terms — for organizations that can provision the hardware (Pro) or that pick Flash (15B active) for feasible on-prem inference.
  • Reproducing and extending a frontier-scale agentic RL loop: the released environments (SW engineering, vulnerability reproduction, knowledge-intensive work, web design/dev) + framework + mini-harnesses give research teams a starting point they previously had to assemble from scratch (months of work). (COMPANY CLAIM on quality/completeness; contents CONFIRMED)
  • Running long-horizon agent workloads (1M context; per AA "~1,500 A4 pages") on open weights with a 99% cache discount — RAG over large codebases, multi-session agent runs, tool traces.
  • Building commercial products on MIT weights with zero license friction (no attribution clause, unlimited redistribution/modification), including fine-tunes, distillation, and serving businesses.
  • Small-team agentic-RL research via Distill-Qwen-9B: a 9B checkpoint documented to gain +5.1 SWE-bench Verified and +15.7 MiMo Cyber Bench from RL over its SFT baseline (company claims) — a startable, GPU-friendly training target.
  • Benchmarking openness: AA's Openness Index components and license data now include a model that is both top-intelligence and maximally permissive — a reference point for procurement "openness" requirements.

◎

Implications

Technical

  • Benchmark standings are time-sensitive: AA's page shows 54.2 t/s and TTFT 2.67 s today, while early press cited 129.7 t/s / 2.17 s — speed and latency figures move as serving is tuned (UltraSpeed deployment, batching). Always capture the date and measurement window when citing AA numbers.
  • 1T-parameter open weights raise the hardware bar: Pro (BF16) essentially requires multi-node tensor/expert parallelism (the card's SGLang recipe assumes --tp 16 --dp 2 --nnodes 2 with EP 16); the practical self-hosted path is Flash (309B total / 15B active, card assumes tp 8) or quantized builds (28 community quantizations already listed for Flash).
  • 1M-context RL training at 3.5–3.7B tokens/step (company claim) is a major systems-engineering data point — attention/communication overhead at that scale is exactly where most labs fail; Xiaomi publishing the framework implies those patterns are now diffusable.
  • MoE router freezing + groupwise agentic grading are concrete, adoptable techniques for stabilizing long-horizon RL — the technical report and code make them falsifiable.
  • MIT on a frontier-scale model resets license negotiation norms for open releases (Apache-2.0 was already common; MIT removes copyleft/notice friction entirely for commercial derivatives).
  • Multi-harness generalization (training across harnesses, claiming transfer to unseen ones) is a strong claim that independent re-runs should probe — currently EARLY RESEARCH/COMPANY CLAIM.

Developer

  • Add MiMo-V2.6-Flash as a default candidate for self-hosted agentic coding — 15B active, 1M context, MIT: vLLM/SGLang one-command recipes on the card (vllm serve XiaomiMiMo/MiMo-V2.6-Flash-RL --tensor-parallel-size 4); sample with temperature 1.0, top_p 0.95 per the card.
  • Pro is a frontier-level option if you have the cluster (tp 16 / 2 nodes per the recipe, or via the API at $0.435/$0.87 with 99% cache discount) — benchmark per-task cost, not sticker price; AA's $0.13/index-task figure is the reference.
  • Pin revisions and license metadata in CI: the HF API license field returns None for these repos even though the card YAML declares MIT — audit from the README YAML or rendered card, and pin the commit; also note the 311B-page-label vs 309B-card-table discrepancy for Flash.
  • UltraSpeed costs 10× per token ($4.35/$8.70 per M for Pro) — only for latency-critical/high-throughput paths; route selectively (e.g., interactive agent loops, not batch jobs).
  • Cache economics are exceptional: 99% cache-read discount (Pro $0.0036/M cached) makes long-context, repeated-prefix workloads (agents, RAG over codebases) dramatically cheaper than competitors — design prompts to maximize cache hits (stable system prompts, conversation reuse).
  • Tooling maturity is high at release: transformers (trust_remote_code), vLLM and SGLang recipes, OpenRouter availability, desktop client, ModelScope mirrors; --reasoning-parser mimo and --tool-call-parser mimo in both major servers mean reasoning/function-call parsing is first-party.
  • Watch the UlTraSpeed/context tradeoffs — 1M-token context with hybrid SWA+GA attention; measure real long-context quality on your docs before trusting the ceiling (AA lists 1M, no degradation data on the page).

Enterprise

  • A credible open-weights alternative for agentic/knowledge workloads: enterprises wary of API lock-in can self-host Flash-class models on-prem or in VPC (15B active fits modest multi-GPU nodes) with MIT terms — renegotiate vendor portfolios that currently assume closed-API-only agentic AI.
  • Cost/supply lever in the current pricing war: AA's $0.13 per index task (Pro via API) is an order of magnitude below closed per-task costs; an internal benchmark set (tokens/task, pass rate, cost) should include MiMo-V2.6 Pro and Flash immediately.
  • Data-governance win: on-prem/self-hosted open weights avoid the data-residency questions inherent in API use for regulated industries (though the model's training provenance and the license still need legal review — MIT covers use, not data lineage).
  • RL-infrastructure adoption risk: the released framework/verl-based stack is a research-grade asset; productionizing it requires MLOps investment — don't treat it as turnkey; treat it as a strong starting point (consulting opportunity).
  • Compliance context: vendor-published safety/refusal metrics and open-weights auditability (weights, code, technical report) become input to the documentation enterprises increasingly need under new AI governance frameworks (e.g., the Sep 18–19 US EO week); MIT + full release is the strongest auditable posture published this window.
  • Deployment reality check: Pro-grade hardware requirements are real; most enterprises will adopt Flash or quantized builds first, with Pro via API.

Strategic

  • Xiaomi moves from phone/hardware AI to a frontier open-weights lab — with a strategy others can copy. The formula: train on verifiable tasks with scaled RL, open the weights under MIT, livestream the run, publish costs, and profit from API/UltraSpeed/desktop/claw/kg ecosystem instead of model secrecy. (INTERPRETATION; components CONFIRMED)
  • The open-weights ceiling (46) now abuts the closed-second-tier (Grok 4.7 at 46) and trails only the closed leaders (53). For customers, the question shifts from "open or closed?" to "how much is +7 index points worth in vendor risk and price?" (INTERPRETATION)
  • The RL-supply-chain play: by open-sourcing environments + framework + Distill-9B, Xiaomi seeds a self-reinforcing ecosystem (researchers train on Xiaomi's stack → improvements flow back into Xiaomi's platform via Orbit/claw/desktop) — cheaper than winning head-to-head benchmark wars, and defensible against closed labs. (INTERPRETATION)
  • China-lab competition is now a three-way open-weights race (Xiaomi MiMo, Z AI GLM-5.3, Kimi K3 at 44/45/46) — the top of the open field is dominated by Chinese labs, with US labs' open entries (Meta/NVIDIA/Thinking Machines) trailing at 17–25 on AA's open-source page. (FACT, AA leaderboard)
  • Expect closed labs to respond with even steeper budget tiers, exclusive harness integrations, or open releases of their own smaller models; open-weights pricing may now set the floor for agentic workloads. (PREDICTION)

⚠

Risks & limitations

Risks
  • Company-claimed benchmark supremacy is not third-party verified on most rows. The DeepSWE/agent tables and design leaders are self-published; AA verified only the composite index behavior. Independent re-runs (SWE-bench Verified, Terminal-Bench, CyberGym) are needed before procurement-grade conclusions. (COMPANY CLAIM status)
  • Hardware-access barrier creates a two-tier open ecosystem: only well-resourced teams can run 1T/42B Pro; Flash (15B active) narrows but does not close that gap; smaller players may be locked into API pricing that quietly drifts. (INTERPRETATION)
  • 1M-context claims at production scale can hide degradation, memory blowups, or extreme TTFT under load — measure on your own data before promising long-context SLAs. (INTERPRETATION; AA's TTFT 2.67 s suggests prefill pressure)
  • RL-stack adoption risk: released training code on top of verl/uni-agent/mini-swe-agent is young; bugs, harness quirks, and undocumented environment behavior are likely — treat as research-grade, not production-grade.
  • Reward-hacking/dual-use exposure: open-sourcing cybersecurity environments and data (vulnerability reproduction, ExploitGym-style tasks on the card) plus MIT weights gives adversaries a free training ground for offensive capabilities — a genuine dual-use consideration for enterprises and regulators. (INTERPRETATION; the open-sourcing facts are CONFIRMED)
  • Supply-chain/trust risk: ungated, MIT, custom-code models demand the usual rigor — pin exact revisions, scan trust_remote_code payloads, and note that the top-level HF API metadata is inconsistent (license: None) which can break automated license-audit pipelines.
  • Provider concentration: AA shows a single API provider (Xiaomi) for hosted Pro; if Xiaomi's platform has an incident or price change, enterprises are exposed unless they self-host or wait for third-party providers (currently none listed on AA's page).

Limitations
  • Not the overall intelligence leader: 46 vs 53 for Claude Fable 5.1 / GPT-6 Astra; Xiaomi's own docs concede the gap; the "tops rankings" claim is specifically open-weights rankings.
  • Below-class-average output speed (54.2 t/s vs 67.1 median per AA's current page) and elevated TTFT (2.67 s vs 2.32 s median) — deliberative, not interactive-streaming, unless UltraSpeed (at 10× price) is used.
  • One API provider on AA's tracker at research time; no third-party hosted options listed yet.
  • All flagship benchmark numbers are vendor-published (company tables); only AA's composite/price/speed measurements are independent.
  • Pro is impractical for most self-hosters (multi-node EP/TP required); the released SGLang recipe assumes 2+ nodes.
  • No published open-weights model card from Xiaomi for the V2.6-Pro base (non-RL) variant — only the RL checkpoints are released; AA notes "a non-reasoning variant may also exist" from other sources.
  • Multimodal claims are architectural, not fully benchmarked: video/audio input support is stated; independent multimodal evals (MMMU-Pro etc.) were not on the fetched pages.
  • Verbal/UX claims (vibe world, music, films) are demos and marketing cases, not standardized measurements — treat as COMPANY CLAIM/illustrative.

?

Open questions

  1. Will independent re-runs confirm DeepSWE 71.9 (Pro), CyberGym 94.0, and the agent-table numbers — and where do they land on AA's Coding Agent Index once measured (not yet on the fetched pages)?
  2. Do the claimed RL-economics (6 days, 30 steps, ~750K trajectories, $2.62M for Pro) hold up to reproduction with the released framework — and what cluster was required?
  3. Why does the HF API return license: None while cards/README/AA all say MIT — metadata migration bug or intentional flag? Will automated license audits be misled?
  4. Is UltraSpeed the documented 129.7 t/s readout (early press) vs the current 54.2 t/s default — and does the 20× claim hold at 1M context?
  5. When will third-party providers (Fireworks/Together/Baseten-class) host Pro/Flash, and at what prices relative to Xiaomi's $0.435/$0.87?
  6. Will Xiaomi release the base (non-RL) V2.6 checkpoints or a V2.6-Pro "non-reasoning" variant for latency-sensitive serving?
  7. What do independent safety/red-team evaluations of the released cyber environments conclude — and does the open cyber kit attract researcher scrutiny or misuse first?
  8. Does the Distill-Qwen-9B recipe (SWE-bench 61.1→66.2, MiMo Cyber Bench 31.3→47.0) reproduce on other 9B bases, and does it generalize outside MiMo's environment suite?
  9. What is the roadmap cadence now — MiMo-V2.7? Integration with MiMo Code/Claw/Orbit ecosystem — and does the API pricing really stay flat?

↗

What happens next?

🎓 For Explorer
  • Days: AA leaderboard refresh (open-source + overall) as more of the 673-model index runs land for Pro/Flash; possible third-party provider announcements; community re-runs of DeepSWE/agent tables begin (early HF discussion threads exist — 8 discussions on Pro, 7 on Flash).
  • Weeks: first independent replication attempts of the RL recipe (environments + framework are fresh downloads); Xiaomi Orbit/Claw/Desktop ecosystem integrations; UltraSpeed beta ends (invitation-only program runs "one more week" per the docs) — watch pricing stability after the beta window; possible base/non-RL variant release.
  • Months: open-weights leaderboard consolidation (MiMo 46 vs GLM-5.3 45 vs Kimi K3 44) as each lab replies; closed-lab budget-tier responses to the $0.13/task cost line; Q4 procurement cycles incorporating MIT-open-weights options; and at some point MiMo-V2.7-era RL-scale transparency becoming the norm for major open releases (PREDICTION).

★

Editorial takeaway

🎓 For Explorer

Xiaomi's MiMo-V2.6 release is the strongest open-weights story of this window for two independent reasons: score — 46 on Artificial Analysis' Intelligence Index, the #1 open-weights result ever measured there, verified by the benchmark owner itself, tied with the proprietary Grok 4.7 and only 7 points behind the closed leaders; and reproduction kit — MIT weights plus 7,000+ RL environments, the full training framework, mini-harnesses, a technical report, and a livestreamed $2.62M RL run that anyone can now attempt to rebuild. The honestly-labiled framing: Xiaomi's own flagship benchmark bullets (DeepSWE, agent tables, the "surpasses Kimi K3 and Qwen3.8 Max" line, the 1/20–1/60 price ratio) are COMPANY CLAIMS awaiting independent replication, and 46 is the open-weights top, not the overall top (Fable 5.1 / GPT-6 Astra lead at 53). The through-line for the weekly: the frontier-for-less pricing war now has an open-weights supply side at $0.13 per index task with a 99% cache discount, and "open-sourcing the training machinery" may matter more long-term than any single benchmark row.

Illustration: frame: a translucent iceberg-vault over a calm sea, packed with tiny glowing modules, its gate standing wide open — an artistic impression of a trillion-parameter open-weight model released permiss…
⌘

Lab: VERIFY

Step 1 — Verify ungated status and event date via the Hugging Face API (read-only)
for repo in "XiaomiMiMo/MiMo-V2.6-Pro-RL" "XiaomiMiMo/MiMo-V2.6-Flash-RL" "XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B"; do
  curl -s "https://huggingface.co/api/models/$repo" | python3 -c "
import json,sys
d = json.load(sys.stdin)
print(d['id'], '| gated:', d.get('gated'), '| created:', d.get('createdAt'),
      '| downloads:', d.get('downloads'), '| license-field:', d.get('license'))"
done

Actual output (2026-09-23):

XiaomiMiMo/MiMo-V2.6-Pro-RL | gated: False | created: 2026-09-21T15:39:33.000Z | downloads: 4070 | license-field: None
XiaomiMiMo/MiMo-V2.6-Flash-RL | gated: False | created: 2026-09-21T15:39:51.000Z | downloads: 13243 | license-field: None
XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B | gated: False | created: 2026-09-21T18:18:40.000Z | downloads: 3253 | license-field: None

Findings:

  • UNGATED — VERIFIED (gated: False on all three).
  • Event date fixed at 2026-09-21 via repository creation timestamps (15:39 UTC) — the weights went live Sep 21 even though the announcement text is dated Sep 22.
  • Metadata quirk caught: the API license field is None on all three, even though the rendered cards and README YAML declare MIT. License-audit pipelines that rely on the API field alone would falsely flag these repos as unlicensed. (Discovery of this quirk is the point of the exercise.)
Step 2 — Verify the license from the authoritative location (README YAML)
curl -s "https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/raw/main/README.md" | head -12
curl -s "https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL/raw/main/README.md" | head -12

Actual output (first lines, both repos identical in the metadata block):

---
license: mit
language:
- en
- zh
tags:
- text-generation
- multimodal
- vision-language
- audio
- agent
- video-understanding
- long-context
- mimo_v2
- transformers
library_name: transformers
---

Finding: MIT license — VERIFIED from card metadata (cross-confirmed by Artificial Analysis' model page and the rendered model-card pages).

Step 3 — Rebuild the open-weights leaderboard from Artificial Analysis' open-source page (fetched 2026-09-23)
RankModelLabLicense classAA Intelligence Index (v4.3.2)
1MiMo-V2.6-ProXiaomiMIT open weights46
2GLM-5.3 (max)Z AIopen weights45
3Kimi K3 (max)Kimiopen weights44
4GLM 5.3 FlashZ AIopen weights42
5DeepSeek V4.1 Flash (Reasoning, Max Effort)DeepSeekopen weights39
6Qwen3.8 27B (xhigh)Alibabaopen weights34
7K2 Horizon 375B A23BInstitute of Foundation Modelsopen weights31
8MiniMax-M3MiniMaxopen weights29
9Inkling (xhigh)Thinking Machinesopen weights25
10Nemotron 3 Ultra 550B A55BNVIDIAopen weights23
11Muse Glimmer (high)Metaopen weights17

Page prose verbatim: "MiMo-V2.6-Pro and GLM-5.3 (max) are the highest intelligence open source models". Closed-model context from the same ecosystem: Grok 4.7 (xhigh) 46 (proprietary), Claude Fable 5.1 / GPT-6 Astra 53 (per Xiaomi docs + S03 research).

Finding: "Tops open-weight rankings at 46" — INDEPENDENTLY VERIFIED (46 > 45 > 44; #1 of 114 in the large open-weights class on the model page). Note the tie with proprietary Grok 4.7 at 46 and the +7 gap to the closed leaders — that is the accurate framing, not "overall #1."

Step 4 — Reproduce the per-intelligence-point economics (pure arithmetic on published figures)
# per_index_point.py — run with plain python3, no deps
rows = [
    # model, origin, index, cost_per_index_task_usd, token_price_in_out
    ("MiMo-V2.6-Pro (open, MIT)", "Xiaomi",   46, 0.13,  "0.435 / 0.87"),
    ("Grok 4.7 xhigh (closed)",   "SpaceXAI", 46, 3.74,  "2.00 / 6.00"),
    ("GPT-6 Astra Max (closed)",  "OpenAI",   53, 3.26,  "n/a"),
    ("GLM-5.3 max (open)",        "Z AI",     45, None,  "n/a"),
]
print(f"{'model':<28}{'idx':>4}{'$/task':>8}{'$/idx-pt':>9}")
for name, lab, idx, task, _ in rows:
    per_pt = round(task / idx, 4) if task and idx else float("nan")
    print(f"{name:<28}{idx:>4}{str(task) if task else '-':>8}{per_pt:>9.4f}")

Expected output on the published numbers:

model                         idx   $/task $/idx-pt
MiMo-V2.6-Pro (open, MIT)      46     0.13    0.0028
Grok 4.7 xhigh (closed)        46     3.74    0.0813
GPT-6 Astra Max (closed)       53     3.26    0.0615
GLM-5.3 max (open)             45        -        -

Reading of the table: MiMo-V2.6-Pro costs ~$0.0028 per index point vs ~$0.0813 for Grok 4.7 — a ~29× per-intelligence-point gap (and ~22× vs GPT-6 Astra Max per index point), driven by the 99% cache discount and low token price. This is the numeric core of the story: on AA's own cost-per-index-task methodology, the top open-weights model is ~1–2 orders of magnitude cheaper per intelligence point than the closed entrants at the same or higher index scores.

Step 5 — Conclude with a verification checklist (reusable for any open-weights release)
ClaimVerdictEvidence
Repos ungatedVERIFIEDHF API gated: False (3 repos)
MIT licenseVERIFIEDREADME YAML license: mit + cards + AA
Weights live Sep 21, 2026VERIFIEDHF createdAt 2026-09-21T15:39Z + AA release date
#1 open-weights score 46VERIFIEDAA model page (#1/114) + AA open-source leaderboard (46>45>44)
Parameter counts 1.02T/42B & 309B/15BVERIFIEDHF model cards + AA; page label "311B" noted as rounding
API pricing $0.435/$0.87, 99% cache, $0.13/taskVERIFIEDAA model page + Unite.AI recap
"Surpasses Kimi K3 and Qwen3.8 Max"PARTIALKimi K3 (max) 44 < 46 independently; "Qwen3.8 Max" naming only in Xiaomi docs
DeepSWE/agent/cyber benchmark rowsNOT YET VERIFIEDCompany-published only (card tables); independent re-runs pending
RL economics ($2.62M, 750K trajectories)NOT YET VERIFIEDCompany-published; framework released so falsifiable
Speed figures (54.2 vs 129.7 t/s)TIME-SENSITIVEAA page updated; capture measurement date when citing

What this lab does NOT do (honest limits)

  • It does not benchmark the model — no API spend, no local inference, no task runs.
  • It treats all vendor benchmark tables as COMPANY CLAIM throughout (evidence discipline).
  • AA index-task cost is methodology-specific (AA's 10-eval composite); real production workloads will differ (prompt mix, caching behavior, effort levels).
  • Leaderboard rows are live data as of 2026-09-23 and will drift as more models are added/re-evaluated.

Result

A one-page, evidence-labelled verification pack proving the story's core claims (ungated MIT, event date, #1 open-weights score, per-point economics) from primary + independent sources, surfacing four caveats worth telling the audience (API license-field quirk, Qwen3.8 Max naming, unverified agent-table numbers, time-sensitive speed figures).

≡

Research sources

Primary Sources (7)
Primary
Hugging Face — REST API model records (queried read-only via curl) - **URL (Pro):** https://huggingface.co/api/models/XiaomiMiMo/MiMo-V2.6-Pro-RL - **URL (Flash):** https://huggingface.co/api/models/XiaomiMiMo/MiMo-V2.6-Flash-RL - **URL (Distill):** https://huggingface.co/api/models/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B** `gated: false` for all three repos (ungated — CONFIRMED); `createdAt` timestamps 2026-09-21T15:39:33Z (Pro), 15:39:51Z (Flash), 18:18:40Z (Distill) — independently fixing the event date at Sep 21 (weights live); pipeline tags (text-generation ×2, image-text-to-text); downloads/likes; file counts (155 / 90 / 17 siblings); and the metadata quirk that the top-level `license` field is `None` even though the cards and README YAML declare MIT — flagged for tooling that audits licenses via the API. — ** FACT (ungated status, event date via repository creation timestamps, metadata inconsistency). ---Date: ** queried 2026-09-23 (records created 2026-09-21)
URL unavailable
Primary
Hugging Face — README YAML front-matter (raw), MiMo-V2.6-Pro-RL and MiMo-V2.6-Flash-RL - **URL (Pro):** https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/raw/main/README.md - **URL (Flash):** https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL/raw/main/README.md** Declared `license: mit` in the model metadata; language en/zh; modality/agent/audio/video-understanding/long-context tags; `library_name: transformers`. This is the authoritative license location (the HF REST API `license` payload field returns `None` — see source 7). — ** FACT (MIT license — primary metadata).Date: ** 2026-09-21 (queried via curl 2026-09-23)
URL unavailable
Primary
Hugging Face — Model card: XiaomiMiMo/MiMo-V2.6-Flash-RL (fetched in full) - **URL:** https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL** License: mit; architecture (sparse MoE 309B total / 15B active; 48 layers, 39 SWA + 9 GA, hidden 4096, 256 routed experts / 8 active, 1M context); same vision/audio encoders and 5-layer MTP drafter; company evaluation table (DeepSWE v1.1 67.9, Terminal Bench 2.1 87.6, Toolathlon-Verified 73.6, CyberGym 95.1, MiMo VisualCoding 71.5, etc.); deployment recipes (SGLang tp8, vLLM tp4); 13.2k downloads last month at research time; 28 community quantizations listed for Flash. — ** FACT (architecture, license, modalities) + COMPANY CLAIM (evaluation table).Date: ** 2026-09-21 (created 2026-09-21T15:39:51Z)
URL unavailable
Primary
Hugging Face — Model card: XiaomiMiMo/MiMo-V2.6-Pro-RL (fetched in full) - **URL:** https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL** License: mit; architecture (sparse MoE 1.02T total / 42B active; 70 layers, 60 SWA + 10 GA, hidden 6144, 384 routed experts / 8 active, SWA window 128, 1M context); modalities (text/image/video/audio input, text output); vision encoder (681M MiMo ViT, 28 layers); audio encoders (308M AudioTokenizer + 127M patch encoder); 5-layer MTP speculative decoder (7 tokens per forward pass); RL method (async GRPO 1,568 prompts × 16 rollouts/step; GRS; GAR; MOPD2; "You Only RL Once"; aligned RL/self-correction cold start); company evaluation table (DeepSWE v1.1 71.9, ProgramBench 26.5, AutomationBench 53.1, Toolathlon-Verified 76.9, GDPval-AA 2.1 1673, Agents' Last Exam 31.6, Terminal Bench 4.0 34.9 / 2.1 89.9, OSWorld-Verified 82.0, JobBench 62.0, CyberGym 94.0, MiMo Cyber Bench 80.2, ExploitGym 17.8, ExploitBench 47.9, SEC Bench Pro 66.3, MiMo VisualCoding 72.3 vs comparators); deployment recipes (SGLang tp16/dp2/2-node, vLLM tp8, reasoning/tool parsers, temperature 1.0/top_p 0.95); availability (AI Studio, MiMo Code, Desktop, Open Platform API, OpenRouter, ModelScope); technical report PDF reference; sample code (transformers pipeline, vLLM/SGLang serve commands). — ** FACT (architecture, license, modalities, deployment) + COMPANY CLAIM (evaluation table).Date: ** 2026-09-21 (created 2026-09-21T15:39:33Z)
URL unavailable
Primary
Hugging Face — "MiMo-V2.6" collection (official XiaomiMiMo org) - **URL:** https://huggingface.co/collections/XiaomiMiMo/mimo-v26** Official release-crate structure: three repositories (MiMo-V2.6-Pro-RL [Text Generation, 1T], MiMo-V2.6-Flash-RL [Text Generation, 311B page label], MiMo-V2.6-Distill-Qwen-9B [Image-Text-to-Text, 9B]); community uptake proxies (downloads/likes: Pro 4.07k/437, Flash 13.2k/416, Distill 3.25k/392 at research time); org ownership (XiaomiMiMo). — ** FACT (release artifacts exist and are live).Date: ** 2026-09-21 (collection updated; entries created Sep 21, updated Sep 22)
URL unavailable
Primary
Xiaomi MiMo — Official blog announcement page "Introducing MiMo-V2.6 series" - **URL:** https://mimo.xiaomi.com/mimo-v2-6** Official announcement URL referenced from both model cards and independent coverage; page title "MiMo-V2.6 | Xiaomi" and meta description "Introducing the MiMo-V2.6 series: frontier intelligence, all the modalities, built in public." (page is a JavaScript app rendering an iframe article served from /mimo-v2-6/article.html; full text mirrored by source 1). — ** Primary announcement artifact (canonical URL; content corroborated by the docs page).Date: ** 2026-09-22 (announcement article; page verified during research)
URL unavailable
Primary
Xiaomi MiMo — Official announcement and documentation: "MiMo-V2.6: Scaling Up Reinforcement Learning for Self-Improvement" (news/docs page) - **URL:** https://mimo.mi.com/docs/en-US/news/latest/v2-6** Official release narrative; series composition (Pro, Flash, Pro-UltraSpeed, Distill-Qwen-9B); AA Intelligence Index 46 claim and "surpasses Kimi K3 and Qwen3.8 Max / most powerful open-source model" wording; admission of gap to Claude Fable 5.1 and GPT-6 Astra; RL-run transparency figures (6 days, 30 steps each, ~750,000 trajectories, ~$0.85M Flash / ~$2.62M Pro training cost, pass-rate improvements 25%/12%, DeepSWE v1.1 48.8→65.7 / 58.4→72.6); RL scaling mechanics (1,568 samples/update, 1M context, 3.5–3.7B tokens/step, frozen MoE router, reward-hacking defense); open-source contents (7k+ RL task environments across software engineering / vulnerability reproduction / knowledge-intensive work / web design and development; end-to-end RL framework on verl/uni-agent/mini-swe-agent; composable mini-harnesses; technical report); Distill-9B benchmark deltas (SWE-bench Verified 61.1→66.2 etc.); unchanged V2.5 API pricing; UltraSpeed up to 20×; Desktop client launch; model names (mimo-v2.6-pro / flash / pro-ultraspeed); demo cases (3D games, Blender, Franka Panda arm, CUA, MOF/PFAS materials research, Lean 4 "Period Three Implies Chaos" formalization, video/music creation, Design Arena parity claim); "1/20 to 1/60 price of overseas models" claim; official open-source link to the Hugging Face collection. — ** FACT baseline (release facts, dates, pricing as published) + COMPANY CLAIM (all performance/economics claims are vendor-published).Date: ** 2026-09-22 (page footer: "Update Time September 22, 2026")
URL unavailable
Independent Sources (4)
Independent
CCLeaks — "Xiaomi MiMo-V2.6 opens Pro, Flash and 7,000+ RL tasks" (Abhishek Tiwari) - **URL:** https://ccleaks.com/news/xiaomi-mimo-v2-6-open-weights-sep-2026** Independent technical analysis: ungated MIT repositories (via HF API/README); Pro 1.02T/42B and Flash 309B/15B; 1M context; four native modalities; 7,000+ RL environments (four task families); end-to-end RL framework + mini-harnesses; Distill-Qwen-9B described as fine-tuned from Qwen3.5-9B on MiMo-generated data; ~750,000 joint trajectories in under six days (30 steps each); an explicit editorial decision to carry *no* AA scores/prices (time-sensitive live data) — a useful adversarial check on score-citing; practical guidance (pin revisions, verify license at the pinned revision, review custom-code payloads). — ** Independent technical reporting; corroborates release-crate facts and provides the "verify at source" discipline used in this research. ---Date: ** 2026-09-21
URL unavailable
Independent
Unite.AI — "Xiaomi's New Flagship Model Leads Open-Weight Rankings With a Score of 46" - **URL:** https://www.unite.ai/xiaomis-new-flagship-model-leads-open-weight-rankings-with-a-score-of-46** Independent recap: AA score 46 highest open-weights result; open-source comparison places MiMo first ahead of GLM-5.3 (45) and Kimi K3 (44); Grok 4.7 (xhigh) also 46 (proprietary); Xiaomi announcement cites 46.32; $0.13 per index task and $206.66 total index evaluation cost; AA release date Sep 21 vs announcement date Sep 22; architecture details (70-layer backbone, 60 SWA + 10 GA, 384/8 experts, 681M ViT, 308M+127M audio, 5-layer MTP); RL-run details (~6 days, 30 steps each, ~750K trajectories, ~$0.85M/$2.62M, DeepSWE 48.8→65.68 / 58.4→72.57); pricing table incl. cache-hit tiers (Pro $0.0036/$0.435/$0.87; Flash $0.0028/$0.14/$0.28; UltraSpeed $0.036/$4.35/$8.70 per M tokens); early AA speed readout 129.7 t/s / 2.17 s TTFT (cf. current 54.2/2.67 on source 8 — flagged in research Section 8); availability (AI Studio, MiMo Code, Desktop, API, OpenRouter); demos (game dev, Blender, Franka Panda arm, video, orchestral piece "Night Road"); research cases (MOF/PFAS, Lean 4 "Period Three Implies Chaos" >6,000 lines kernel-verified); DeepSWE comparison rows incl. DeepSeek V4.1 Flash 74.2 / Claude Opus 5 74.0 / GPT 6 Astra 74.0. — ** Independent reporting; corroborates official claims and supplies the earliest third-party readout of AA data.Date: ** 2026-09-21 (published; by Jonas Reeve, AI-generated research analyst with disclosed editorial review)
URL unavailable
Independent
Artificial Analysis — "Comparison of Open Source AI Models" (open-source leaderboard page, fetched in full) - **URL:** https://artificialanalysis.ai/models/open-source** INDEPENDENTLY VERIFIED open-weights ordering on the Intelligence Index: **MiMo-V2.6-Pro 46** (top) > GLM-5.3 (max) 45 > Kimi K3 (max) 44 > GLM 5.3 Flash 42 > DeepSeek V4.1 Flash 39 > Qwen3.8 27B (xhigh) 34 > K2 Horizon 31 > MiniMax-M3 29 > Inkling 25 > Nemotron 3 Ultra 23 > Muse Glimmer 17; page prose "MiMo-V2.6-Pro and GLM-5.3 (max) are the highest intelligence open source models"; MiMo-V2.6-Pro row (1.0T / 42B active / 1M ctx / 54 t/s / weights link to HF); large-model class definition (>150B) that makes the "#1/114" rank meaningful. — ** Independent evaluation platform — verifies the "tops open-weight rankings at 46" headline claim.Date: ** consulted 2026-09-23 (leaderboard rows are live data)
URL unavailable
Independent
Artificial Analysis — "MiMo-V2.6-Pro — Intelligence, Performance & Price Analysis" (model page, fetched in full) - **URL:** https://artificialanalysis.ai/models/mimo-v2-6-pro** INDEPENDENTLY VERIFIED Intelligence Index 46 on v4.3.2, ranked **#1 of 114** in its class (large open-weights reasoning models); released September 21, 2026; open weights; MIT license; 1.0T total / 42B active; 1M context (~1,500 A4 pages); text/image/speech/video input, text output; reasoning model; output speed 54.2 t/s (#42/114, below class median 67.1); TTFT 2.67 s (vs 2.32 s median); pricing $0.435/M in / $0.87/M out, 99% cache discount, $0.13 cost per Intelligence Index task; 140M output tokens on the index (median 140M); single API provider (Xiaomi); precise FAQ answers (license, release date, params, modalities, provider count). — ** Independent technical evaluation — the key third-party confirmation of the headline score, license, size, and economics.Date: ** release date listed 2026-09-21; page consulted 2026-09-23 (figures are live/updated)
URL unavailable