News Weekly
LV 10 XP
0% read
S18research
#18 Issue #1Confirmed

MLPerf Inference v6.1 released with 30 submitters, Vera Rubin NVL72 debut, new E2E-RAG and Edge-Agentic workloads

MLCommons released MLPerf Inference v6.1 — the industry-standard inference benchmark suite — on September 16, 2026, a record round that added two brand-new evaluation workloads. Essential facts (all CONFIRMED against the official announcement unless labelled otherwise): Record participation: a record-high 30 participating organizations submitted: AMD, ASUSTeK, Atlas Inference, Cisco, CoreWeave, Crusoe, Dell, Fujitsu, GigaComputing, Google, HPE, Intel, Inventec, KRAI, Lambda, MangoBoost, Microsoft Azure, MiTAC, Nebius, NVIDIA, Oracle, Orrick, Quanta Cloud Technology, Red Hat, ScitiX, Supermicro, Telecommunications Technology Association, VibeHPC, Wiwynn, and individual contributor Naeem Khoshnevis (Harvard's Kempner Institute, single-H200 Llama 3.1-8B entry). The prior round (v6.0, April 2026) drew 24 organizations (TechTimes). Six were first-time submitters: Atlas Inference, Crusoe, Orrick Industries LLC, ScitiX, VibeHPC, and Naeem Khoshnevis. Scale: 120 systems across the Datacenter and Edge suites in both Closed and Open divisions (chairs analysis), generating 486 individual datacenter and edge results (StorageReview). Several were joint submissions (Dell_AMD, Dell_MangoBoost, RedHat_Intel, RedHat_Supermicro). Two new benchmarks (FACT) — the structural headline: an End-to-End Retrieval-Augmented Generation (E2E-RAG) test for the datacenter and an Edge Agentic Inference test for the edge. These are the first standardized MLPerf measurements of multi-component RAG pipelines and of multi-turn agentic tool-calling inference. Vera Rubin NVL72 debut (FACT): NVIDIA Rubin and NVIDIA Vera Rubin NVL72 appeared in MLPerf's preview availability category — the first peer-reviewed, consortium-verified performance numbers for the Rubin architecture. Nebius independently filed a second Vera Rubin preview entry (its VR200 NVL72 on 36 GPUs across nine nodes), so the round carries two independent Rubin submissions. Three other new platforms appeared in the available category: AMD Ryzen AI Max+ 395, AMD Instinct MI350P, and Intel Arc Pro B70 (first peer-reviewed results for all five). Headline performance gains (FACT as reported by MLCommons): best per-accelerator DeepSeek-R1 Server result 5.7× better than v5.1 one year earlier; best per-accelerator VLM (Qwen3-VL) Server result 2.99× better than v6.0 six months earlier. The chairs analysis attributes both step-changes primarily to the Vera Rubin preview submissions on those two models. Largest system ever submitted: a 512-accelerator submission (Crusoe entries 6.1-0026/6.1-0027, AMD MI355X), which also set a token-throughput record of almost 5.8M tokens/second on gpt-oss-120b Offline (chairs analysis; 5,749,440 tok/s per AMD's technical blog). The round also set a record of 16 multi-node submissions. Two novel heterogeneous systems (FACT): the first cross-vendor heterogeneous accelerator deployment — Cisco unifying eight NVIDIA H200 and eight AMD Instinct MI350X into one inference pool over a Cisco G200 network — and the first geographically distributed system spanning the Pacific — MangoBoost serving four sites on two continents as one endpoint at 97% scaling efficiency. Benchmark suite updates (FACT): v6.1 comprises 10 Datacenter and 6 Edge benchmarks; a new Interactive scenario for VLM (Qwen3-VL-based); speculative decoding support added to the GPT-OSS-120B interactive scenario; and — a market signal — gpt-oss-120b (MoE) became the most-submitted datacenter benchmark, displacing Llama2-70B for the first time (chairs analysis). The Endpoints inflection (FACT): over 50% of submitters (16 of 30) used MLPerf's new API-centric harness, the foundation of the next-generation MLPerf Endpoints suite (vs a single open-division submitter in v6.0). David Kanter, Head of MLPerf: "Moving forward, MLPerf Endpoints will replace Inference in our family of benchmarks for the datacenter." StorageReview adds that Endpoints opens on-demand rolling submissions in October 2026 and replaces Inference as the datacenter benchmark in 2027, with normalized results and expanded agentic workloads planned for Endpoints v1.0.

Dozens of identical measurement rigs fill a vast hall beneath one broad canopy that gathers their light into a single scored marker.
How do you want to read this?

Tailored emphasis while keeping the full article available.

Best for you · Explorer

🎓 Start with the story, why it matters, and where it goes next.

At a glance

The essential information in 30 seconds

What happened

MLCommons released MLPerf Inference v6.1 — the industry-standard inference benchmark suite — on September 16, 2026, a record round that added two brand-new evaluation workloads. Essential facts (all CONFIRMED against the official announcement unless labelled otherwise):

  • Record participation: a record-high 30 participating organizations submitted: AMD, ASUSTeK, Atlas Inference, Cisco, CoreWeave, Crusoe, Dell, Fujitsu, GigaComputing, Google, HPE, Intel, Inventec, KRAI, Lambda, MangoBoost, Microsoft Azure, MiTAC, Nebius, NVIDIA, Oracle, Orrick, Quanta Cloud Technology, Red Hat, ScitiX, Supermicro, Telecommunications Technology Association, VibeHPC, Wiwynn, and individual contributor Naeem Khoshnevis (Harvard's Kempner Institute, single-H200 Llama 3.1-8B entry). The prior round (v6.0, April 2026) drew 24 organizations (TechTimes). Six were first-time submitters: Atlas Inference, Crusoe, Orrick Industries LLC, ScitiX, VibeHPC, and Naeem Khoshnevis.
  • Scale: 120 systems across the Datacenter and Edge suites in both Closed and Open divisions (chairs analysis), generating 486 individual datacenter and edge results (StorageReview). Several were joint submissions (Dell_AMD, Dell_MangoBoost, RedHat_Intel, RedHat_Supermicro).
  • Two new benchmarks (FACT) — the structural headline: an End-to-End Retrieval-Augmented Generation (E2E-RAG) test for the datacenter and an Edge Agentic Inference test for the edge. These are the first standardized MLPerf measurements of multi-component RAG pipelines and of multi-turn agentic tool-calling inference.
  • Vera Rubin NVL72 debut (FACT): NVIDIA Rubin and NVIDIA Vera Rubin NVL72 appeared in MLPerf's preview availability category — the first peer-reviewed, consortium-verified performance numbers for the Rubin architecture. Nebius independently filed a second Vera Rubin preview entry (its VR200 NVL72 on 36 GPUs across nine nodes), so the round carries two independent Rubin submissions. Three other new platforms appeared in the available category: AMD Ryzen AI Max+ 395, AMD Instinct MI350P, and Intel Arc Pro B70 (first peer-reviewed results for all five).
  • Headline performance gains (FACT as reported by MLCommons): best per-accelerator DeepSeek-R1 Server result 5.7× better than v5.1 one year earlier; best per-accelerator VLM (Qwen3-VL) Server result 2.99× better than v6.0 six months earlier. The chairs analysis attributes both step-changes primarily to the Vera Rubin preview submissions on those two models.
  • Largest system ever submitted: a 512-accelerator submission (Crusoe entries 6.1-0026/6.1-0027, AMD MI355X), which also set a token-throughput record of almost 5.8M tokens/second on gpt-oss-120b Offline (chairs analysis; 5,749,440 tok/s per AMD's technical blog). The round also set a record of 16 multi-node submissions.
  • Two novel heterogeneous systems (FACT): the first cross-vendor heterogeneous accelerator deployment — Cisco unifying eight NVIDIA H200 and eight AMD Instinct MI350X into one inference pool over a Cisco G200 network — and the first geographically distributed system spanning the Pacific — MangoBoost serving four sites on two continents as one endpoint at 97% scaling efficiency.
  • Benchmark suite updates (FACT): v6.1 comprises 10 Datacenter and 6 Edge benchmarks; a new Interactive scenario for VLM (Qwen3-VL-based); speculative decoding support added to the GPT-OSS-120B interactive scenario; and — a market signal — gpt-oss-120b (MoE) became the most-submitted datacenter benchmark, displacing Llama2-70B for the first time (chairs analysis).
  • The Endpoints inflection (FACT): over 50% of submitters (16 of 30) used MLPerf's new API-centric harness, the foundation of the next-generation MLPerf Endpoints suite (vs a single open-division submitter in v6.0). David Kanter, Head of MLPerf: "Moving forward, MLPerf Endpoints will replace Inference in our family of benchmarks for the datacenter." StorageReview adds that Endpoints opens on-demand rolling submissions in October 2026 and replaces Inference as the datacenter benchmark in 2027, with normalized results and expanded agentic workloads planned for Endpoints v1.0.

Substance requiring careful label-splitting (COMPANY CLAIM unless otherwise verified):

  • NVIDIA (submitter blog, Sep 16): Vera Rubin NVL72 delivered "up to 3.7× higher throughput than GB300 NVL72" on Qwen3-VL (using vLLM with NVIDIA Dynamo) and "up to 2.5×" on DeepSeek-R1 (using TensorRT-LLM), entries 6.1-0106 and 6.1-0074; a four-rack (288-GPU) GB300 NVL72 DeepSeek-R1 submission hit 99% scaling efficiency (Offline); GB300 Qwen3-VL improved up to 1.6× over v6.0 via software (lower KV-cache precision, kernel fusion, disaggregated serving); post-deadline gains on GPT-OSS-120B/DLRMv3 are explicitly unverified by MLCommons. Independent reading of the same public entries (diyai.io) shows the Interactive scenario driving the 3.7× (≈3.74×) while Offline/Server gains are ≈1.7–2.0× — the headline multiplier is scenario- and stack-specific.
  • AMD and partners (supplemental + AMD ROCm blog + WCCFTech): Crusoe's 512-GPU MI355X cluster led DeepSeek-R1 and GPT-OSS-120B at extreme scale (5.75M tok/s GPT-OSS Offline; 2.90M DeepSeek-R1); AMD states MI355X led B200/B300 GPT-OSS-120B at 8 GPUs and GB200 at 72 GPUs, with 95% scale efficiency at 72 GPUs; MI350P led selected RTX PRO 6000/H200 NVL results in its first round.
  • Intel (supplemental): Xeon 6 expanded from two SKUs/24 results to five SKUs/35 results — the only standalone server CPU represented; a Xeon 6787P + 4× Arc Pro B70 node (128 GB VRAM) made the round's only E2E-RAG submissions, with CPU and GPU both on the critical path in a single measured run; gpt-oss-120b Server improved 36% / Offline 27% on identical four-GPU Arc Pro B70 hardware since v6.0.
  • Lambda (submitter, Open division): first agentic inference workload on datacenter hardware in MLPerf — swapped the reference Qwen3.6-27B for Kimi K2.6 (Moonshot AI, >1 trillion parameters, MoE), served FP8 on an HGX B200 with SGLang at TP=8; 1,007/1,007 replay turns completed with zero failures; mean per-turn latency 770.8 ms (median 361.0 ms), mean TTFT 179.7 ms; BFCL v4 accuracy 86.83% (inside the field's 86.13–87.94% band) — also the first >1T-parameter model ever deployed in an MLPerf submission.
  • Atlas Inference (supplemental): edge agentic entries on NVIDIA DGX Spark (GB10) and AMD Strix Halo at ~20 tok/s (1,007 turns in under 64 minutes on DGX Spark).
Why it matters

MLPerf Inference v6.1 matters because the industry's measurement standard finally caught up with what production AI actually is: multi-model RAG pipelines, multi-turn agents, and rolling API-centric evaluation. A record 30 submitters, 120 systems and 486 results delivered the first neutral, peer-reviewed numbers for Vera Rubin NVL72 (preview), a 512-accelerator system at ~5.8M tokens/second, cross-vendor and trans-Pacific heterogeneous deployments, and documented software-only gains of 8-36% per cycle. Two new workload categories (E2E-RAG and Edge Agentic Inference) become official benchmarks for the first time, and MLCommons committed to MLPerf Endpoints replacing Inference as the datacenter benchmark. The durable consequence is that procurement-grade, apples-to-apples AI performance measurement has arrived for pipelines and agents, so benchmark-led buying becomes the rational default.

Evidence

CONFIRMED

0 sources · 72 min read
Story identity
  • Story ID: S18
  • Title: MLPerf Inference v6.1 released with 30 submitters, Vera Rubin NVL72 debut, new E2E-RAG and Edge-Agentic workloads
  • Organization: MLCommons (open engineering consortium, 130+ members and affiliates; home of the MLPerf benchmark family)
  • Category: research (industry-standard AI benchmarking)
  • Event date: 2026-09-16 — CONFIRMED, in-window (2026-09-10 ≤ 2026-09-16 ≤ 2026-09-17 ✓). MLCommons published the MLPerf Inference v6.1 results and announcement on September 16, 2026 (official announcement dated Sep 16, 2026; GlobeNewswire press release timestamped 2026-09-16T15:00:00Z; corroborated by StorageReview, WCCFTech and Unite.AI articles all dated Sep 16, 2026).
  • Announcement date: 2026-09-16 (official announcement + press release). A follow-up chairs analysis, "Where the Industry Is Investing: A Look at MLPerf Inference v6.1," was published 2026-09-17 — also inside the window.
  • Article dates: 2026-09-16 (MLCommons, GlobeNewswire, StorageReview, WCCFTech, Unite.AI, diyai.io, theclarity.today, Lambda, NVIDIA blog) and 2026-09-17 (MLCommons chairs analysis, TechTimes, IoT Tech News, agenccy.ai). Discovery record's article_dates (2026-09-16) is correct for the primary coverage.
  • Evidence status: CONFIRMED (FACT) at the level of the release itself. The v6.1 results, the record 30-submitter field, the two new workload categories (E2E-RAG, Edge Agentic), the preview classification of NVIDIA Rubin / Vera Rubin NVL72, the headline improvement figures (best per-accelerator DeepSeek-R1 Server result 5.7× better than v5.1 a year ago; best per-accelerator VLM Server result 2.99× better than v6.0 six months earlier), the 512-accelerator system, the two heterogeneous systems, and the commitment that MLPerf Endpoints will replace Inference for the datacenter are all stated in MLCommons' own official announcement and chairs analysis, and independently corroborated by multiple outlets. COMPANY CLAIM at the level of vendor self-characterizations. NVIDIA's "up to 3.7× / up to 2.5×" Vera-Rubin-vs-GB300 comparisons, AMD's 512-GPU leadership framing, and the "lower cost per token" economics in NVIDIA's post are submitter interpretations built on top of MLCommons-verified raw entries. Per AGENTS.md source preference and evidence discipline, the NVIDIA blog is treated as submitter material: the raw numbers pass MLPerf peer review (FACT), the vendor's comparative/economics spin is COMPANY CLAIM.
  • Discovery-record note: discovery named "The Register" and "VentureBeat" as independent sources; searches in this pass found no v6.1-specific articles from either outlet (only older MLPerf coverage, 2019–2024). No content from either was used; independent coverage actually used comes from StorageReview, WCCFTech, TechTimes, Unite.AI, IoT Tech News, diyai.io, theclarity.today and agenccy.ai (see sources/S18.md). The discovery record itself was otherwise accurate in every checkable detail.
✓

What happened?

🎓 For Explorer

MLCommons released MLPerf Inference v6.1 — the industry-standard inference benchmark suite — on September 16, 2026, a record round that added two brand-new evaluation workloads. Essential facts (all CONFIRMED against the official announcement unless labelled otherwise):

  • Record participation: a record-high 30 participating organizations submitted: AMD, ASUSTeK, Atlas Inference, Cisco, CoreWeave, Crusoe, Dell, Fujitsu, GigaComputing, Google, HPE, Intel, Inventec, KRAI, Lambda, MangoBoost, Microsoft Azure, MiTAC, Nebius, NVIDIA, Oracle, Orrick, Quanta Cloud Technology, Red Hat, ScitiX, Supermicro, Telecommunications Technology Association, VibeHPC, Wiwynn, and individual contributor Naeem Khoshnevis (Harvard's Kempner Institute, single-H200 Llama 3.1-8B entry). The prior round (v6.0, April 2026) drew 24 organizations (TechTimes). Six were first-time submitters: Atlas Inference, Crusoe, Orrick Industries LLC, ScitiX, VibeHPC, and Naeem Khoshnevis.
  • Scale: 120 systems across the Datacenter and Edge suites in both Closed and Open divisions (chairs analysis), generating 486 individual datacenter and edge results (StorageReview). Several were joint submissions (Dell_AMD, Dell_MangoBoost, RedHat_Intel, RedHat_Supermicro).
  • Two new benchmarks (FACT) — the structural headline: an End-to-End Retrieval-Augmented Generation (E2E-RAG) test for the datacenter and an Edge Agentic Inference test for the edge. These are the first standardized MLPerf measurements of multi-component RAG pipelines and of multi-turn agentic tool-calling inference.
  • Vera Rubin NVL72 debut (FACT): NVIDIA Rubin and NVIDIA Vera Rubin NVL72 appeared in MLPerf's preview availability category — the first peer-reviewed, consortium-verified performance numbers for the Rubin architecture. Nebius independently filed a second Vera Rubin preview entry (its VR200 NVL72 on 36 GPUs across nine nodes), so the round carries two independent Rubin submissions. Three other new platforms appeared in the available category: AMD Ryzen AI Max+ 395, AMD Instinct MI350P, and Intel Arc Pro B70 (first peer-reviewed results for all five).
  • Headline performance gains (FACT as reported by MLCommons): best per-accelerator DeepSeek-R1 Server result 5.7× better than v5.1 one year earlier; best per-accelerator VLM (Qwen3-VL) Server result 2.99× better than v6.0 six months earlier. The chairs analysis attributes both step-changes primarily to the Vera Rubin preview submissions on those two models.
  • Largest system ever submitted: a 512-accelerator submission (Crusoe entries 6.1-0026/6.1-0027, AMD MI355X), which also set a token-throughput record of almost 5.8M tokens/second on gpt-oss-120b Offline (chairs analysis; 5,749,440 tok/s per AMD's technical blog). The round also set a record of 16 multi-node submissions.
  • Two novel heterogeneous systems (FACT): the first cross-vendor heterogeneous accelerator deployment — Cisco unifying eight NVIDIA H200 and eight AMD Instinct MI350X into one inference pool over a Cisco G200 network — and the first geographically distributed system spanning the Pacific — MangoBoost serving four sites on two continents as one endpoint at 97% scaling efficiency.
  • Benchmark suite updates (FACT): v6.1 comprises 10 Datacenter and 6 Edge benchmarks; a new Interactive scenario for VLM (Qwen3-VL-based); speculative decoding support added to the GPT-OSS-120B interactive scenario; and — a market signal — gpt-oss-120b (MoE) became the most-submitted datacenter benchmark, displacing Llama2-70B for the first time (chairs analysis).
  • The Endpoints inflection (FACT): over 50% of submitters (16 of 30) used MLPerf's new API-centric harness, the foundation of the next-generation MLPerf Endpoints suite (vs a single open-division submitter in v6.0). David Kanter, Head of MLPerf: "Moving forward, MLPerf Endpoints will replace Inference in our family of benchmarks for the datacenter." StorageReview adds that Endpoints opens on-demand rolling submissions in October 2026 and replaces Inference as the datacenter benchmark in 2027, with normalized results and expanded agentic workloads planned for Endpoints v1.0.

Substance requiring careful label-splitting (COMPANY CLAIM unless otherwise verified):

  • NVIDIA (submitter blog, Sep 16): Vera Rubin NVL72 delivered "up to 3.7× higher throughput than GB300 NVL72" on Qwen3-VL (using vLLM with NVIDIA Dynamo) and "up to 2.5×" on DeepSeek-R1 (using TensorRT-LLM), entries 6.1-0106 and 6.1-0074; a four-rack (288-GPU) GB300 NVL72 DeepSeek-R1 submission hit 99% scaling efficiency (Offline); GB300 Qwen3-VL improved up to 1.6× over v6.0 via software (lower KV-cache precision, kernel fusion, disaggregated serving); post-deadline gains on GPT-OSS-120B/DLRMv3 are explicitly unverified by MLCommons. Independent reading of the same public entries (diyai.io) shows the Interactive scenario driving the 3.7× (≈3.74×) while Offline/Server gains are ≈1.7–2.0× — the headline multiplier is scenario- and stack-specific.
  • AMD and partners (supplemental + AMD ROCm blog + WCCFTech): Crusoe's 512-GPU MI355X cluster led DeepSeek-R1 and GPT-OSS-120B at extreme scale (5.75M tok/s GPT-OSS Offline; 2.90M DeepSeek-R1); AMD states MI355X led B200/B300 GPT-OSS-120B at 8 GPUs and GB200 at 72 GPUs, with 95% scale efficiency at 72 GPUs; MI350P led selected RTX PRO 6000/H200 NVL results in its first round.
  • Intel (supplemental): Xeon 6 expanded from two SKUs/24 results to five SKUs/35 results — the only standalone server CPU represented; a Xeon 6787P + 4× Arc Pro B70 node (128 GB VRAM) made the round's only E2E-RAG submissions, with CPU and GPU both on the critical path in a single measured run; gpt-oss-120b Server improved 36% / Offline 27% on identical four-GPU Arc Pro B70 hardware since v6.0.
  • Lambda (submitter, Open division): first agentic inference workload on datacenter hardware in MLPerf — swapped the reference Qwen3.6-27B for Kimi K2.6 (Moonshot AI, >1 trillion parameters, MoE), served FP8 on an HGX B200 with SGLang at TP=8; 1,007/1,007 replay turns completed with zero failures; mean per-turn latency 770.8 ms (median 361.0 ms), mean TTFT 179.7 ms; BFCL v4 accuracy 86.83% (inside the field's 86.13–87.94% band) — also the first >1T-parameter model ever deployed in an MLPerf submission.
  • Atlas Inference (supplemental): edge agentic entries on NVIDIA DGX Spark (GB10) and AMD Strix Halo at ~20 tok/s (1,007 turns in under 64 minutes on DGX Spark).
Δ

What changed?

  • Before (through v6.0, April 2026): MLPerf Inference measured models one at a time, single-shot. No standardized benchmark existed for (a) a full multi-model RAG pipeline or (b) multi-turn agentic tool-calling inference, in the datacenter or at the edge. Participation: 24 organizations. The flagship hardware story was Blackwell (GB200/GB300); Vera Rubin had no consortium-verified numbers — NVIDIA's AgentX results released at its AI Infra Summit on Sep 15, 2026 were self-hosted benchmark runs on early silicon, a different evidentiary class. The API-centric harness had a single open-division user in v6.0. MLPerf Endpoints existed only as a v0.7 foundation release (July 28, 2026) with initial results from CoreWeave, Google, Intel, KRAI and NVIDIA.
  • Change (v6.1, Sep 16, 2026): the suite formally acknowledged that production AI is pipelines and agents, not isolated models: E2E-RAG (multi-hop, multi-model, up to 5 retrieval rounds, vector DB + reranker + two LLM tiers + judge) and Edge Agentic Inference (multi-turn coding trajectories, tool calling, single-stream latency) became official benchmarks. Participation jumped to a record 30 organizations / 120 systems / 486 results. First peer-reviewed numbers appeared for Vera Rubin NVL72 (preview, NVIDIA + Nebius), MI350P, Arc Pro B70 and Ryzen AI Max+ 395. Sixteen of 30 submitters adopted the API-centric harness, and MLCommons officially committed to Endpoints replacing Inference for the datacenter. Aggregate measurement mix shifted: gpt-oss-120b (MoE) overtook Llama2-70B as the most-submitted workload; 16 multi-node and two heterogeneous submissions became the scale-out evidence base.
  • After (expected): procurement-grade evidence now exists for RAG pipelines and agentic edge devices; RFPs and vendor scorecards will include E2E-RAG and Agentic numbers; the datacenter benchmark regime transitions to rolling, API-centric, normalized Endpoints results (October 2026 sign-up; 2027 replacement); Vera Rubin moves from preview to the available category (preview obligates resubmission) as it ships; software-only gains (8–36% on identical hardware within one cycle) become a measurable, citable baseline for "optimize before you buy."
↔

Before → Change → After

🎓 For Explorer
DimensionBefore (through Sep 15, 2026)Change (Sep 16, 2026 — v6.1 release)After (expected)
Benchmark scope10 DC + 6 edge workloads, all single-model/single-shot+ E2E-RAG (multi-model pipeline, up to 5 hops) and Edge Agentic (1,007-turn replay)"Real workload" defined as pipelines/agents; single-model numbers recontextualized
Participation24 orgs (v6.0)Record 30 orgs, 120 systems, 486 results; 6 first-timersNeoclouds/ODMs/software companies regular submitters; Endpoints lowers the bar further
Rubin evidenceNone peer-reviewed; only NVIDIA-hosted AgentX results (Sep 15)First MLPerf-verified Vera Rubin NVL72 preview numbers (NVIDIA + Nebius)Rubin reappears in available category as GA ships; buyers compare verified racks
RAG/agentic metricsNoneFRAMES-based E2E-RAG accuracy+throughput; BFCL-v4-gated agentic accuracy+latencyRFP checkboxes include pipeline/agentic columns
Measurement regimeFixed biannual rounds; harness mostly vendor-embedded16/30 use API-centric client/server harness; Endpoints succession committedRolling submissions (Oct 2026), normalized Pareto curves; Endpoints replaces Inference (2027)
Workload mix signalLlama2-70B most-submittedgpt-oss-120b (MoE) most-submitted; DeepSeek-R1 and VLM draw heavilyMoE + reasoning + multimodal as procurement norm
Scale-out evidenceRecord 288 accelerators (v6.0)512 accelerators (Crusoe), ~5.8M tok/s; 16 multi-node; cross-vendor + trans-Pacific heterogeneityRack-scale and federated inference normalized in buying decisions
Software vs hardwareSoftware gains anecdotal (vendor self-reports)Neutral-source evidence: CoreWeave +19.8% (GB200, identical hw), Lambda +8.85% (Blackwell Ultra, Server), Intel +36% (Arc Pro B70, Server)"Optimize-then-buy" becomes data-backed standard practice
⚙

How it works

MLPerf Inference mechanics. Systems run fixed model/dataset workloads in the Closed division (model mathematically equivalent to the reference — apples-to-apples hardware/framework comparison) or Open division (any model/harness), across scenarios (Datacenter: Offline, Server, Interactive; Edge: SingleStream, MultiStream, Offline). Availability categories: Available (buyable now) vs Preview (expected commercially available by the next round; the vendor must resubmit as Available later). Results pass working-group peer review; submitters publish reproducibility artifacts in the public results repository.

End-to-End RAG (datacenter, new). Measures a complete production-style RAG system rather than a single model. Two workloads: ingestion — parse and chunk a frozen corpus of 2,515 Wikipedia HTML files into ~107,484 passages of 768 characters (32-char overlap), embed with e5-base-v2 (110M), index into a FAISS HNSW graph; and QnA — the scored loop: query rewriter (GPT-OSS-120B) decomposes the query into up to 3 sub-queries → embed → retrieve → rerank (ColBERTv2) → grade documents (GPT-OSS-20B) → sufficiency check (GPT-OSS-120B) → answer generation (GPT-OSS-120B), iterating up to 5 hops across 824 multi-hop FRAMES queries. Metrics: documents/second (ingestion) and tasks/second (QnA — one task can involve a dozen or more LLM calls, which is why tasks/sec rather than tokens/sec). Accuracy: a Llama-3.1-8B judge (deliberately a different model family to avoid self-preference bias) scores final answers; reference answer accuracy is 35% over the full set and a submission is valid at ≥97% of reference; a DB-integrity check (retrieval of the required Wikipedia URLs) verifies submitters built equivalent vector databases. Compliance TEST09 guards performance-mode output truncation (mean output token length within ±10% of the 273.81-token reference). Reference implementation is open source (mlcommons/inference e2e-rag/, ~283 GB of models/data, frozen corpus mandated, CUDA/ROCm/XPU/CPU supported).

Edge Agentic Inference (edge, new). Targets the on-device coding-assistant pattern: one interactive user, one accelerator, fixed memory/power, growing context. Reference: Qwen3.6-27B (Apache-2.0, released Apr 22, 2026) as a Q4_K_M GGUF (~16.5 GB) served by llama.cpp, 32K served context, reasoning off, temperature 0, max_new_tokens 1024. Two gated components: (1) accuracy — Berkeley Function Calling Leaderboard v4 (BFCL v4) single-turn set (~995 samples across non_live/live/hallucination categories), pass threshold 0.97× of reference (reference Overall 86.23% → pass ≥83.64%; Normalized 87.96% → ≥85.32%); (2) performance — deterministic replay of 20 recorded agentic-coding trajectories (1,007 turns) from SWE-bench-style repositories (peak input ~23.5K tokens, inside the 32K window so no trajectory overflows), served single-stream (target_concurrency 1), reporting TTFT, TPOT and end-to-end latency per turn (p50/p90/p99/max) plus a zero-cost inline multiset-IoU tool-call check during the latency run. The design adapts the methodology of the forthcoming datacenter MLPerf Agentic benchmark to the edge budget. First-round field: 5 systems — NVIDIA Jetson AGX Thor (TensorRT Edge-LLM: 52.33 tok/s, median TTFT 247.12 ms, median TPOT 14.68 ms, full suite in 24m36s vs 2h37m for the llama.cpp Q4_K_M reference on the same board, BFCL v4 87.94% — vendor-published, MLCommons-verified entries), HPE ProLiant DL345 Gen12 (4× RTX PRO 4500) and DL145 Gen11 (1× RTX PRO 4500), NVIDIA DGX Spark (GB10), plus Lambda's Open-division HGX B200/Kimi K2.6 run.

API-centric harness and the Endpoints succession. A decoupled client drives a live model-serving endpoint over standard APIs (HTTP/gRPC); the system under test is "simply a URL." Sixteen of 30 v6.1 submitters used it (vs one in v6.0). It underlies MLPerf Endpoints, which replaces fixed operating points with a swept-load Pareto curve, moves to rolling submissions, and is slated to replace Inference for the datacenter in 2027 (v1.0 with normalization and expanded agentic workloads; rolling submissions open October 2026 per StorageReview).

!

Why it matters

🎓 For Explorer

MLPerf Inference v6.1 matters because the industry's measurement standard finally caught up with what production AI actually is: multi-model RAG pipelines, multi-turn agents, and rolling API-centric evaluation. A record 30 submitters, 120 systems and 486 results delivered the first neutral, peer-reviewed numbers for Vera Rubin NVL72 (preview), a 512-accelerator system at ~5.8M tokens/second, cross-vendor and trans-Pacific heterogeneous deployments, and documented software-only gains of 8-36% per cycle. Two new workload categories (E2E-RAG and Edge Agentic Inference) become official benchmarks for the first time, and MLCommons committed to MLPerf Endpoints replacing Inference as the datacenter benchmark. The durable consequence is that procurement-grade, apples-to-apples AI performance measurement has arrived for pipelines and agents, so benchmark-led buying becomes the rational default.

✦

What became possible?

🎓 For Explorer

Enterprises, engineers and investors can now:

  • Compare RAG pipelines and agentic edge devices with vendor-neutral numbers: E2E-RAG measures documents/second and tasks/second against FRAMES with a fixed Llama-3.1-8B judge, gated at 97% of reference; Edge Agentic measures TTFT/TPOT and BFCL-v4-gated accuracy over a deterministic 1,007-turn replay.
  • See verified first numbers for Vera Rubin NVL72 (preview, NVIDIA + Nebius), AMD MI350P, Intel Arc Pro B70 and Ryzen AI Max+ 395.
  • Put RFP checkboxes around pipeline/agentic metrics instead of single-model rows, and use neutral-source software-gain evidence (CoreWeave +19.8%, Lambda +8.85%, Intel +36%) as leverage in vendor negotiations.
  • Prepare for the Endpoints regime: rolling submissions open October 2026 and Endpoints replaces Inference as the datacenter benchmark in 2027.
◎

Implications

Technical

The suite now measures multi-component systems, not single models. E2E-RAG runs a five-hop pipeline (query rewriter, embeddings, FAISS HNSW retrieval, ColBERTv2 reranker, two-token-graded LLM tiers and a judge) over 2,515 Wikipedia pages chunked into ~107,484 passages; Edge Agentic runs Qwen3.6-27B Q4_K_M GGUF via llama.cpp against a 1,007-turn replay with a zero-cost inline multiset-IoU tool-call check. The API-centric harness (16 of 30 submitters) decouples the client from the serving endpoint, making the system under test "simply a URL" — the foundation for rolled/normalized end-to-end sweep evaluation. Preview-category hardware (Vera Rubin NVL72) obligates resubmission as the Available class, and post-deadline artifacts are explicitly not verified by MLCommons.

Developer

Developers get concrete, public reference implementations to reproduce: the E2E-RAG reference in mlcommons/inference (e2e-rag/, ~283 GB, CUDA/ROCm/XPU/CPU) and the Edge-Agentic replay harness (Qwen3.6-27B + llama.cpp, 32K context, BFCL v4 gate). Recommended: build a lab/baseline from the agentic replay first (one accelerator is enough), instrument TTFT/TPOT per turn, then optionally run the heavier E2E-RAG ingestion. Treat preview numbers as unverified for procurement.

Enterprise

Enterprise buyers and cloud/platform procurement can now request E2E-RAG and Edge-Agentic numbers in RFPs, require preview-label disclosure, and track Endpoints-normalized curves once rolling submissions open. Neutral software-gain baselines support "optimize-then-buy" negotiation leverage. Standing caveat: preview labels obligate vendor resubmission, so buyers should not lock in preview-generation hardware without re-baselining.

Strategic

The round de-risks the Rubin narrative within RAG/agentic-heavy workloads while simultaneously giving AMD its strongest cross-round scale-out evidence (512-GPU MI355X, 5.8M tok/s, ~31%/~27% software gains on same clusters). Ecosystem-level, the shift to normalized, rolling, API-centric Endpoints evaluation changes how AI performance is sold: winners accrue to hardware that generalizes across stacks and to software/optimization specialists who monetize measured gains.

⚠

Risks & limitations

Risks

Preview-label risks: Vera Rubin NVL72 numbers are pre-GA with post-deadline software, and power-submission status is unproven. Benchmark-design risks: the FRAMES judge threshold (valid at 97% of a 35% reference) could flatten small-model differentiation claims, and a small accuracy band plus degradation guards (TEST09 length check) are not yet hardened into public debate. Leadership claims are model- and scenario-selected; each round's precedence table gets refreshed — the reason MLCommons standardized normalization.

Limitations

Assumptions and weaknesses recorded in research: vendor-origin "up to" multipliers are COMPANY CLAIM (verified raw, comparative framing vendor-authored; diyai.io shows the headline 3.7x is Interactive-scenario-bound, ~1.74-2.0x Offline/Server). Endpoints timing specifics come from StorageReview's write-up, consistent with but not directly on MLCommons' announcement page. Some secondary outlets are mid/low-authority aggregators used only for structural framing. Not all 486 result rows were individually audited; headline numbers and repeated claims were traced to the announcement/chairs analysis/vendor material as labelled. Known unknowns: exact Endpoints v1.0 normalization, Rubin GA dates, preview power-status, and independent normalization of the 5.8M tok/s record.

?

Open questions

  1. When does Vera Rubin NVL72 move from preview to available, and do GA numbers match the preview 3.7×/2.5× Interactive-heavy claims at Offline/Server workloads?
  2. Will the verified 5.8M tok/s gpt-oss-120b Offline record hold under independent normalization (same-hardware, same-batch)?
  3. How do the new E2E-RAG tasks/sec and Edge-Agentic TTFT/TPOT numbers converge across Accelerator→CPU/XPU and bed-of-NVIDIA lanes in the next round?
  4. Do Intel's E2E-RAG exclusivity and 36% software gains convert into enterprise CPU+GPU procurement share, or stay a novelty?
  5. What fraction of the "AMD leads the largest systems" positioning survives a round that includes NVIDIA's full GA Vera Rubin fleet and Rubin-optimized stacks?
  6. Will BFCL v4 gate tighten beyond the 0.97× reference, and will judge-methodology changes (LLaMA-3.1-8B, FRAMES) generalize to enterprise RAG data?
  7. Which submitters convert Endpoints rolling entry (October 2026) into standing normalized presence, and does it cannibalize Inference rounds or complement them?
  8. What is the power/energy submission status of the preview Vera-Rubin and 512-GPU systems — pure-performance or efficiency-accounted entries?
↗

What happens next?

🎓 For Explorer
  • October 2026: MLPerf Endpoints rolling submissions open (static results available now; normalized excluding-power Pareto curves to follow); first Fall 2026 round submissions windowed; industry webinars/workshops publish follow-up analyses of v6.1 tables.
  • Late 2026 / 2027: Fall 2026 round results (expected: broader Rubin fleet resubmission in available category, more E2E-RAG and Edge-Agentic entries, first Endpoints v1.0 normalized outputs); the MMA (MLPerf Agentic) datacenter benchmark advancing; potential tightening of BFCL-v4 gates and expansion of E2E-RAG to multi-model families.
  • 2027: MLPerf Endpoints replaces Inference as the datacenter benchmark per Kanter's statement — the end of the fixed-round regime for datacenter claims.
  • Ongoing watch items: preview→available transitions for Rubin/MI355X-class silicon; power/energy accounting maturity; judge-model selection and accuracy-gate debates; adoption of E2E-RAG/Edge-Agentic numbers in real RFPs.
★

Editorial takeaway

🎓 For Explorer

MLPerf Inference v6.1 is one of the most consequential benchmark releases in the suite's history — not because any single vendor "won," but because the industry standard finally caught up with what production AI actually is: multi-model RAG pipelines, multi-turn agents, and rolling API-centric evaluation. A record 30 submitters, 120 systems and 486 results delivered the first neutral, peer-reviewed numbers for Vera Rubin NVL72 (preview), a 512-accelerator system at ~5.8M tokens/sec, cross-vendor and trans-Pacific heterogeneous deployments, and documented software-only gains of 8–36% per cycle. The headline 5.7× (DeepSeek-R1) and 2.99× (VLM) per-accelerator improvements, and NVIDIA's "up to 3.7×" Vera-vs-GB300 claim, are real but scenario- and stack-specific — preview labels on the most explosive numbers mean the durable commercial substance is the new workloads (E2E-RAG, Edge Agentic) and the Endpoints succession, which will reshape how AI performance is measured, marketed and bought from 2027 onward. Enterprise readers should take away a single practical line: procurement-grade, apples-to-apples AI performance measurement has arrived for pipelines and agents — benchmark-led buying is now the rational default.

A question draws document fragments from a shelf that assemble into a grounded answer block, clamped at checkpoints and hung with citation tabs.
⌘

Lab: VERIFY