News Weekly
LV 10 XP
0% read
Your progress · 0/5 chapters
About 6 min total
ModelsISSUE #2 · STORY 11 OF 20Sep 18, 2026CONFIRMED

Alibaba's new Qwen model hears, sees and costs very little

Alibaba released Qwen3.8-Omni-Flash, a model that reads text, images, audio and video at one very low price. It ships with free tools you can install today.

Illustration: a single bright flash of light passes through a wide abstract sensor-array — clean geometric ear, lens and antennae forms arranged in a ring — an artistic impression of a real-time omni-modal model…

Read it your way

CHAPTER 1 · THE 60-SECOND VERSIONPicked for Explorers

A cheap model that sees and hears

Alibaba's Qwen team launched Qwen3.8-Omni-Flash on Sep 18, 2026. It accepts four kinds of input and returns written answers. The twist is its price, which sits at the level of plain text models.

One model, four inputsQwen3.8-Omni-Flash takes in text, images, audio and video, then returns written answers.
Priced like plain textThe international list price is $0.15 per million input and $0.47 per million output tokens.
A million token memoryIt carries a 1M token context window to hold long documents and hours of media.
Open tools, closed brainThe model is API only, but Alibaba released free Apache-2.0 harness tools.
Finish this chapter for +15 XP
Flip the switch

From premium omni to commodity

YOU GETOne flat token priceAudio, video and text all bill at the same low rate.
YOU GETA full realtime stackA model, a live API and a desktop harness shipped together.
YOU GETOpen tools to installApache-2.0 harness and plugins run on your own machine.
Play with the numbers · +10 XP

What your input might cost

DRAG THE SLIDER
Qwen3.8-Omni-Flash
$3
about $0.15 per million
Older omni model
$8
was $0.40 per million
You'd save
$5
every month

International input prices per million tokens from the official listing. The compare row uses the previous Qwen3.5-Omni-Flash rate.

Your next move · as a Explorer

Try the tooling, wait on the API

1Verify specs and prices from the official pages
2Rebuild the per-token cost math yourself
3Install the open harness only if you have budget

Switch your reading mode at the top to see a different next move.

Tap to open

Things to keep an eye on

Pop quiz · unlock the Omni Scout badge

Did it stick?

0/3
What can Qwen3.8-Omni-Flash output on its main API?+20 XP
Which input price is listed for the international tier?+20 XP
What is true about the model's weights?+20 XP
Your call · +5 XP

Should you test a cheap omni model on your own media workloads?

Deep dive

The full research, labeled and sourced

CONFIRMED20 sources · 88 min
Story identity
FieldValue
Story IDS11
TitleAlibaba Cloud releases Qwen3.8-Omni-Flash: 1M-context real-time omni model at $0.15/$0.47
OrganizationAlibaba Cloud / Qwen (Qwen Team; announced via the Qwen official blog and Qwen's X account @Alibaba_Qwen)
CategoryModel Release
Event date2026-09-18 (Qwen blog "Qwen3.8-Omni-Flash: Omni Senses. Agentic Delivery." dated 2026/09/18; Alibaba Cloud Model Studio docs page "Last Updated: Sep 18, 2026"; Gate News flash 2026-09-18 07:03 UTC; TechNode and MarkTechPost both dated 2026-09-18)
Announcement date2026-09-18 (same as event date; companion releases inside the window: Qwen3.8-Omni-Flash-Realtime + Qwen-Live-Harness v1.0.0 on 2026-09-21 per the harness repo's own News entry; arXiv technical report 2609.25611 submitted 2026-09-22)
Article dates2026-09-18 (MarkTechPost, TechNode, Neowin, El Economista, DataCamp, Gate News ×2), 2026-09-19 (Pandaily publish-date metadata, Oflight), 2026-09-20 (Alibaba Cloud Community blog mirror), 2026-09-21 (liteLLM, QwenCloud listing, EmpirioLabs realtime page, sofar/bot), 2026-09-22 (modelscale.dev, Compras…)
Alternative datingSeveral registry/directory sources list the release as 2026-09-17 (CloudPrice "Released 2026-09-17"; EmpirioLabs "Released Sep 17, 2026"; OpenCode model stats "Sep 17, 2026"; Alibaba docs' Traditional-Chinese help page shows "更新時間 Sep 17"). This is the UTC-side of the Beijing-time (UTC+8) announcement: Sep 18 before ~08:00 Beijing time is still Sep 17 UTC. Both dates are inside the window, so eligibility is unaffected; the story's consensus event date is 2026-09-18
Window check2026-09-18 ∈ [2026-09-18, 2026-09-22] — eligible (the 2026-09-21 Realtime/Harness release and the 2026-09-22 arXiv report are also inside the window)
Evidence statusCONFIRMED (official Qwen blog announcement text captured via full mirror on Alibaba Cloud Community + search-index excerpts of qwen.ai; official Model Studio spec page fetched in full; official Realtime API docs fetched in full; official QwenCloud pricing page fetched in full; arXiv abstract fetched in full; both GitHub repos fetched in full; four independent English-language coverages fetched/captured; three additional independent relays; two secondary price directories)

Corrections / refinements to the discovery record (important):

  • The "$0.15/$0.47" is the international (Singapore) tier, not a single global price. Official Model Studio pricing tables list the omni-flash SKU at $0.15 input / $0.016 implicit-cache / $0.47 output per 1M tokens for the Singapore "International" scope, and $0.113 / $0.014 / $0.382 for the other "Global" scopes (Beijing, Hong Kong, Frankfurt, Tokyo, Virginia); on the mainland the published row is CNY 0.8 / 0.1 (cache hit) / 2.7 (modelscale.dev's transcription of the official pricing page; Oflight: "Bailian's multimodal input price is ¥0.8 per million tokens, down sharply from roughly ¥18 for the predecessor"). The discovery record's "$0.15 in / $0.47 out" matches the QwenCloud/liteLLM/Vercel/CloudPrice international listing — treat "$0.15/$0.47" as the international headline rate.
  • "Real-time omni model" conflates two artifacts. The base qwen3.8-omni-flash (Sep 18) is a non-realtime HTTP API model (Chat Completions/Responses) with text-only output. The realtime variant qwen3.8-omni-flash-realtime (WebSocket/WebRTC, text and audio output, voice cloning, MCP, multichannel input) was released with Qwen-Live-Harness v1.0.0 on 2026-09-21 per the harness repo's own News entry. The story spans both; this research records both dates, both inside the window.
  • "Roughly 89% cheaper on video input than Qwen3.5-Omni-Plus" is a COMPANY CLAIM (Qwen X post / PANews relay), not an independently derivable number. Qwen's own blog states the audio-visual per-hour input price falls "by more than 93%" and audio per-hour "by more than 98%" using a disclosed methodology (hourly price = 30 × the input cost of two minutes of source material; 720p at 1 fps). The X post and PANews/Gate News relay "~89%" for video input. Independently checkable flat per-token math vs the previous generation's Flash omni (CloudPrice: Qwen3.5-Omni-Flash $0.40 in / $3.00 out) gives an input cut of 62.5% and an output cut of 84.3% — smaller than the hourly-media claim, which embeds a token-consumption-per-hour drop, not just a per-token price drop. Treat 89%/93%/98% as vendor-methodology claims; 62.5%/84.3% as the verifiable per-token floor.
  • Two small consistency frictions between the official blog and independent recaps: (a) the blog says "across 29 evaluations… average score improves by more than 25%" while TechNode reports "more than 26% across 30 evaluations"; (b) the blog's OmniVideoBench agentic-mode token reduction is "145,736 → 79,117 (approximately 45.7%)" while Gate News (relaying PANews) writes "agentic perception mode consumed 51.8% fewer tokens." Both differences are small and likely reflect rounding/eval-set scoping; the blog's exact table is the anchor.
  • No open weights were announced at launch. MarkTechPost and Oflight both flag that the model is API-only ("No open weights were announced at launch, so self-hosting is not an option"; "weights are closed and proprietary"). This contrasts with Qwen3.8-Flash (Aug 2026), which paired the API with the open-weight Qwen3.8-Flash-Next preview — the architecture on which the omni-flash is built (MarkTechPost). The open parts of this launch are the tooling: Qwen-MM-Plugins and Qwen-Live-Harness, both Apache-2.0.

✓

What happened?

🎓 For Explorer

On September 18, 2026, the Qwen Team (Alibaba) launched Qwen3.8-Omni-Flash, described in the official blog as "our next-generation native omnimodal model" built to move omni models "from understanding omnimodal content to planning tasks, calling tools, and completing creative work" (FACT, official blog). It went live the same day on the Qianwen AI Platform, QwenCloud, Qwen Studio, and Alibaba Cloud Model Studio, with an OpenAI-compatible Chat Completions + Responses API (FACT; multiple sources).

What shipped (FACT, official docs + QwenCloud + blog):

  • Input: text, image, audio, video; output: text only (non-realtime API). 113 input audio languages/dialects; multichannel (2-ch stereo, 4-ch FOA) spatial audio input via use_multichannel.
  • 1M-token context window; max input 991,808 tokens (non-thinking) / 983,616 (thinking); max output 131,072; max reasoning (chain-of-thought) 262,144.
  • Sparse MoE architecture inherited from Qwen3.8-Next (per the arXiv abstract), i.e. built on the Qwen3.8-Flash-Next lineage that Alibaba open-weighted in August 2026 (MarkTechPost).
  • Thinking on by default with reasoning_effort = xhigh/medium/low and preserve_thinking default-on; reasoning_effort: none disables thinking.
  • Tool calling (custom), built-in web search (Responses tool web_search), context caching (automatic implicit + Responses Session), structured outputs, batch calls, DashScope + OpenAI protocols.
  • Regions: Beijing, Singapore, Hong Kong, Tokyo, Frankfurt, US Virginia.
  • Pricing (international): $0.15/M input, $0.016/M implicit-cache input, $0.47/M output (QwenCloud, liteLLM, Vercel, CloudPrice); mainland China ¥0.8/¥0.1/¥2.7.
  • API-only; no open weights at launch (MarkTechPost, Oflight).

Companion releases inside the window (FACT):

  • Qwen3.8-Omni-Flash-Realtime (WebSocket + WebRTC; text+audio output; up to 100 audio turns / 50 video turns; 600 s audio / 240 s video retained in context; 196,608 max input tokens; 120-minute session cap; MCP tool support; voice cloning; voices incl. longanlingxin, default Tina) — official Realtime docs, page confirmed live Sep 23 with the model referenced throughout.
  • Qwen-Live-Harness v1.0.0 (2026-09-21 per the repo's own News entry): Apache-2.0 open-source desktop harness for the Realtime API — camera/mic interaction, background task delegation (Qwen Code, Qoder CLI, Codex, Claude Code, Gemini CLI via ACP), proactive audio/visual monitors, long-term memory; macOS 12+; npm install -g qwen-live-harness.
  • Qwen-MM-Plugins: Apache-2.0 plugin suite ("Make any agent harness multimodal-native") for Claude Code, CodeBuddy, Codex, Qoder, OpenClaw, Qwen Code, Gemini CLI + manual installs; capabilities incl. core, api, search, omni-chatcut (Music-to-MV/commentary/video translation), omni-video2note, omni-skill-creator, omni-memory, video-edit, blender, freecad. Documented gap: "most harnesses cannot yet feed audio to the main model natively — audio is handled through the API for now."
  • arXiv 2609.25611 "Qwen3.8-Omni: Towards Native Omni-Modal Agents" (Qwen Team), submitted 2026-09-22: introduces the model, the native multimodal co-training strategy, Qwen-MM-Plugins, and Qwen-Live-Harness.

Company-published performance claims (COMPANY CLAIM — MarkTechPost explicitly notes "All figures here come from Qwen. Independent results were not available at publication"):

  • Average score improvement >25% across 29 evaluations vs Qwen3.5-Omni-Plus (TechNode's recap rounds to ">26% across 30 evaluations").
  • WildClawBench-MM +36.5, AgenticVBench +22.3, UniClawBench 69.6; audio/audio-visual core gains: LongAudioSpan +8.3, OmniVideoBench +9.6, OmniCap-IF CSR +8.5 / ISR +14.1; AliMeeting DER 88.11 → 3.35, cpWER 89.61 → 17.18.
  • "Audio-visual performance close to Gemini 3.8 Flash and overall audio performance that exceeds Gemini 3.8 Flash" (Qwen-run cross-model evals, not independent).
  • Agentic-vs-static understanding on OmniVideoBench: 63.4 → 67.8 accuracy with tokens 145,736 → 79,117 (~45.7% fewer tokens); the Qwen X post (status 2100785962414702599) and PANews/Gate News relay ~89% video-input cost reduction and +19.5-pt average agentic gain.
  • Qwen2.5-Omni-3B self-improvement demo: within 12 hours the model lowered the 3B model's Sichuan-dialect CER 25.79% → 15.30% (~40.7% relative) by building 3,413 training examples across 4 experiment rounds.

Independent measurements (INDEPENDENT EVIDENCE, partial): modelscale.dev (2026-09-22) shows observed benchmark-aligned capability rows for the hosted API — Coding 54.9, Knowledge 53.9, Multimodal & Grounded 85.1, Instruction Following 87.7 (BenchLM-sourced; no composite score). No Artificial Analysis-style leaderboard entry for the omni-flash was found as of Sep 23; no independent replication of the audio/video benchmark claims exists yet.


Δ

What changed?

  • Omni (audio/video) capability was brought to the same $0.15/$0.47 price tier as text-and-image Flash models. Qwen3.5-Omni-Flash was $0.40/$3.00 and Qwen3.5-Omni-Plus charged separate (much higher) audio and video-input rates; the omni-flash bills all input modalities at one token rate (official pricing note via modelscale: "one token rate for every modality"). Omni understanding is now priced like ordinary Flash text inference — the cost floor for voice/video agents reset. (FACT on prices; the strategic reading is INTERPRETATION.)
  • The previous generation's realtime audio input premium collapsed. Qwen3-Omni-Flash-Realtime (Dec 2025) charged $4.57/M audio input on QwenCloud; the new realtime model's audio input lists at ~$1.86/M audio tokens (EmpirioLabs mirror of official rates) — a ~59% per-token cut, before the much-larger per-hour claims (>98% audio / >93% audio-visual) that embed token-consumption-per-hour improvements.
  • The reference competitor for Chinese omni is now Gemini 3.8 Flash, and Qwen is claiming wins on audio. Qwen's own cross-model evals claim "overall audio performance exceeds Gemini 3.8 Flash" with audio-visual "close to" it; the launch-week Anglo press echoed it (Neowin: "undercuts Gemini on audio"; El Economista: "surpassing Google's Gemini 3.8 Flash in voice processing"). (COMPANY CLAIM + relay; no independent verification yet.)
  • Agentic perception became a first-class documented workload: the model plans, watches/listens selectively, and gathers evidence over several coarse-to-fine rounds rather than ingesting whole files (blog; OmniVideoBench 63.4 → 67.8 with ~45.7% fewer tokens). (COMPANY CLAIM, direction is the launch thesis.)
  • Alibaba open-sourced the layers around the closed model (Apache-2.0 Qwen-MM-Plugins + Qwen-Live-Harness), reversing the usual open-model/closed-tooling split: the model itself is API-only while the harnesses are free and community-driven. (FACT, licenses verified on both repos.)
  • A technical report landed within the same week (arXiv, Sep 22) — faster than most Chinese-lab launches, and it names the architecture (sparse MoE from Qwen3.8-Next, 1M context) and the co-training strategy. (FACT.)

↔

Before → Change → After

🎓 For Explorer

Before (through Sep 17, 2026):

  • Alibaba's omni line: Qwen3.5-Omni-Plus and Qwen3.5-Omni-Flash (interactive, audio+video input; audio output on realtime variants; 262K context; $1.15/$8.75 per-M audio-heavy billing for Plus in Beijing tier, $0.40/$3.00 Flash international); Qwen3-Omni-Flash (2025-12) with realtime audio at $4.57/M and tiny 8-dialog-turn sessions.
  • Qwen3.8-Flash (Aug 26) had already set the $0.15/$0.47 tier for text/image/video understanding with the open-weight Flash-Next preview, but it was not an omni model (no audio input).
  • No 1M-context omni API at ~$0.15/M input existed from any vendor; old omni models charged per-modality input premiums.
  • Qwen had no dedicated realtime agent harness; Western realtime voice offerings (e.g., Gemini Live / GPT-realtime-class APIs) were the default references.

Change (Sep 18–22, 2026):

  • Qwen3.8-Omni-Flash released: text/image/audio/video in, text out, 1M context, thinking default-on, tool calling, web search, caching — at $0.15/$0.016/$0.47 international (¥0.8/0.1/2.7 mainland).
  • Sep 21: Qwen3.8-Omni-Flash-Realtime + Qwen-Live-Harness v1.0.0 (Apache-2.0 desktop harness, ACP delegation, memory, proactive monitors).
  • Qwen-MM-Plugins (Apache-2.0) expanded with omni capabilities (chatcut, video2note, skill-creator, memory).
  • Sep 22: arXiv report 2609.25611 published (Qwen Team) detailing co-training, MoE, harness architecture.
  • Independent record: catalogues mirrored pricing within hours (Vercel, LiteLLM day-0 support on Sep 18–21, CloudPrice, EmpirioLabs); capability registry (modelscale) filled in observed rows; no independent benchmark replication of audio/video claims yet.

After:

  • "1M-context omni at Flash text prices" is now a shelf item, and the Qwen realtime stack (Realtime API + Live-Harness + MM-Plugins) is a coherent competitor to Western realtime voice/vision agent stacks — a Chinese-lab full offering from model to harness.
  • Enterprises evaluating omni agentic workloads now compare against a $0.15/$0.47 baseline with 1M context instead of per-modality-premium pricing; the old generation's pricing looks locked to 2025.
  • Open-weight positioning questions shifted to the tooling: the model is closed, the harness is open — a split that is now a talking point for procurement and for the open-AI community. (INTERPRETATION)

⚙

How it works

  • Architecture (FACT as declared): sparse mixture-of-experts inherited from Qwen3.8-Next, extended to a 1M-token context (arXiv abstract); the Qwen3.8-Flash-Next base (125B total / 6B active per token, 51B N-gram embedding parameters, Gated DeltaNet + Qwen Sparse Attention, per August coverage) is the publicly-described ancestor. Exact omni-specific parameter counts were not disclosed at launch (COMPANY CLAIM territory; unverifiable without weights).
  • Native multimodal co-training (COMPANY CLAIM, per the arXiv abstract): the model was co-trained across modalities "preserving strong text-domain capabilities while facilitating the transfer of agentic capabilities from text to audio and video tasks" — the mechanism behind the "1M context without text degradation" claim.
  • Thinking/effort control (FACT, docs): thinking enabled by default; reasoning_effort = xhigh (default) / medium / low; preserve_thinking default-on; none disables thinking. Reasoning tokens count against the 262K max chain-of-thought and are billed as output tokens.
  • Agentic omni understanding (COMPANY CLAIM): for hours-long media, the model starts from the question and performs coarse-to-fine evidence gathering — deciding what to watch/listen to rather than processing everything; blog reports 63.4 → 67.8 accuracy on OmniVideoBench at ~45.7% fewer tokens (145,736 → 79,117).
  • Realtime pipeline (FACT, docs): WebSocket (native + DashScope/Java SDKs) or WebRTC (browser, UDP, echo cancellation); server-side VAD or manual turn submission; streaming input of PCM audio (16 kHz) and video frames (recommended 1 fps); output audio at 24 kHz PCM; session cap 120 min; audio 600 s / video 240 s retained; 100 audio / 50 video turns; multichannel audio (1/2/4 channels with raw_mic_array/foa_ambix); video aggregation flag representation_compact; MCP Streamable-HTTP tools with approval; voice list incl. longanlingxin + voice cloning.
  • Qwen-Live-Harness (FACT, repo): two local processes — Node daemon (models, tools, memory, delegation) + Electron Host (UI, capture, playback) over a local WebSocket; background harnesses (Qwen Code, Qoder, Codex, Claude Code, Gemini CLI) over ACP/REST/SSE; Proactive monitors (1 fps, 2-second audio chunks — explicitly "not a safety-critical alarm system"); local memory libraries with cloud consolidation.
  • Qwen-MM-Plugins (FACT, repo): skills + optional MCP servers per capability; core lets the main model read local images/video natively; omni capabilities route media understanding through the DashScope API (harnesses can't feed audio to the main model natively yet).

!

Why it matters

🎓 For Explorer
  1. The cost floor for omni/voice/video agents just moved to Flash-text pricing. A 1M-context model that ingests audio and video at $0.15/M input (¥0.8/M on the mainland) removes the per-modality premium that made omni agents expensive to operate at scale; at these rates, whole-hour media understanding becomes a background utility cost, not a line item. (FACT on pricing; significance is INTERPRETATION.)
  2. Qwen's realtime stack is now a full vertical — model (Realtime API), desktop runtime (Live-Harness), and agent-harness plugins (MM-Plugins), all Apache-2.0 except the model — a credible counterweight to Western realtime voice/vision offerings in the same week (Gemini Flash Live line, GPT realtime-class APIs). (FACT on artifacts; the competitive reading is INTERPRETATION.)
  3. The "89%/93%/98% cheaper" framing is the kind of claim buyers can now audit — the flat per-token cut vs the previous Flash omni is 62.5%/84.3%, while the hourly-media claims (>98% audio, >93% audio-visual, ~89% video) embed token-consumption-per-hour gains with a disclosed methodology. Teams planning budgets should quote the per-token numbers and treat the hourly numbers as directionally vendor. (Arithmetic + INTERPRETATION.)
  4. Agentic media perception is the product thesis, not captioning. The launch is explicitly positioned as agent-with-tools work on audio/video (video editing, MV creation, film commentary, meeting action-items, video deep research) — a step-change in how Chinese labs talk about omni models. The OmniVideoBench 63.4→67.8-at-45.7%-fewer-tokens figure is the load-bearing demo. (COMPANY CLAIM, but it is the strategic core of the launch.)
  5. Chinese labs now set the audio/omni price benchmark. Within the same window as StepFun's Step 5 Preview (S10) and Xiaomi's MiMo-V2.6 (S09), Alibaba re-anchors the "capable omni at commodity prices" position — reinforcing the week's pattern that the price-setting frontier is increasingly in China. (INTERPRETATION grounded in S09/S10 data within this project.)

✦

What became possible?

🎓 For Explorer
  • Voice/video agent front-ends at text-tier token prices: 1-hour meeting transcription + summarization, video deep-research reports, media asset pipelines at $0.15/M-class input rates instead of per-modality premia.
  • Long-context omni reasoning over ~1M-token media+text contexts — hours of audio/video (file limits: video up to ~2 h/2 GB by URL, audio up to ~3 h per MarkTechPost's reading of the docs; blog cites "up to one hour" for the flagship meeting scenario — confirm per-region when quoting).
  • A desktop realtime agent with eyes and ears, installed in one command: npm install -g qwen-live-harness on macOS 12+ gives camera/mic conversational control of Claude Code/Codex/Qwen Code/Gemini CLI background agents, proactive monitor triggers, and persistent memory — no local GPU needed (cloud inference).
  • Turning videos into reusable skills and documents on open tooling: omni-skill-creator (demo → Skill.md), omni-video2note (tutorial → illustrated PDF), omni-chatcut (MV/commentary/translation), omni-memory — Apache-2.0, installable into everyday agent harnesses.
  • Multichannel spatial audio understanding (2-ch/4-ch FOA) — new ground for API omni models, enabling directional/localization applications ("locating targets by sound" per the blog).
  • Not possible: self-hosting the model (no weights announced), audio output on the non-realtime API (text only), and fully-native audio in third-party harnesses (routed via API for now).

◎

Implications

Technical

  • Billing granularity changed: one token rate for every input modality (per the official pricing note), so audio-heavy prompts no longer carry a separate "audio" price column — cost modeling simplifies but hides per-modality tokenization ratios. Watch for per-region discrepancies ($0.15 Singapore vs $0.113 other Global vs ¥0.8 mainland).
  • Dated registry drift is present across sources: some catalogues still say "Released Sep 17" (UTC-side dating; CloudPrice, EmpirioLabs, OpenCode pages) — pin capture dates when citing release dates.
  • Realtime constraints are concrete: 196,608-token input cap on the realtime endpoint (vs 1M on the offline model), 120-minute session ceiling, 600 s audio / 240 s video context retention, 100/50-turn caps, workspace-scoped WebSocket endpoints (wss://{WorkspaceId}.…maas.aliyuncs.com/api-ws/v1/realtime). Long-horizon realtime designs must engineer around these.
  • Web search and tool calling are mutually exclusive on the realtime endpoint (docs), and turn-detection/VAD options differ by model family (semantic VAD available on 3.5-era realtime, server VAD elsewhere) — API-parity traps for code written against older Qwen realtime models.
  • No weights ⇒ no local verification of the MoE config, co-training claims, or audio/video benchmark numbers; the open artifacts are harnesses (deterministic to inspect) rather than the core model.
  • The 29-vs-30-eval and 45.7%-vs-51.8% figure frictions across official/relay coverage show that vendor tables and media recaps diverge slightly; independent replication is the missing layer.

Developer

  • Add qwen3.8-omni-flash to routing matrices this week: 1M context, text/image/audio/video in, text out, $0.15/$0.016/$0.47 (international), OpenAI-compatible Chat Completions + Responses, DashScope SDK; day-0 proxy support via LiteLLM (dashscope/qwen3.8-omni-flash, PR #41754) and Vercel AI Gateway (alibaba/qwen3.8-omni-flash).
  • Remember the output is text-only: if your product needs spoken replies, chain qwen3.8-omni-flash-realtime (audio out) for live interaction or bolt on a TTS stage for the offline API (Oflight's pipeline guidance).
  • Tune reasoning_effort (xhigh/medium/low) and preserve_thinking — a direct cost/depth dial; thinking tokens are billed as output and count against the 262K CoT cap.
  • Exploit the cache: $0.016/M implicit cache reads make stable-prefix long-context loops (~1M media-adjacent prompts) genuinely cheap.
  • Try Qwen-MM-Plugins now (Apache-2.0, guided installer for Claude Code/Codex/Qoder/OpenClaw/Qwen Code/Gemini CLI) — install core + api + an omni capability; note the audio-native gap ("audio is handled through the API for now") before promising audio workflows inside your harness.
  • Realtime integration cautions: workspace-scoped endpoints and region-locked API keys (Beijing vs Singapore, non-interchangeable); MCP tools need approval; representation_compact must be set before the first audio segment; voice changes to the 3.8 voice list (e.g. longanlingxin) don't carry over from earlier voice sets.
  • Eligibility/version hygiene: confirm region-specific max file sizes and the 1-hour language/meeting claims against the live docs; snapshot pricing (promo vs original) before quoting in procurement.

Enterprise

  • A benchmarked reference price for omni agent workloads: enterprises can now cost out voice/video agent pilot economies at $0.15/$0.47 with 1M context — a strong negotiating reference for any omni vendor conversation (regardless of adoption).
  • Candidate workflows: meeting-minutes-and-action-items (speaker diarization across audio+video, DER 88.11→3.35 claimed), video knowledge bases (video2note → PDF), content localization (short-drama translation with voice cloning via the realtime line + omni-chatcut), customer-service realtime agents (Skills injection + MCP).
  • Procurement posture: API-only (no weights, no self-host, no fine-tuning listed; fine-tune "Unsupported" on comparable Qwen docs rows), data flows to Alibaba Cloud regions (Beijing or Singapore) — document data-residency posture for regulated environments; a US (Virginia) region exists, but compliance review still applies.
  • Vendor dependency risk: single-vendor stack (model + harness + plugins all Qwen), regional pricing variance (¥0.8 vs $0.15 vs $0.113 tiers), and the usual Chinese-cloud terms/TOS review (data use, retention, export-control context).
  • Keep the S9/S10 pattern in mind: enterprise comparisons this week should include MiMo-V2.6-Pro/GLM-5.3/Kimi/Step 5 (text/coding agentic) and Qwen3.8-Omni-Flash (omni) as distinct rows — per-category cost floors, not a single leaderboard.

Strategic

  • Qwen's play is "omni at commodity price + open harness moat." The model is API-rented while the ecosystem glue (Live-Harness, MM-Plugins, Apache-2.0) accumulates community contributions, stars (harness 96, plugins ~3k as of Sep 23) and workflow lock-in around a closed core. (FACT on licenses/usage; the moat reading is INTERPRETATION.)
  • The Gemini 3.8 Flash comparison is deliberate. Qwen's own evaluations benchmark against Gemini 3.8 Flash and claim audio leadership — Alibaba is publicly picking a fight with Google on audio, the modal front where Western vendors still lead. (COMPANY CLAIM + INTERPRETATION.)
  • Realtime voice is consolidating into full verticals: Qwen now ships model+harness+plugins; the same pattern is visible across the Western realtime landscape — an architectural arms race around "always-on multimodal agents," not just model scores.
  • Expect a pricing-response cycle: with omni at $0.15/$0.47, competing omni/video-understanding APIs face pressure to match; next Qwen releases (likely a Plus/Max omni sibling and a non-realtime "flash-next" open-weight omni, mirroring the Aug Flash/Flash-Next split) are plausible within the quarter. (PREDICTION.)
  • The open/closed split becomes a policy talking point: a Chinese lab open-sourcing the harness layer while keeping the highest-value model closed is a new shape for the "open weights" debate; regulators and enterprises will weigh ecosystem capture vs vendor concentration. (INTERPRETATION + PREDICTION.)

⚠

Risks & limitations

Risks
  • All headline benchmark claims are vendor-run, some on in-house benchmarks (UniClawBench, AgenticVBench, WildClawBench-MM are Qwen-family evals). No independent replication existed as of Sep 23; the "exceeds Gemini 3.8 Flash on audio" claim rests on Qwen's own harnesses. (Risks of over-commitment on unverified scores.)
  • API-only with no weights announcement — no self-host path, no community verification of architecture or safety-relevant behavior, and full dependency on Alibaba Cloud availability/terms.
  • Data-governance exposure for regulated users: audio/video content (meeting recordings, customer calls) flows to Alibaba Cloud regions; combined with cross-border ambiguity for China-hosted data — the most sensitive enterprise workload here is exactly the one with the newest data-handling questions.
  • Pricing volatility: intro/single-rate structure may not survive; region-tier variance (¥0.8 vs $0.15 vs $0.113) invites billing surprises in multi-region deployments; "one token rate for every modality" could be revised once real usage data lands.
  • Realtime reliability unknowns: session caps (120 min), context-retention truncation (600 s/240 s), and VAD/turn-detection maturity in noisy production environments — the official docs' own monitoring disclaimer for the harness ("not a safety-critical alarm system") generalizes to the stack.
  • Ecosystem-supply risk: Live-Harness is macOS-only at v1.0.0 (Windows/Linux "in progress"); harness code is young (96 stars, 64 commits); MCP approvals default-on adds friction and an approval UX burden.

Limitations
  • No independent verification of any audio/video capability claim — no AA-style leaderboard entry, no third-party benchmark re-runs as of research date (Sep 23); modelscale.dev's observed rows (Coding 54.9 / Knowledge 53.9 / Multimodal & Grounded 85.1 / IF 87.7) are registry aggregates, not deep evals of the omni claims.
  • The 89%/93%/98% cost-reduction figures are methodology-dependent (vendor's per-hour estimation: 30 × 2-minute-input cost; 720p@1fps) and exceed the flat per-token cuts (62.5%/84.3% vs the previous Flash omni) — quote with the floor numbers.
  • The base model outputs text only — "omni" here means omni input for the offline API; audio output requires the realtime variant or a TTS stage.
  • Runtimes/limits constrain the "1M context" headline in realtime mode (196,608-token input cap on the realtime endpoint), and the file/duration caps for media input need per-region confirmation.
  • No technical-depth public disclosure at launch — exact params, tokenization ratios per modality, and training data were not published with the announcement; the arXiv report (Sep 22) is the intended technical anchor but was not deeply reviewed in this research beyond the abstract.
  • Ecosystem immaturity: harness macOS-only; harnesses cannot pass audio natively to the main model (MM-Plugins routes via API); MCP/approval flows are new.
  • Dating drift across sources (Sep 17 vs Sep 18) and Chinese-source relay variance (Gate News' 51.8% vs the blog's 45.7%) mean exact figures should be re-pinned against the official pages at quote time.

?

Open questions

  1. Which price actually applies where — and does the "one token rate for every modality" survive Q3 billing statements (audio tokenization ratios per hour of media are undisclosed)?
  2. Will Alibaba ship an open-weight omni checkpoint (Flash-Next-style split) — and would it be Apache-2.0 like the August preview?
  3. Do independent evals confirm "overall audio performance exceeds Gemini 3.8 Flash," or is it harness/protocol-specific (Qwen ran both models; prompt/params/methodology matter)?
  4. What are the real per-hour-of-media token counts for the new model (the 45.7% vs 51.8% token-consumption figures differ), and what do they imply for the 89/93/98% hourly-cost claims on real workloads?
  5. How do long-context (near-1M) media prompts behave on the offline API — TTFT, memory, coherence over multi-hour inputs — in third-party workloads?
  6. Does the realtime endpoint hold up under load (turn detection, latency, cost per 10-minute conversation with voice cloning + MCP), and what are true per-call bills?
  7. Will the UI/web-search/audio integration gaps in third-party harnesses (MM-Plugins' audio routing) close quickly — i.e., is the ecosystem moat real or a patch?
  8. Where is this on Alibaba's strategic map vs Qwen3.8-Max/Flash and the Apsara-conf week announcements (full-stack AI roadmap coverage adjacent to this launch)? (Open for the synthesis phase.)

↗

What happens next?

🎓 For Explorer
  • Days (to end of September): registry mirrors converge on the Sep 17/18 dating; first community reviews of Live-Harness (macOS-only limits) and MM-Plugins; liteLLM/Vercel and other gateways finish billing-map rollouts; early pilot reports on real per-hour audio/video token counts start surfacing; Qwen3.8-LiveTranslate follow-up posts sit adjacent in the Qwen blog feed.
  • Weeks (October): expect (a) the first third-party audio-visual benchmark re-runs vs Gemini 3.8 Flash — the decisive check on the flagship claim; (b) pricing responses from Western omni/realtime vendors to the $0.15/$0.47 row; (c) possible expansion of the omni family (a Plus/Max omni sibling and/or an open-weight omni Flash-Next split, mirroring the August pattern); (d) IaaS-bill post-mortems from the first production pilots. (PREDICTION.)
  • Structural watch (quarter): whether Alibaba turns the open-harness/closed-model shape into an ecosystem standard for Chinese-lab omni releases, and whether enterprise governance teams build an explicit "weights vs harnesses" evaluation axis. (INTERPRETATION + PREDICTION.)

★

Editorial takeaway

🎓 For Explorer

Qwen3.8-Omni-Flash is this window's price-setting omni release: 1M context, text/image/audio/video input, thinking, tools, and web search at $0.15/$0.47 international (¥0.8/0.1/2.7 mainland) — with the $0.016/M cache tier making long media loops genuinely cheap. The verified core is real (specs, prices, docs, repos, paper all checked), and the companion releases (Realtime API on Sep 21, Apache-2.0 Qwen-Live-Harness and Qwen-MM-Plugins, arXiv report Sep 22) make it the week's most complete Chinese-lab omni vertical. The honest caveats are equally real: every headline capability claim — the >25%-over-Omni-Plus average, "overall audio exceeds Gemini 3.8 Flash," even the 89/93/98% cost-reduction numbers — is vendor-run or methodology-dependent; the flat, verifiable per-token cut against the previous Flash omni is 62.5% input / 84.3% output, which is already a story without the hourly-media claims. And the model is API-only: the open parts are the harnesses, not the brain. The through-line for the weekly: Alibaba has turned omni from a premium per-modality product into a commodity Flash-tier row, set the reference price for realtime voice/video agents, and made "who runs the benchmark" the next fight — because with claims this big and no weights to inspect, independent replication is the only referee.

Illustration: frame: a bright flash passes through a ring of abstract sensor forms — geometric ear, lens and antenna shapes — an artistic impression of a real-time omni-modal model taking in many senses at once.
⌘

Lab: VERIFY

Step 1 — Verify the official spec and pricing from the provider's own pages (all fetched in full 2026-09-23)

From https://www.alibabacloud.com/help/en/model-studio/qwen3-8-omni-flash, https://www.qwencloud.com/models/qwen3.8-omni-flash, https://www.alibabacloud.com/help/en/model-studio/model-pricing, and https://www.alibabacloud.com/help/en/model-studio/realtime:

Claim in circulationOfficial pages sayVerdict
1M-token contextContext window 1M; max input 991,808 (non-thinking) / 983,616 (thinking); max output 131,072; max CoT 262,144VERIFIED
Omni (text/image/audio/video input)Input text/image/audio/video; output text only on the offline API; 113 audio languages; 2ch/4ch spatial audioVERIFIED — with the "text-only output" nuance most headlines missed
"$0.15 / $0.47"Singapore "International" scope $0.15 in / $0.016 implicit-cache / $0.47 out per 1M; other "Global" scopes $0.113 / $0.014 / $0.382; mainland CNY 0.8 / 0.1 / 2.7VERIFIED — with the region-tier correction (see Step 2)
Thinking on by defaultreasoning_effort xhigh (default) / medium / low; preserve_thinking default-on; none disablesVERIFIED
Tools + web search + cachingTool calling (custom), Responses web_search, automatic implicit cache + Session, structured outputsVERIFIED
Realtime companionqwen3.8-omni-flash-realtime: WebSocket/WebRTC, text+audio out, 120-min sessions, 196,608 max input, 100/50-turn caps, MCP, voice cloningVERIFIED (docs; released 2026-09-21)
Open weightsNot announced at launch (MarkTechPost + Oflight independently)VERIFIED — API-only
Step 2 — Reproduce the cost claims (pure arithmetic on published prices; script executed 2026-09-23)
# s11_cost.py — reproduce Qwen3.8-Omni-Flash cost claims from published per-token prices
old_in, old_out = 0.40, 3.00   # Qwen3.5-Omni-Flash (CloudPrice intl)
new_in, new_out = 0.15, 0.47   # Qwen3.8-Omni-Flash (intl tier)
print(f"input cut:  {1 - new_in/old_in:.1%}   output cut: {1 - new_out/old_out:.1%}")
print(f"implicit cache discount: {1 - 0.016/0.15:.1%}")
old_audio, new_audio = 4.57, 1.86   # Qwen3-Omni-Flash-Realtime vs qwen3.8-omni-flash-realtime audio input
print(f"realtime audio cut: {1 - new_audio/old_audio:.1%}")
# agent-loop cache scenario, 1M-token prefix, 10 reads
uncached = 1_000_000/1e6 * 0.15 * 10
cached   = 1_000_000/1e6 * (0.15 + 9*0.016)
print(f"10 uncached reads ${uncached:.2f} vs warm+9 cached ${cached:.2f} = {1-cached/uncached:.1%} cheaper")

Actual output:

input cut:  62.5%   output cut: 84.3%
implicit cache discount: 89.3%
realtime audio cut: 59.3%
10 uncached reads $1.50 vs warm+9 cached $0.29 = 80.4% cheaper

Findings — the story's most important numerics:

  • The verifiable per-token floor: 62.5% input / 84.3% output cheaper than Qwen3.5-Omni-Flash (pure ratio of two official price rows). This is the number that survives independent audit.
  • The headline "~89% video / >93% audio-visual / >98% audio cheaper" figures do NOT reproduce on per-token math — they are vendor-methodology claims that embed a per-hour token-consumption drop (agentic mode watching/listening selectively: 145,736 → 79,117 tokens in the blog's own demo, ~45.7%) on top of the per-token cut. The blog discloses its hourly methodology (hourly input price = 30 × the input cost of 2 minutes of source material; 720p at 1 fps); the X post's ~89%-video figure is a third, different framing. Correct quoting discipline: per-token 62.5%/84.3% as facts; 89/93/98% as labeled company claims.
  • Realtime audio: ~$1.86/M vs $4.57/M for the Dec-2025 Qwen3-Omni-Flash-Realtime → ~59% per-token cut (via the EmpirioLabs mirror of official rates) — again smaller than the >98% hourly claim.
  • The cache is the practical agent win: a 1M-token media-adjacent prefix read 10× costs $0.29 warm vs $1.50 uncached (80.4% cheaper; $0.016/M implicit-cache reads).
  • Regional variance is real: $0.15/$0.47 (Singapore "International") vs $0.113/$0.014/$0.382 (other "Global" scopes) vs CNY 0.8/0.1/2.7 (mainland). Quotes must pin the region; the discovery record's single "$0.15/$0.47" is the international headline rate only.
Step 3 — COMPARE: the previous generation and the current Flash tier (official/listed prices, 2026-09-23)
ModelContextInput typesOutputIn $/MCache $/MOut $/MNotes
Qwen3-Omni-Flash-Realtime (Dec 2025)262K-classaudio/video/textaudio+textaudio $4.57/M——8-dialog-turn sessions; tiny realtime contexts
Qwen3.5-Omni-Plus262Kaudio/video/texttext (audio on realtime)$1.15 (Beijing tier; heavier audio billing)—$8.75predecessor premium omni
Qwen3.5-Omni-Flash262Ktext/image/audio/videotext$0.40—$3.00previous Flash omni (CloudPrice intl)
Qwen3.8-Flash (Aug 2026)1Mtext/image/video, no audiotext$0.15—$0.47the tier-setter; not omni; had open-weight Flash-Next preview
Qwen3.8-Omni-Flash (Sep 18)1Mtext/image/audio/videotext only$0.15$0.016$0.47113 audio languages; 2ch/4ch; thinking default-on; closed weights
Qwen3.8-Omni-Flash-Realtime (Sep 21)196K input captext/image/audio/videoaudio+textaudio ~$1.86/M——120-min sessions; MCP; voice cloning; Live-Harness

Reading of the table (INTERPRETATION): the omni-flash closes the gap between "Flash tier" and "omni tier" — 1M context + audio/video input now cost the same as text/image Flash inference, and the realtime line resets the audio-input premium (~$1.86 vs $4.57). The strategic benchmark named by Qwen is Gemini 3.8 Flash (Qwen-run evals claim audio leadership; Neowin/El Economista relay it) — but that comparison is COMPANY CLAIM with no independent replication as of Sep 23 (research/S11.md Section 13).

Step 4 — Verify the open-tooling claims on the repos (both fetched in full 2026-09-23)
ClaimRepo evidenceVerdict
Qwen-Live-Harness Apache-2.0, v1.0.0License file Apache-2.0; News entry dates v1.0.0 2026-09-21VERIFIED
macOS 12+ only at v1.0.0Repo states macOS 12+; Windows/Linux "in progress"VERIFIED (constraint)
Delegates to Claude Code / Codex / Qwen Code / Gemini CLIACP backends listed in repo (Qwen Code, Qoder CLI, Codex, Claude Code, Gemini CLI)VERIFIED
Proactive monitors "not safety-critical"Repo's own disclaimerVERIFIED (scope limit)
Qwen-MM-Plugins Apache-2.0License file Apache-2.0; ~3k starsVERIFIED
Audio-native gap"most harnesses cannot yet feed audio to the main model natively — audio is handled through the API for now"VERIFIED (documents the ecosystem immaturity)
Step 5 — Verification checklist (claim-by-claim, with evidence labels)
ClaimVerdictEvidence
Released Sep 18, 2026 (event date)VERIFIEDQwen blog + Model Studio docs "Last Updated: Sep 18" + Gate News flash 07:03 UTC + MarkTechPost/TechNode; catalogues saying Sep 17 are the UTC-side of Beijing time — both inside the window
1M context; text/image/audio/video in; text outVERIFIEDOfficial docs (source 2)
Pricing $0.15/$0.016/$0.47 internationalVERIFIEDQwenCloud + Model Studio pricing page (sources 4–5); regional tiers differ
Thinking default-on, effort xhigh/medium/lowVERIFIEDOfficial docs
API-only, no open weightsVERIFIEDMarkTechPost + Oflight independently; no HF/repo release found
Realtime variant + Live-Harness v1.0.0 on Sep 21VERIFIEDOfficial Realtime docs + harness repo News entry
>25% avg across 29 evals (→ TechNode: >26%/30)COMPANY CLAIMQwen blog mirror; no independent replication; figure friction noted
"Audio exceeds Gemini 3.8 Flash"COMPANY CLAIMQwen-run cross-model evals; relayed by Neowin/El Economista; no third-party rerun
~89% video / >93% AV / >98% audio cheaperCOMPANY CLAIM (methodology-dependent)X post + blog (methodology disclosed); flat per-token floor is 62.5%/84.3% (arithmetic on official prices)
OmniVideoBench 63.4→67.8 at ~45.7% fewer tokens (145,736→79,117)COMPANY CLAIMBlog table; Gate News relays "51.8%" (friction noted); no independent rerun
MoE from Qwen3.8-Next, co-trainingCOMPANY CLAIM (declared)arXiv 2609.25611 abstract

What this lab does NOT do (honest limits)

  • It does not call either API (no DashScope key/budget) and does not run the model on real audio/video workloads.
  • It treats every vendor benchmark row, the Gemini comparison, and the 89/93/98% hourly cost claims as COMPANY CLAIM (evidence discipline). "Multimodal & Grounded 85.1" from modelscale.dev is a registry aggregate, not a deep evaluation.
  • The per-hour-cost methodology (30 × two-minute input; 720p@1fps) is Qwen's; real workloads have different token consumption and caching, so per-token floor vs hourly claim spread is expected to vary.
  • Per-token arithmetic uses the international tier; mainland and other "Global" scopes differ (CNY 0.8/0.1/2.7; $0.113/$0.014/$0.382).

Result

A one-page, evidence-labelled verification pack proving: (1) the official spec (1M context; omni input; text-only output; thinking default-on) and pricing ($0.15/$0.016/$0.47 international; region tiers differ) match the launch sheet; (2) the story-worthy arithmetic is the flat per-token cut — 62.5% input / 84.3% output vs Qwen3.5-Omni-Flash, 59.3% realtime-audio cut, 80.4% cheaper warm agent loops — while the headline 89/93/98% hourly figures are methodology-dependent company claims, never bare facts; (3) both open-tooling repos verify as Apache-2.0 with documented constraints (macOS-only harness; audio routed via API); and (4) the only independent measurement found (modelscale multimodal row 85.1) is a registry aggregate — the decisive open question remains third-party replication of the audio/video claims.

≡

Research sources

Primary Sources (10)
Primary
Qwen on X (@Alibaba_Qwen) — launch post (status 2100785962414702599) - **URL:** https://x.com/Alibaba_Qwen/status/2100785962414702599** Origin of the **~89% video-input cost reduction** figure vs Qwen3.5-Omni-Plus and the +19.5-pt average agentic gain (captured via search-index excerpts; the post underlies the PANews/Gate News relay of "~89%"). The canonical blog figures are >93% audio-visual / >98% audio per the disclosed methodology — the X-post number is a different (smaller) framing and is treated as COMPANY CLAIM. — ** COMPANY CLAIM (primary-source origin of the 89% figure; not independently derivable). ---Date: ** 2026-09-18
URL unavailable
Primary
GitHub — QwenLM/Qwen-MM-Plugins (fetched in full) - **URL:** https://github.com/QwenLM/Qwen-MM-Plugins** Apache-2.0 license verified; "Make any agent harness multimodal-native": plugins `core`, `api`, `search`, `mhs`, `video-memory`, `video-edit`, `video-spatio`, `blender`, `freecad`, `edu-agent`, `omni-chatcut` (Music-to-MV/commentary/video translation), `omni-video2note`, `omni-skill-creator`, `omni-memory`; guided installer for Claude Code, CodeBuddy, Codex, Qoder, OpenClaw, Qwen Code, Gemini CLI; **documented gap: "most harnesses cannot yet feed audio to the main model natively — audio is handled through the API for now"**. — ** FACT (license, capabilities, and the audio-routing caveat that limits the ecosystem-moat story).Date: ** consulted 2026-09-23 (~3k stars)
URL unavailable
Primary
GitHub — QwenLM/Qwen-Live-Harness (fetched in full) - **URL:** https://github.com/QwenLM/Qwen-Live-Harness** Apache-2.0 license verified; desktop harness for the Realtime API; macOS 12+ (Windows/Linux "in progress"); camera/mic interaction; background-task delegation (Qwen Code, Qoder CLI, Codex, Claude Code, Gemini CLI via ACP); proactive audio/visual monitors (1 fps, 2-second audio chunks; explicitly "not a safety-critical alarm system"); long-term memory with local-first storage; `npm install -g qwen-live-harness`; ~96 stars / young codebase (v1.0.0). — ** FACT (existence, license, macOS-only constraint, capability claims verified on the repo).Date: ** v1.0.0 released 2026-09-21 per the repo's own News entry; consulted 2026-09-23
URL unavailable
Primary
arXiv — "Qwen3.8-Omni: Towards Native Omni-Modal Agents" (Qwen Team), 2609.25611 (fetched in full) - **URL:** https://arxiv.org/abs/2609.25611v1** The official technical anchor: sparse MoE architecture inherited from Qwen3.8-Next; 1M context; native multimodal co-training strategy ("preserving strong text-domain capabilities while facilitating the transfer of agentic capabilities from text to audio and video tasks"); introduces Qwen-MM-Plugins and Qwen-Live-Harness. (Abstract reviewed in full; deeper body review is out of scope for this research.) — ** FACT (report exists, authors, submission date, declared architecture) + COMPANY CLAIM (co-training claims are the authors' own).Date: ** submitted 2026-09-22 (inside the window)
URL unavailable
Primary
Alibaba Cloud Model Studio — official Realtime API documentation (fetched in full) - **URL:** https://www.alibabacloud.com/help/en/model-studio/realtime** The companion `qwen3.8-omni-flash-realtime` artifact: WebSocket + WebRTC; text **and audio** output; up to 100 audio turns / 50 video turns; 600 s audio / 240 s video retained in context; **196,608 max input tokens**; 120-minute session cap; MCP tool support with approval; voice cloning; voice list incl. `longanlingxin`, default `Tina`; multichannel audio (1/2/4 ch); server-side VAD; workspace-scoped endpoints (`wss://{WorkspaceId}.…maas.aliyuncs.com/api-ws/v1/realtime`); web search and tool calling mutually exclusive on this endpoint; region-locked API keys. — ** FACT (realtime variant specs and constraints — basis for the "1M context ≠ realtime input cap" technical nuance).Date: ** page "Last Updated" Sep 23, 2026 (consulted 2026-09-23); model referenced throughout
URL unavailable
Primary
Alibaba Cloud Model Studio — official model pricing page - **URL:** https://www.alibabacloud.com/help/en/model-studio/model-pricing** Regional pricing tiers for the omni-flash SKU: Singapore "International" scope **$0.15 / $0.016 / $0.47**; other "Global" scopes (Beijing, Hong Kong, Frankfurt, Tokyo, Virginia) **$0.113 / $0.014 / $0.382**; mainland China published row **CNY 0.8 input / 0.1 cache-hit / 2.7 output** (per modelscale.dev's transcription of this page; Oflight confirms ¥0.8 vs ~¥18 predecessor). Together with source 4 it resolves the "$0.15/$0.47 is international, not a single global price" correction. — ** FACT (official pricing table — the discovery record's flat "$0.15/$0.47" is refined to "international/Singapore tier" by this page).Date: ** consulted 2026-09-23 (live page)
URL unavailable
Primary
QwenCloud — official model page "qwen3.8-omni-flash" (fetched in full) - **URL:** https://www.qwencloud.com/models/qwen3.8-omni-flash** Official international pricing: **Input $0.15 / 1M, Implicit Cache $0.016 / 1M, Output $0.47 / 1M**; 1M context; max reasoning 262K; TPM 2M; RPM 30K; built on Qwen3.8-Flash-Next architecture; 2ch/4ch spatial audio; DashScope + OpenAI protocols. This is the source of the story's headline "$0.15/$0.47" (also mirrored by LiteLLM/Vercel/CloudPrice) — ** FACT (official price list — anchor for all cost arithmetic).Date: ** listing live when consulted 2026-09-23 (release Sep 18)
URL unavailable
Primary
Alibaba Cloud Community — full mirror of the Qwen blog announcement (fetched in full) - **URL:** https://www.alibabacloud.com/blog/603580** Complete readable text of the official announcement (the only full-text rendering available, since qwen.ai is a JS SPA): the >25%-across-29-evals claim; WildClawBench-MM +36.5 / AgenticVBench +22.3 / UniClawBench 69.6 / LongAudioSpan +8.3 / OmniVideoBench +9.6 / OmniCap-IF CSR +8.5 / ISR +14.1 / AliMeeting DER 88.11→3.35 cpWER 89.61→17.18; "audio-visual performance close to Gemini 3.8 Flash and overall audio performance that exceeds Gemini 3.8 Flash" (Qwen-run cross-model evals); OmniVideoBench agentic-vs-static 63.4 → 67.8 with tokens 145,736 → 79,117 (**~45.7% fewer**); pricing methodology footnote; the "up to one hour" meeting scenario; multichannel (2-ch/4-ch FOA) input; Qwen2.5-Omni-3B Sichuan-dialect self-improvement demo (CER 25.79% → 15.30% via 3,413 examples / 4 rounds in 12 hours). — ** FACT (it is the official text, mirrored) + COMPANY CLAIM (all benchmark/demo figures are vendor-published).Date: ** 2026-09-20 (community blog; mirrors the Sep 18 Qwen blog)
URL unavailable
Primary
Alibaba Cloud Model Studio — official model documentation: "Qwen3.8-Omni-Flash" (fetched in full) - **URL:** https://www.alibabacloud.com/help/en/model-studio/qwen3-8-omni-flash** Canonical spec sheet: input types text/image/audio/video, **output text only** (non-realtime API); **1M-token context window**; max input **991,808 tokens (non-thinking) / 983,616 (thinking)**; max output **131,072**; max reasoning (CoT) **262,144**; thinking enabled by default with `reasoning_effort` xhigh/medium/low, `preserve_thinking` default-on, `none` disables thinking; tool calling (custom), built-in web search (Responses `web_search`), automatic implicit context caching + Responses Session, structured outputs, batch; DashScope + OpenAI-compatible protocols; regions Beijing, Singapore, Hong Kong, Tokyo, Frankfurt, US Virginia. — ** FACT (authoritative spec — resolves the "omni" question: omni *input*, text-only *output* on the offline API).Date: ** "Last Updated: Sep 18, 2026" (page consulted 2026-09-23)
URL unavailable
Primary
Qwen (official blog) — "Qwen3.8-Omni-Flash: Omni Senses. Agentic Delivery." - **URL:** https://qwen.ai/blog?id=qwen3.8-omni-flash** Official launch narrative: "next-generation native omnimodal model," moving omni models "from understanding omnimodal content to planning tasks, calling tools, and completing creative work"; availability on Qianwen AI Platform, QwenCloud, Qwen Studio, Alibaba Cloud Model Studio; architecture lineage (Qwen3.8-Flash-Next); 1M context; thinking support; the agentic-perception framing (coarse-to-fine evidence gathering); the audio/audio-visual per-hour cost claims (>93% audio-visual / >98% audio, with disclosed methodology: hourly price = 30 × the input cost of two minutes of source material; 720p at 1 fps); performance claims (>25% average across 29 evaluations vs Qwen3.5-Omni-Plus). — ** FACT baseline (release date, model name, availability, scope) + COMPANY CLAIM (all performance/cost-reduction figures). Note: the page is a JS-rendered SPA — direct webfetch returns only the page shell ("Qwen"); the full text was captured via search-index excerpts and cross-checked against the full mirror on Alibaba Cloud Community (source 3).Date: ** 2026-09-18 (blog dated 2026/09/18)
URL unavailable
Independent Sources (6)
Independent
EmpirioLabs — "qwen3.8-omni-flash-realtime" model page - **URL:** https://empiriolabs.ai/models/qwen3-8-omni-flash-realtime** Independent mirror of the realtime model's official rates: audio input ~**$1.86/M** audio tokens (vs Qwen3-Omni-Flash-Realtime's $4.57/M on QwenCloud — a ~59% per-token cut); release-date drift evidence (their "Sep 17" vs the official Sep 18 — UTC/Beijing timezone documentable difference). — ** Independent directory — source of the realtime audio per-token comparison and a second instance of the dating drift. ---Date: ** listing added 2026-09-21 (release dated "Sep 17, 2026" in their field — UTC-side dating note)
URL unavailable
Independent
Oflight — "Qwen3.8-Omni-Flash API: 1M context…" (Japanese vendor analysis) - **URL:** https://www.oflight.co.jp/en/columns/qwen3-8-omni-flash-api-1m-context-2026** Independent technical read: **"weights are closed and proprietary"** (confirms API-only, second independent source); mainland Bailian multimodal input ¥0.8/M ("down sharply from roughly ¥18 for the predecessor"); one-token-rate-for-every-modality pricing note; text-only output and the chain-a-TTS-stage pipeline guidance; the realtime line differentiation. — ** Independent technical reporting — corroborates API-only status, mainland pricing backdrop, and the text-output constraint.Date: ** 2026-09-19 (article)
URL unavailable
Independent
Pandaily — "Qwen3.8-Omni-Flash: …native omni-modal model" (China-tech beat) - **URL:** https://pandaily.com/qwen3-8-omni-flash-native-omni-model** Independent English-language China-tech recap of the launch (specs, positioning, pricing direction). Page is JS-rendered — direct webfetch returns title only; content corroborated via search-index excerpts. — ** Independent reporting (relayed content) — corroborates the launch facts from the China-tech beat.Date: ** 2026-09-19 (publish-date metadata)
URL unavailable
Independent
Gate News (flash) — "Alibaba's Qwen Launches Qwen3.8-Omni-Flash Model to Enhance AI Agent" - **URL:** https://www.gate.com/news/detail/alibabas-qwen-launches-qwen38-omni-flash-model-to-enhance-ai-agent-24386332** Second relay: **51.8% fewer tokens** in agentic perception mode (relaying PANews) vs the blog's 45.7% (145,736 → 79,117) — one of the documented figure frictions; audio/audio-visual cost claims. Direct webfetch returns HTTP 403 — content captured via full search-index excerpts. — ** Independent relay — documents the small official-vs-relay figure variance (45.7% vs 51.8%) that this research flags for quote-time re-pinning.Date: ** 2026-09-18 (flash)
URL unavailable
Independent
Gate News (flash) — "Alibaba Qwen Releases Qwen3.8-Omni-Flash With 89% Reduction in Video Input" - **URL:** https://www.gate.com/news/detail/alibaba-qwen-releases-qwen38-omni-flash-with-89-reduction-in-video-input-24368598** Time-stamped launch relay (2026-09-18 07:03 UTC): 1M context; the ~89% video-input cost reduction (relaying the Qwen X post/PANews); agentic perception framing; Qwen3.8-Flash-Next lineage. Direct webfetch returns HTTP 403 (bot protection) — content captured via full search-index excerpts. — ** Independent relay (Chinese crypto-media aggregation of PANews) — corroborates launch date/framing; the 89% figure is a relay of a company claim, labeled as such.Date: ** 2026-09-18 (flash; relayed PANews)
URL unavailable
Independent
MarkTechPost — "Alibaba Qwen Releases Qwen3.8-Omni-Flash" (Asif Razzaq; fetched in full) - **URL:** https://www.marktechpost.com/2026/09/18/alibaba-qwen-releases-qwen3-8-omni-flash/** Independent recap: **"No open weights were announced at launch, so self-hosting is not an option"** (key API-only confirmation, echoed by Oflight); text-only output; thinking default-on with `reasoning_effort`=xhigh; DashScope + OpenAI protocols; video via URL up to ~2 h / 2 GB, audio up to ~3 h (docs reading); 113 audio languages; 6 regions; explicit note "All figures here come from Qwen. Independent results were not available at publication"; Qwen-MM-Plugins Apache-2.0 skills; the X post ~89% claim. — ** Independent reporting — the principal independent confirmation of the API-only status and the vendor-figure caveat.Date: ** 2026-09-18
URL unavailable
Secondary Sources (4)
Secondary
Relayed launch-week press (Neowin, El Economista, DataCamp, TechNode, Ground News aggregation) - **No URL available** (all consulted through search-index excerpts / the Ground News aggregation page surfaced during web research; individual article URLs not captured in this environment)** Neowin: "undercuts Gemini on audio" framing; El Economista: "surpassing Google's Gemini 3.8 Flash in voice processing" (both relay the Qwen-run Gemini comparison — COMPANY CLAIM in circulation); TechNode: rounds the official average claim to "more than 26% across 30 evaluations" (vs the blog's >25% / 29 — documented friction) and confirms the Gemini 3.8 Flash comparison framing; DataCamp: standard spec recap (1M context, text/image/audio/video in, $0.15/$0.47); Ground News: coverage-map corroboration. — ** Secondary press relays — establish the story's media spread and the minor official-vs-press figure frictions; none add independent benchmark data. ---Date: ** 2026-09-18 (Neowin, El Economista, DataCamp, TechNode; Ground News aggregation of 5 articles incl. TechNode, El Economista, Neowin, MarkTechPost)
URL unavailable
Secondary
modelscale.dev — "Qwen3.8 Omni Flash" capability registry - **No URL available** (consulted via search-index excerpts of modelscale.dev during web research; exact page URL not captured)** Observed benchmark-aligned capability rows for the hosted API: Coding 54.9 / Knowledge 53.9 / **Multimodal & Grounded 85.1** / Instruction Following 87.7 (BenchLM-sourced) — the *only* partial independent measurement of the model found as of Sep 23 (registry aggregate, not a deep evaluation of the omni claims); also transcribes the official mainland pricing row (CNY 0.8/0.1/2.7) and lists the release date Sep 17 (dating drift). — ** Secondary registry — the only third-party capability readout found; explicitly not an AA-style leaderboard.Date: ** 2026-09-22 (page date)
URL unavailable
Secondary
LiteLLM — day-0 provider support for `dashscope/qwen3.8-omni-flash` - **No URL available** (found via search-index excerpts referencing LiteLLM PR #41754 and the DashScope provider docs; not fetched directly in this environment)** Independent ecosystem corroboration that the model was routable through a standard proxy on launch week (OpenAI-compatible Chat Completions + Responses via DashScope), i.e. not a walled-garden API despite being API-only. — ** Secondary (ecosystem) — corroborates API compatibility claims in Section 9 of research/S11.md.Date: ** 2026-09-18 to 2026-09-21 (day-0 support window)
URL unavailable
Secondary
CloudPrice — Qwen3.5-Omni-Flash pricing & specs (predecessor cross-check) - **No URL available** (consulted via search-index excerpts of cloudprice.net during web research; the exact page URL was not captured in this environment)** The previous-generation *Flash* omni international pricing used for the flat per-token comparison: **$0.40 input / $3.00 output** per 1M → the verifiable floor cost cut for Qwen3.8-Omni-Flash is **62.5% input / 84.3% output** (pure arithmetic on the two official price rows, sources 4–5). Also listed release date 2026-09-17 for the new model (dating-drift instance). — ** Secondary directory — pricing cross-check for the predecessor and an instance of the Sep 17 dating drift.Date: ** data current Sep 2026 (predecessor listing `qwen3.5-omni-flash`)
URL unavailable