A cheap model that sees and hears
Alibaba's Qwen team launched Qwen3.8-Omni-Flash on Sep 18, 2026. It accepts four kinds of input and returns written answers. The twist is its price, which sits at the level of plain text models.
Alibaba released Qwen3.8-Omni-Flash, a model that reads text, images, audio and video at one very low price. It ships with free tools you can install today.

Alibaba's Qwen team launched Qwen3.8-Omni-Flash on Sep 18, 2026. It accepts four kinds of input and returns written answers. The twist is its price, which sits at the level of plain text models.
| Field | Value |
|---|---|
| Story ID | S11 |
| Title | Alibaba Cloud releases Qwen3.8-Omni-Flash: 1M-context real-time omni model at $0.15/$0.47 |
| Organization | Alibaba Cloud / Qwen (Qwen Team; announced via the Qwen official blog and Qwen's X account @Alibaba_Qwen) |
| Category | Model Release |
| Event date | 2026-09-18 (Qwen blog "Qwen3.8-Omni-Flash: Omni Senses. Agentic Delivery." dated 2026/09/18; Alibaba Cloud Model Studio docs page "Last Updated: Sep 18, 2026"; Gate News flash 2026-09-18 07:03 UTC; TechNode and MarkTechPost both dated 2026-09-18) |
| Announcement date | 2026-09-18 (same as event date; companion releases inside the window: Qwen3.8-Omni-Flash-Realtime + Qwen-Live-Harness v1.0.0 on 2026-09-21 per the harness repo's own News entry; arXiv technical report 2609.25611 submitted 2026-09-22) |
| Article dates | 2026-09-18 (MarkTechPost, TechNode, Neowin, El Economista, DataCamp, Gate News ×2), 2026-09-19 (Pandaily publish-date metadata, Oflight), 2026-09-20 (Alibaba Cloud Community blog mirror), 2026-09-21 (liteLLM, QwenCloud listing, EmpirioLabs realtime page, sofar/bot), 2026-09-22 (modelscale.dev, Compras…) |
| Alternative dating | Several registry/directory sources list the release as 2026-09-17 (CloudPrice "Released 2026-09-17"; EmpirioLabs "Released Sep 17, 2026"; OpenCode model stats "Sep 17, 2026"; Alibaba docs' Traditional-Chinese help page shows "更新時間 Sep 17"). This is the UTC-side of the Beijing-time (UTC+8) announcement: Sep 18 before ~08:00 Beijing time is still Sep 17 UTC. Both dates are inside the window, so eligibility is unaffected; the story's consensus event date is 2026-09-18 |
| Window check | 2026-09-18 ∈ [2026-09-18, 2026-09-22] — eligible (the 2026-09-21 Realtime/Harness release and the 2026-09-22 arXiv report are also inside the window) |
| Evidence status | CONFIRMED (official Qwen blog announcement text captured via full mirror on Alibaba Cloud Community + search-index excerpts of qwen.ai; official Model Studio spec page fetched in full; official Realtime API docs fetched in full; official QwenCloud pricing page fetched in full; arXiv abstract fetched in full; both GitHub repos fetched in full; four independent English-language coverages fetched/captured; three additional independent relays; two secondary price directories) |
Corrections / refinements to the discovery record (important):
qwen3.8-omni-flash (Sep 18) is a non-realtime HTTP API model (Chat Completions/Responses) with text-only output. The realtime variant qwen3.8-omni-flash-realtime (WebSocket/WebRTC, text and audio output, voice cloning, MCP, multichannel input) was released with Qwen-Live-Harness v1.0.0 on 2026-09-21 per the harness repo's own News entry. The story spans both; this research records both dates, both inside the window.On September 18, 2026, the Qwen Team (Alibaba) launched Qwen3.8-Omni-Flash, described in the official blog as "our next-generation native omnimodal model" built to move omni models "from understanding omnimodal content to planning tasks, calling tools, and completing creative work" (FACT, official blog). It went live the same day on the Qianwen AI Platform, QwenCloud, Qwen Studio, and Alibaba Cloud Model Studio, with an OpenAI-compatible Chat Completions + Responses API (FACT; multiple sources).
What shipped (FACT, official docs + QwenCloud + blog):
use_multichannel.reasoning_effort = xhigh/medium/low and preserve_thinking default-on; reasoning_effort: none disables thinking.web_search), context caching (automatic implicit + Responses Session), structured outputs, batch calls, DashScope + OpenAI protocols.Companion releases inside the window (FACT):
longanlingxin, default Tina) — official Realtime docs, page confirmed live Sep 23 with the model referenced throughout.npm install -g qwen-live-harness.core, api, search, omni-chatcut (Music-to-MV/commentary/video translation), omni-video2note, omni-skill-creator, omni-memory, video-edit, blender, freecad. Documented gap: "most harnesses cannot yet feed audio to the main model natively — audio is handled through the API for now."Company-published performance claims (COMPANY CLAIM — MarkTechPost explicitly notes "All figures here come from Qwen. Independent results were not available at publication"):
Independent measurements (INDEPENDENT EVIDENCE, partial): modelscale.dev (2026-09-22) shows observed benchmark-aligned capability rows for the hosted API — Coding 54.9, Knowledge 53.9, Multimodal & Grounded 85.1, Instruction Following 87.7 (BenchLM-sourced; no composite score). No Artificial Analysis-style leaderboard entry for the omni-flash was found as of Sep 23; no independent replication of the audio/video benchmark claims exists yet.
Before (through Sep 17, 2026):
Change (Sep 18–22, 2026):
After:
reasoning_effort = xhigh (default) / medium / low; preserve_thinking default-on; none disables thinking. Reasoning tokens count against the 262K max chain-of-thought and are billed as output tokens.raw_mic_array/foa_ambix); video aggregation flag representation_compact; MCP Streamable-HTTP tools with approval; voice list incl. longanlingxin + voice cloning.daemon (models, tools, memory, delegation) + Electron Host (UI, capture, playback) over a local WebSocket; background harnesses (Qwen Code, Qoder, Codex, Claude Code, Gemini CLI) over ACP/REST/SSE; Proactive monitors (1 fps, 2-second audio chunks — explicitly "not a safety-critical alarm system"); local memory libraries with cloud consolidation.core lets the main model read local images/video natively; omni capabilities route media understanding through the DashScope API (harnesses can't feed audio to the main model natively yet).npm install -g qwen-live-harness on macOS 12+ gives camera/mic conversational control of Claude Code/Codex/Qwen Code/Gemini CLI background agents, proactive monitor triggers, and persistent memory — no local GPU needed (cloud inference).wss://{WorkspaceId}.…maas.aliyuncs.com/api-ws/v1/realtime). Long-horizon realtime designs must engineer around these.qwen3.8-omni-flash to routing matrices this week: 1M context, text/image/audio/video in, text out, $0.15/$0.016/$0.47 (international), OpenAI-compatible Chat Completions + Responses, DashScope SDK; day-0 proxy support via LiteLLM (dashscope/qwen3.8-omni-flash, PR #41754) and Vercel AI Gateway (alibaba/qwen3.8-omni-flash).qwen3.8-omni-flash-realtime (audio out) for live interaction or bolt on a TTS stage for the offline API (Oflight's pipeline guidance).reasoning_effort (xhigh/medium/low) and preserve_thinking — a direct cost/depth dial; thinking tokens are billed as output and count against the 262K CoT cap.core + api + an omni capability; note the audio-native gap ("audio is handled through the API for now") before promising audio workflows inside your harness.representation_compact must be set before the first audio segment; voice changes to the 3.8 voice list (e.g. longanlingxin) don't carry over from earlier voice sets.Impact: An individual developer gets a 1M-context omni API that ingests audio/video at $0.15/$0.47 — enough to prototype voice/video agents and media pipelines this week — plus Apache-2.0 tooling (Live-Harness on macOS, MM-Plugins in existing harnesses) that makes the stack testable without waiting for weights that may never come.
Action: Run the labs/S11.md VERIFY+COMPARE exercise (read-only): confirm specs/prices from the official pages, rebuild the cost arithmetic (62.5%/84.3% per-token cuts vs Qwen3.5-Omni-Flash; realtime audio $4.57 → ~$1.86; methodology caveats on 89/93/98%), and inspect both GitHub repos' licenses and feature claims. Then, if a DashScope key + budget exist, run a $5–15 pilot: transcribe/summarize a 1-hour meeting via the offline API and run one 10-minute realtime conversation via qwen3.8-omni-flash-realtime, logging tokens and first-audio latency. The decision to build on the tooling is safe (Apache-2.0); the decision to build product on the closed API should wait for one billing cycle of real per-hour costs.
Impact: Teams/companies gain a new price floor for omni agent workloads and a realtime voice/video stack anchored by a Chinese vendor — with the governance caveats that audio/video data flows to Alibaba Cloud, no weights mean no self-host, and all capability claims are vendor-run so far.
Action: Add Qwen3.8-Omni-Flash to the evaluation matrix with three columns: (1) flat per-token cost ($0.15/$0.016/$0.47 international; ¥0.8-tier mainland) for cost modeling, (2) vendor-claimed hourly-media metrics (>98%/93%/89%) clearly labeled as methodology-dependent, and (3) region/data-flow posture for the specific workloads (meeting minutes, customer-service realtime, media localization). Run one non-production pilot (meeting summarization or video2note) before any volume commitment; negotiate on the basis of the per-token floor; document that audio output requires the realtime endpoint and that self-hosting is not available.
Impact: Alibaba's launch completes this window's Chinese price-setting pattern (MiMo-V2.6-Pro, GLM-5.3, Step 5 Preview, now Qwen3.8-Omni-Flash) and extends the "frontier-for-less" war into the omni/realtime modality — where Gemini Flash Live-class Western offers previously set the tone. It also introduces the closed-model + open-harness shape to the open-AI debate.
Action: Track three things over the next month: (1) independent omni benchmarks vs Gemini 3.8 Flash (the first third-party audio-visual evals will settle the flagship claim); (2) whether an open-weight omni checkpoint appears (Flash-Next-style split) and under which license; (3) real-world per-hour billing reports from pilot users, which will validate or puncture the 89/93/98% hourly claims. In procurement guidance, add "open harness vs open weights" as a distinct axis from "open source." (PREDICTION: expect Western omni pricing responses within 2–4 weeks of Qwen's $0.15/$0.47 row.)
VERIFY (with a COMPARE component) — see labs/S11.md. Without an API key or GPU, the highest-value hands-on work is verifying the launch's load-bearing claims against the public record: (1) confirm the 1M-context spec and $0.15/$0.016/$0.47 pricing from the official pages (done in this research — QwenCloud + Model Studio docs), (2) reproduce the cost arithmetic — 62.5%/84.3% per-token cuts vs Qwen3.5-Omni-Flash, realtime audio $4.57 → ~$1.86, and the methodology caveats behind the 89/93/98% hourly claims, (3) inspect both GitHub repos for license/feature claims (done — Apache-2.0; harness macOS-only; MM-Plugins audio gap documented), and (4) map which catalogues still date the release Sep 17 (timezone drift awareness). A paid TEST pilot on both the offline and realtime APIs is the natural follow-up.
Qwen3.8-Omni-Flash is this window's price-setting omni release: 1M context, text/image/audio/video input, thinking, tools, and web search at $0.15/$0.47 international (¥0.8/0.1/2.7 mainland) — with the $0.016/M cache tier making long media loops genuinely cheap. The verified core is real (specs, prices, docs, repos, paper all checked), and the companion releases (Realtime API on Sep 21, Apache-2.0 Qwen-Live-Harness and Qwen-MM-Plugins, arXiv report Sep 22) make it the week's most complete Chinese-lab omni vertical. The honest caveats are equally real: every headline capability claim — the >25%-over-Omni-Plus average, "overall audio exceeds Gemini 3.8 Flash," even the 89/93/98% cost-reduction numbers — is vendor-run or methodology-dependent; the flat, verifiable per-token cut against the previous Flash omni is 62.5% input / 84.3% output, which is already a story without the hourly-media claims. And the model is API-only: the open parts are the harnesses, not the brain. The through-line for the weekly: Alibaba has turned omni from a premium per-modality product into a commodity Flash-tier row, set the reference price for realtime voice/video agents, and made "who runs the benchmark" the next fight — because with claims this big and no weights to inspect, independent replication is the only referee.

From https://www.alibabacloud.com/help/en/model-studio/qwen3-8-omni-flash, https://www.qwencloud.com/models/qwen3.8-omni-flash, https://www.alibabacloud.com/help/en/model-studio/model-pricing, and https://www.alibabacloud.com/help/en/model-studio/realtime:
| Claim in circulation | Official pages say | Verdict |
|---|---|---|
| 1M-token context | Context window 1M; max input 991,808 (non-thinking) / 983,616 (thinking); max output 131,072; max CoT 262,144 | VERIFIED |
| Omni (text/image/audio/video input) | Input text/image/audio/video; output text only on the offline API; 113 audio languages; 2ch/4ch spatial audio | VERIFIED — with the "text-only output" nuance most headlines missed |
| "$0.15 / $0.47" | Singapore "International" scope $0.15 in / $0.016 implicit-cache / $0.47 out per 1M; other "Global" scopes $0.113 / $0.014 / $0.382; mainland CNY 0.8 / 0.1 / 2.7 | VERIFIED — with the region-tier correction (see Step 2) |
| Thinking on by default | reasoning_effort xhigh (default) / medium / low; preserve_thinking default-on; none disables | VERIFIED |
| Tools + web search + caching | Tool calling (custom), Responses web_search, automatic implicit cache + Session, structured outputs | VERIFIED |
| Realtime companion | qwen3.8-omni-flash-realtime: WebSocket/WebRTC, text+audio out, 120-min sessions, 196,608 max input, 100/50-turn caps, MCP, voice cloning | VERIFIED (docs; released 2026-09-21) |
| Open weights | Not announced at launch (MarkTechPost + Oflight independently) | VERIFIED — API-only |
# s11_cost.py — reproduce Qwen3.8-Omni-Flash cost claims from published per-token prices
old_in, old_out = 0.40, 3.00 # Qwen3.5-Omni-Flash (CloudPrice intl)
new_in, new_out = 0.15, 0.47 # Qwen3.8-Omni-Flash (intl tier)
print(f"input cut: {1 - new_in/old_in:.1%} output cut: {1 - new_out/old_out:.1%}")
print(f"implicit cache discount: {1 - 0.016/0.15:.1%}")
old_audio, new_audio = 4.57, 1.86 # Qwen3-Omni-Flash-Realtime vs qwen3.8-omni-flash-realtime audio input
print(f"realtime audio cut: {1 - new_audio/old_audio:.1%}")
# agent-loop cache scenario, 1M-token prefix, 10 reads
uncached = 1_000_000/1e6 * 0.15 * 10
cached = 1_000_000/1e6 * (0.15 + 9*0.016)
print(f"10 uncached reads ${uncached:.2f} vs warm+9 cached ${cached:.2f} = {1-cached/uncached:.1%} cheaper")
Actual output:
input cut: 62.5% output cut: 84.3%
implicit cache discount: 89.3%
realtime audio cut: 59.3%
10 uncached reads $1.50 vs warm+9 cached $0.29 = 80.4% cheaper
Findings — the story's most important numerics:
| Model | Context | Input types | Output | In $/M | Cache $/M | Out $/M | Notes |
|---|---|---|---|---|---|---|---|
| Qwen3-Omni-Flash-Realtime (Dec 2025) | 262K-class | audio/video/text | audio+text | audio $4.57/M | — | — | 8-dialog-turn sessions; tiny realtime contexts |
| Qwen3.5-Omni-Plus | 262K | audio/video/text | text (audio on realtime) | $1.15 (Beijing tier; heavier audio billing) | — | $8.75 | predecessor premium omni |
| Qwen3.5-Omni-Flash | 262K | text/image/audio/video | text | $0.40 | — | $3.00 | previous Flash omni (CloudPrice intl) |
| Qwen3.8-Flash (Aug 2026) | 1M | text/image/video, no audio | text | $0.15 | — | $0.47 | the tier-setter; not omni; had open-weight Flash-Next preview |
| Qwen3.8-Omni-Flash (Sep 18) | 1M | text/image/audio/video | text only | $0.15 | $0.016 | $0.47 | 113 audio languages; 2ch/4ch; thinking default-on; closed weights |
| Qwen3.8-Omni-Flash-Realtime (Sep 21) | 196K input cap | text/image/audio/video | audio+text | audio ~$1.86/M | — | — | 120-min sessions; MCP; voice cloning; Live-Harness |
Reading of the table (INTERPRETATION): the omni-flash closes the gap between "Flash tier" and "omni tier" — 1M context + audio/video input now cost the same as text/image Flash inference, and the realtime line resets the audio-input premium (~$1.86 vs $4.57). The strategic benchmark named by Qwen is Gemini 3.8 Flash (Qwen-run evals claim audio leadership; Neowin/El Economista relay it) — but that comparison is COMPANY CLAIM with no independent replication as of Sep 23 (research/S11.md Section 13).
| Claim | Repo evidence | Verdict |
|---|---|---|
| Qwen-Live-Harness Apache-2.0, v1.0.0 | License file Apache-2.0; News entry dates v1.0.0 2026-09-21 | VERIFIED |
| macOS 12+ only at v1.0.0 | Repo states macOS 12+; Windows/Linux "in progress" | VERIFIED (constraint) |
| Delegates to Claude Code / Codex / Qwen Code / Gemini CLI | ACP backends listed in repo (Qwen Code, Qoder CLI, Codex, Claude Code, Gemini CLI) | VERIFIED |
| Proactive monitors "not safety-critical" | Repo's own disclaimer | VERIFIED (scope limit) |
| Qwen-MM-Plugins Apache-2.0 | License file Apache-2.0; ~3k stars | VERIFIED |
| Audio-native gap | "most harnesses cannot yet feed audio to the main model natively — audio is handled through the API for now" | VERIFIED (documents the ecosystem immaturity) |
| Claim | Verdict | Evidence |
|---|---|---|
| Released Sep 18, 2026 (event date) | VERIFIED | Qwen blog + Model Studio docs "Last Updated: Sep 18" + Gate News flash 07:03 UTC + MarkTechPost/TechNode; catalogues saying Sep 17 are the UTC-side of Beijing time — both inside the window |
| 1M context; text/image/audio/video in; text out | VERIFIED | Official docs (source 2) |
| Pricing $0.15/$0.016/$0.47 international | VERIFIED | QwenCloud + Model Studio pricing page (sources 4–5); regional tiers differ |
| Thinking default-on, effort xhigh/medium/low | VERIFIED | Official docs |
| API-only, no open weights | VERIFIED | MarkTechPost + Oflight independently; no HF/repo release found |
| Realtime variant + Live-Harness v1.0.0 on Sep 21 | VERIFIED | Official Realtime docs + harness repo News entry |
| >25% avg across 29 evals (→ TechNode: >26%/30) | COMPANY CLAIM | Qwen blog mirror; no independent replication; figure friction noted |
| "Audio exceeds Gemini 3.8 Flash" | COMPANY CLAIM | Qwen-run cross-model evals; relayed by Neowin/El Economista; no third-party rerun |
| ~89% video / >93% AV / >98% audio cheaper | COMPANY CLAIM (methodology-dependent) | X post + blog (methodology disclosed); flat per-token floor is 62.5%/84.3% (arithmetic on official prices) |
| OmniVideoBench 63.4→67.8 at ~45.7% fewer tokens (145,736→79,117) | COMPANY CLAIM | Blog table; Gate News relays "51.8%" (friction noted); no independent rerun |
| MoE from Qwen3.8-Next, co-training | COMPANY CLAIM (declared) | arXiv 2609.25611 abstract |
A one-page, evidence-labelled verification pack proving: (1) the official spec (1M context; omni input; text-only output; thinking default-on) and pricing ($0.15/$0.016/$0.47 international; region tiers differ) match the launch sheet; (2) the story-worthy arithmetic is the flat per-token cut — 62.5% input / 84.3% output vs Qwen3.5-Omni-Flash, 59.3% realtime-audio cut, 80.4% cheaper warm agent loops — while the headline 89/93/98% hourly figures are methodology-dependent company claims, never bare facts; (3) both open-tooling repos verify as Apache-2.0 with documented constraints (macOS-only harness; audio routed via API); and (4) the only independent measurement found (modelscale multimodal row 85.1) is a registry aggregate — the decisive open question remains third-party replication of the audio/video claims.