News Weekly
LV 10 XP
0% read
Your progress · 0/5 chapters
About 6 min total
ModelsISSUE #1 · STORY 9 OF 53Sep 15, 2026CONFIRMED

Google's new voice AI can talk, see and get work done at the same time

Gemini 3.8 Live and its Extended Thinking sibling keep the conversation flowing while they look things up, speak 97 languages and cost a fraction of the competition.

An unbroken luminous ribbon of sound sweeps across the frame while a pale plane of image light folds into it and is carried onward.

Read it your way

CHAPTER 1 · THE 60-SECOND VERSIONPicked for Explorers

Two new voices from Google

On September 15, Google launched two real-time voice models. Gemini 3.8 Live is built to be fast and cheap at scale. Gemini 3.8 Live Extended Thinking is built for harder jobs: it keeps reasoning in the background while it goes on talking to you.

Talks in 97 languagesIt can switch languages mid-conversation without missing a beat.
Sees what you seePoint your phone camera at something and chat about it in near real time.
Works while it talksIt runs tools and looks things up in the background instead of going silent.
Cheap to runAbout $1.38 for an hour of back-and-forth conversation at list price.
Finish this chapter for +15 XP
Flip the switch

Voice assistants, before and after

YOU ASKOne model hears youListening, thinking and speaking happen inside a single model.
IT WORKSTools run in the backgroundIt keeps talking and tells you what it is doing while tasks finish.
YOU GETA conversation that flowsFaster turns, camera awareness and 97 languages in one session.
PLAY WITH THE NUMBERS · +10 XP

What would a voice agent cost you?

DRAG THE SLIDER
Gemini 3.8 Live
$2,760
about $1.38 per hour
OpenAI GPT-Live-1
$6,000
from $3.00 per hour
You'd keep
$3,240
every month

List prices at launch: $0.005/min in, $0.018/min out for Gemini; $0.05/min for GPT-Live-1's front end. Extra reasoning and video input cost more.

Your next move · as a Explorer

See it for yourself in two minutes.

1Open the Google app and tap the Live icon to start Search Live.
2Point your camera at something around you and ask about it out loud.
3Ask a question that needs a lookup and listen to how it talks you through the wait.

Switch your reading mode at the top to see a different next move.

Tap to open

Things to keep an eye on

Pop quiz · unlock the Voice Wrangler badge

Did it stick?

0/3
What makes Extended Thinking special?+20 XP
About how much is an hour of chat at list price?+20 XP
How many languages can it switch between?+20 XP
Your call · +5 XP

Will OpenAI cut its voice prices before the year is out?

Deep dive

The full research, labeled and sourced

CONFIRMED24 sources · 55 min
Story identity

CONFIRMED. Google (Google DeepMind) introduced two new native speech-to-speech live dialogue models — Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking — on September 15, 2026, via the official Google blog ("Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking", by Tom Ouyang (Principal Engineer) and Malini Jaganathan (Member of Technical Staff) for the Gemini Audio Team), a companion developer post, updated Gemini API documentation, and a DeepMind model card. Rollout began the same day across the Gemini API, Google AI Studio, Search Live, Gemini Live and Google Workspace surfaces. The date is corroborated by numerous independent outlets (9to5Google, The Decoder, Search Engine Land, TechTarget, Thurrott, Jetstream, AI Buzz Wire, TechRepublic). Within the story, capability claims (97 languages, background tool execution, benchmark scores) are COMPANY CLAIM unless independently verified by third-party measurement (partially independently measured by Artificial Analysis, reported by datanorth.ai and others).

Discovery framing verified: the release is widely read as Google's direct answer to OpenAI's GPT-Live-1 (launched September 10, 2026, five days earlier) and extends the Gemini 3.8 family (Flash and Flash Cyber, September 2, 2026).

Fact

Key deltas versus the previous state:

  • Latency architecture: single native speech-to-speech model vs. cascaded pipelines that Google/third parties estimate add 450–1,000 ms; Google positions sub-200 ms end-to-end (COMPANY CLAIM / EARLY RESEARCH — not independently measured at this time).
  • Reasoning model split: a low-latency volume SKU (3.8 Live, interleaved reasoning) and a high-reasoning SKU (Extended Thinking, background reasoning + narrated progress) — a deliberate two-SKU product architecture replacing the earlier gemini-3.1-flash-live-preview in the Live API.
  • Conversational state machine: with Extended Thinking, turnComplete no longer signals the end of an interaction; clients must track interaction_status (IN_PROGRESS / IDLE). This is a new client-side contract for voice-agent developers.
  • Async-by-default tools: non-blocking NON_BLOCKING function calling is the default on 3.8 Live and the only mode on Extended Thinking (blocking returns a hard error) — tools run while audio streams.
  • Visual grounding in consumer surfaces: Search Live (Google app "Live" icon) lets users point the camera at something and ask questions conversationally, with transcript links and AI Mode history.
  • Pricing: $0.005/min in / $0.018/min out (~$1.38/hour for overlapped conversation using published rates), undercutting OpenAI's GPT-Live-1 front-end at $0.05/min (≥$3.00/hour) — COMPANY CLAIM pricing confirmed by independent outlets (The Decoder, TPS, BitsMinds, BibiGPT), and roughly consistent with Artificial Analysis hourly cost measures reported by datanorth.ai ($0.84/hr for 3.8 Live, $3.50/hr for Extended Thinking input-audio, vs $5.83/hr GPT-Live-1 Astra).
✓

What happened?

🎓 For Explorer

On September 15, 2026, Google launched two real-time voice models:

  • Gemini 3.8 Live — "built for scale and cost efficiency," combining conversational intelligence, fluid dialogue, visual grounding (near real-time video/image input), automatic mid-conversation switching across 97 languages, and background execution of tools/API calls while the conversation continues. Priced at $0.005/minute audio input and $0.018/minute audio output. Rollout: Gemini API, Google AI Studio (developers); private preview in Gemini Enterprise (enterprises); Search Live in the Google app (consumers).
  • Gemini 3.8 Live Extended Thinking — "built for high-complexity tasks," adding configurable background reasoning (thinking levels low/medium/high) that runs while the model keeps speaking ("reasons and speaks simultaneously"), asynchronous non-blocking tool calls only, and live progress narration ("Let me check that…") that eliminates conversational dead air. Google states it holds the #1 spot on Artificial Analysis' Speech-to-Speech Quality Index (82.6). Rollout: Gemini API / AI Studio; private preview Gemini Enterprise; Gemini Live app; Workspace Docs Live (Google AI Pro/Ultra subscribers); Gmail Live and Keep Live (all Google AI subscribers).

Both are available via the Gemini Live API (stateful WebSocket, raw 16kHz PCM audio in, JPEG frames ~1fps, streaming audio out), with day-one integrations from Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel and Vision Agents, and launch partners including Salesforce, Genspark and Lumeris. All AI-generated audio is watermarked with SynthID.

Δ

What changed?

FACT: Google now offers production-grade, natively speech-to-speech voice models — no ASR→LLM→TTS cascade required — that can see, reason and act inside a single live session.

Key deltas versus the previous state:

  • Latency architecture: single native speech-to-speech model vs. cascaded pipelines that Google/third parties estimate add 450–1,000 ms; Google positions sub-200 ms end-to-end (COMPANY CLAIM / EARLY RESEARCH — not independently measured at this time).
  • Reasoning model split: a low-latency volume SKU (3.8 Live, interleaved reasoning) and a high-reasoning SKU (Extended Thinking, background reasoning + narrated progress) — a deliberate two-SKU product architecture replacing the earlier gemini-3.1-flash-live-preview in the Live API.
  • Conversational state machine: with Extended Thinking, turnComplete no longer signals the end of an interaction; clients must track interaction_status (IN_PROGRESS / IDLE). This is a new client-side contract for voice-agent developers.
  • Async-by-default tools: non-blocking NON_BLOCKING function calling is the default on 3.8 Live and the only mode on Extended Thinking (blocking returns a hard error) — tools run while audio streams.
  • Visual grounding in consumer surfaces: Search Live (Google app "Live" icon) lets users point the camera at something and ask questions conversationally, with transcript links and AI Mode history.
  • Pricing: $0.005/min in / $0.018/min out (~$1.38/hour for overlapped conversation using published rates), undercutting OpenAI's GPT-Live-1 front-end at $0.05/min (≥$3.00/hour) — COMPANY CLAIM pricing confirmed by independent outlets (The Decoder, TPS, BitsMinds, BibiGPT), and roughly consistent with Artificial Analysis hourly cost measures reported by datanorth.ai ($0.84/hr for 3.8 Live, $3.50/hr for Extended Thinking input-audio, vs $5.83/hr GPT-Live-1 Astra).
↔

Before → Change → After

🎓 For Explorer

Before (Sep 10–14, 2026): OpenAI's GPT-Live-1 (Sep 10) set the frontier for real-time voice with full-duplex conversation at $0.05/min (backend billed separately); Google's best live option was gemini-3.1-flash-live-preview and Gemini 3.8 Flash (Sep 2) — a multimodal speed model, not a dialogue model; most voice agents still ran cascaded ASR+LLM+TTS stacks that stalled ("dead air") during tool calls; voice assistants were largely turn-based.

Change (Sep 15, 2026): Two stable native speech-to-speech model IDs (gemini-3.8-live, gemini-3.8-live-extended-thinking) in the Live API with in-session reasoning, visual grounding, 97-language auto-switching, background async tools, narrated progress, and a ~7x cheaper hourly price point than the OpenAI front-end (on Artificial Analysis' input-audio measure).

After: Developers can build voice agents that talk, see, reason and act concurrently; Google's default consumer surfaces (Search Live, Gemini app, Workspace Docs/Gmail/Keep) run on the new models; enterprises can private-preview the models in Gemini Enterprise; the industry's voice-agent default UX shifts from "silence while thinking" to "narrated background work."

⚙

How it works

FACT (from official docs and model card): Both models are natively multimodal audio models within the Gemini 3 series ("Gemini 3.8 Audio"), processing continuous streams of audio, video and text to deliver spoken responses in real time. Technical properties:

  • Inputs/outputs: text, images, audio, video in; text + audio out. 131,072-token input window (model card: "up to 128K"), 65,536-token output.
  • Transport: Gemini Live API over a stateful WebSocket; raw 16 kHz PCM audio in, JPEG images (up to ~1 frame/second), text; ~24 kHz audio out (per third-party documentation summaries).
  • 3.8 Live: interleaved reasoning optimized for latency; NON_BLOCKING async function calling is default (blocking mode available for backward compatibility); full-session send_client_content with explicit user/model roles; video frames included in turns by default.
  • 3.8 Live Extended Thinking: background reasoning controlled by thinking_config (thinking_level: low/medium/high; MINIMAL not supported); async non-blocking tools only (function scheduling not supported); proactive audio permanently enabled; clients monitor interaction_status (IN_PROGRESS → IDLE) because turnComplete only ends an utterance; the model can emit multiple spoken segments (acknowledgements, progress narration, final answer) within one interaction.
  • Safety: every AI-generated audio output is watermarked with SynthID; model card notes the Frontier Safety Framework assessment (based on Gemini 3.7 Flash results) finds no Tracked/Critical Capability Levels likely; knowledge cutoff January 2025.

COMPANY CLAIM: "sub-200 ms" end-to-end latency positioning and the specific benchmark scores below.

Benchmarks (COMPANY CLAIM unless noted): Extended Thinking — Artificial Analysis Speech-to-Speech Quality Index 82.6 (#1 overall), τ-Voice agentic task completion 68.6%, Sierra τ-Voice-banking 35.1%, Big Bench Audio 97.7%; 3.8 Live — 2nd place in Speech Agent Arena; ServiceNow EVA-Bench Pareto-frontier results for complex workflows. Independent measurement (Artificial Analysis, as reported by datanorth.ai): Extended Thinking 82.6 vs GPT-Live-1 Astra 81.5; τ-Voice 68.6% vs 67.9%; τ-Voice-banking 35.1% vs 32.0%; Full Duplex Bench 91.9% vs 94.9% (OpenAI still leads); time-to-first-audio 1.35 s vs 1.34 s; hourly input-audio cost $3.50 vs $5.83 (base 3.8 Live: ~$0.84/hr, 76.0 index score). Note: datanorth.ai flags the GPT-Live-1 row as a single trial (most models get three), i.e., the OpenAI figure is noisier.

!

Why it matters

🎓 For Explorer
  1. The voice-interface race just got a price anchor. Google has undercut OpenAI's GPT-Live-1 (and xAI's Grok Voice Think Fast 2.0) on cost by a wide margin while roughly matching or edging ahead on quality indices — precisely the "scale + cost" play Google has run successfully before.
  2. "Dead air" is now a solved-looking problem. Extended Thinking's narrated background reasoning changes user expectations for every voice agent; competitors' UX will be judged against it.
  3. Distribution. This upgrades the default assistant experience across Google's surfaces (Search, Android, Workspace, Gemini app) — billions of users — turning voice-agent capability into a default feature, consistent with the discovery's "billions of Android/Google users" framing.
  4. Architectural statement. Keeping speech, reasoning and tool execution in one stateful session (vs. OpenAI's split front-end/backend design) is a bet that simplicity wins developer adoption.
  5. Week's context: lands five days after OpenAI's GPT-Live-1, two days after Apple's Siri AI event (Sep 14), in the same week as xAI's Voice Agent API ($0.08/min, Sep 17) — the voice-agent platform war is now the week's most concentrated competitive story.
✦

What became possible?

🎓 For Explorer
  • Voice agents that do real work while talking: booking flows, account investigations, refunds, troubleshooting — tools run async while the model narrates ("Let me check that…"), so complex tasks no longer require silence or "hold music."
  • Vision-in-the-loop conversation: pointing a phone camera at a problem (Search Live demo: a leaky pipe) and holding a natural conversation about what the camera sees, in near real time.
  • Mixed-language conversations on one model: automatic detection + switching across 97 languages mid-conversation with accent consistency.
  • Production voice agents at previously impossible unit economics: e.g., ~$13,800/month for 10,000 hours of voice traffic vs $30,000+ on GPT-Live-1 front-end pricing (TPS estimate).
  • In-session reasoning without backend orchestration: developers no longer need to build sideband context-passing to a separate reasoning model for every hard question.
◎

Implications

Technical

  • Stateful async dialog as a first-class API contract: interaction_status, turnComplete semantics, and async-only tool execution on Extended Thinking force a new client state machine (listening / speaking / awaiting-tool / recovering / idle). Clients designed for turn-based APIs will misbehave (INTERPRETATION, strongly supported by Google's own docs and third-party guides).
  • Native speech-to-speech displaces cascades: ASR→LLM→TTS pipelines (450–1,000 ms penalty) become an inferior default for voice-agent workloads.
  • Modal complexity: both models are multimodal-live (voice primary, vision context); structured outputs, code execution and caching are NOT supported on the Live endpoints — teams needing those must orchestrate around the session.
  • Duplex gap: independent benchmark data shows GPT-Live-1 leads on Full Duplex Bench (94.9 vs 91.9) — Google's "fluid dialogue with interruptions" may not yet match OpenAI's true simultaneous listen-and-speak. EARLY RESEARCH / needs hands-on verification.
  • Preview caveat: model IDs are stable, but the Live API itself remains in Preview (per multiple developer-facing write-ups) — pin model names, log events, and rerun canaries after SDK changes.

Developer

  • Switch path: existing gemini-3.1-flash-live-preview users get a documented migration (default NON_BLOCKING tools, new send_client_content semantics, video-by-default in turns).
  • Cost curve: $0.005/min in / $0.018/min out; Extended Thinking additionally bills reasoning tokens and extra inputs (video); budget for the reasoning-tax on long interactions — the "$0.84/hr" figure for base Live vs "$3.50/hr" for Extended Thinking is an input-audio measure, not a total bill.
  • Ecosystem readiness: day-one integrations for Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, Vision Agents mean existing voice stacks can adopt without replacing orchestration.
  • State management is now a correctness issue: on Extended Thinking, fire-and-forget turnComplete logic will drop final audio; monitor interaction_status and keep streaming the session until IDLE.
  • Session limits: reported 15-minute audio-only / 2-minute audio+video session caps (single secondary source — UNVERIFIED) imply session-resumption design for long workflows.

Enterprise

  • Customer-experience agents: Salesforce, Genspark, Lumeris, ServiceNow and 11Sight are named launch/testimonial partners; Extended Thinking's narrated tool use fits support, sales and claims workflows (τ-Voice-banking 35.1% vs 32.0% GPT-Live-1 Astra — independent measurement).
  • Gemini Enterprise: private preview at launch; Gemini Enterprise for Customer Experience and Workspace business tiers "coming soon" — enterprises should not assume GA parity yet.
  • Cost-sensitive scale: high-volume voice deployments (support lines, telephony bots) get dramatically better unit economics vs. OpenAI's front-end pricing.
  • Compliance surface: SynthID watermarking supports AI-disclosure obligations (EU AI Act transparency, state chatbot laws); but enterprise voice agents trigger the week's dominant governance themes — agent autonomy, audit trails and misalignment risk (see S01, S02).
  • Vendor strategy: enterprises now have a credible second full-stack realtime-voice vendor; multi-model voice architectures (Gemini for cost/vision, GPT-Live-1 for full-duplex feel) become defensible procurement positions.

Strategic

  • Google's playbook: win on cost + distribution + integration rather than raw conversational fidelity — same pattern as previous audio/video releases (The Decoder, TPS). Google is betting the default assistant experience beats the best-in-class demo.
  • OpenAI reaction risk: GPT-Live-1 launched Sep 10 at 10x Google's voice-layer price; expect a price cut or cheaper tier within months (PREDICTION, flagged by TPS independently).
  • xAI's counter: xAI shipped its Voice Agent API at $0.08/min and a no-code builder on Sep 17 — a different lane (telephony/agentic), showing the voice layer is fragmenting into "conversation model" vs "agent platform."
  • Benchmark wars: Google, OpenAI and xAI each cite different tests (Speech-to-Speech Quality Index vs Full Duplex Bench vs internal agentic suites); none of these is a neutral production-reliability measure (INTERPRETATION shared by multiple independent outlets).
  • Voice as the next OS surface: with Google (Search Live + Gemini app), Apple (Siri AI, Sep 14) and OpenAI (GPT-Live-1) all shipping realtime voice in one week, voice-first interaction is moving from demo to default — the strategic prize is conversational primacy on the phone.
⚠

Risks & limitations

Risks
  • Agentic blast radius: background async tool execution (payments, email, refunds) inside a spoken interface raises the stakes of misconfiguration and prompt injection via audio/video input — directly relevant to the week's agent-safety disclosures (Anthropic S01, RubyGems S14). Corporate claims of safety do not eliminate this.
  • Authoritative-sounding narration: a model that says "Let me check that…" while a tool fails or hallucinates can sound more reliable than it is; progress narration is generated, not a status log.
  • Voice deepfakes/spoofing: mass-deployed realtime voice generation increases impersonation surface; SynthID mitigates detection but not real-time abuse.
  • Interruption orphan work: when users change course mid-interaction, background tool calls may continue — who cancels them?
  • Price-led race to the bottom: Google's pricing pressure may commoditize voice-agent margins across the industry, hitting smaller voice platforms.
  • Duplex quality gap: if 3.8 Live's conversational naturalness trails GPT-Live-1 in real use, "cheap but robotic" could tarnish Google's flagship assistant story (The Decoder's critique).
Limitations
  • Knowledge cutoff January 2025 (model card) — stale for current-events conversations.
  • No structured outputs, no code execution, no caching on the Live endpoints.
  • Extended Thinking constraints: blocking tools unsupported (hard error), no function scheduling, MINIMAL thinking unsupported, proactive audio cannot be disabled.
  • Live API still in Preview despite stable model IDs — breaking changes possible.
  • Context/session caps: 131k input tokens; reported 15-min audio / 2-min audio+video session limits (UNVERIFIED, single secondary source).
  • Benchmark caveats: Google-cited scores are company claims; Artificial Analysis' GPT-Live-1 comparison row rests on a single trial; benchmarks measure voice-agent quality, not production reliability or meeting-summary quality.
  • Enterprise coverage is staged: private preview only; Caveat: Workspace business and CX availability "coming soon."
  • Latency claim unverified: sub-200 ms end-to-end is positioning, not yet independently measured at launch.
?

Open questions

  1. Does 3.8 Live's "fluid dialogue, interruptions" match GPT-Live-1's full-duplex behavior in head-to-head use (Full Duplex Bench says no, so far: 91.9 vs 94.9)?
  2. What does a real-world latency distribution look like under load (the sub-200 ms claim)?
  3. When does the Live API leave Preview, and what do GA prices look like (Gemini 3.8 Flash pricing reportedly ~doubles in 2027 per Gigazine via tech-insider.org — EARLY RESEARCH)?
  4. How does OpenAI respond on price, and how quickly?
  5. Will Gemini Enterprise GA include agent-platform parity (EVA-Bench was run on the Live API + Gemini Enterprise Agent Platform)?
  6. How do session limits evolve for video-in conversations?
  7. Does background-reasoning billing (reasoning tokens + video inputs) meaningfully erode the headline price advantage?
↗

What happens next?

🎓 For Explorer
  • Near term: OpenAI answers on price or a cheaper GPT-Live tier (PREDICTION); xAI extends its agent-platform lane; independent latency/duplex evaluations of 3.8 Live appear; Live API GA with final pricing; Gemini Enterprise CX + Workspace business rollout.
  • Mid term: voice-agent stacks standardize on "narrated background work" UX; Google extends Search Live-style camera conversations across surfaces; Multi-modal voice + agent platforms converge into the enterprise "voice operator" category.
  • Watch items: session-limit changes for video; Gemini 3.8 Flash 2027 price increases signalling the current Live pricing is land-grab territory; the duplex-quality debate resolving via head-to-head arena results.
★

Editorial takeaway

🎓 For Explorer

Google didn't win this week's voice race on naturalness — the independent duplex data, and several reviewers, say OpenAI still leads on conversational feel. Google won on the two things it has always weaponized: price (roughly 7x cheaper per hour) and distribution (Search Live, the Gemini app, Workspace, Android). The deeper story is architectural and systemic: with Extended Thinking, Google turned the voice agent into a stateful worker that talks while it works, and forced every developer to re-learn what "the conversation" means. In a week dominated by agent-misalignment alarms, the most important question in any of these launches is the one Google's model card answers with a page of safety attestations: the same background-tool autonomy that makes a voice agent feel alive is exactly the capability labs are still learning to contain. Speed, price and safety are now racing each other — and this launch put Google firmly in the lead on the first two.


Top-down view of a sealed tool chamber cycling requests on a loop, sending results along one narrow channel to a receiving box while a separate steady band of light runs past above.
·

Evidence label register (per AGENTS.md)

  • FACT: announcement date/rollout, model IDs, pricing figures as published, context/session limits as documented, partner list, SynthID watermarking, API semantics (async-only on Extended Thinking, interaction_status), knowledge cutoff.
  • COMPANY CLAIM: "most advanced live dialogue models," benchmark scores Google quoted (82.6 #1, τ-Voice 68.6%, τ-Voice-banking 35.1%, Big Bench Audio 97.7%, Speech Agent Arena 2nd, EVA-Bench Pareto), "sub-200 ms" latency, 97-language mid-conversation switching robustness, "built for scale and cost efficiency."
  • INDEPENDENT EVIDENCE: Artificial Analysis measurements as reported by datanorth.ai (82.6 vs 81.5; 91.9 vs 94.9 Full Duplex; 1.35 s vs 1.34 s TTF; $0.84/$3.50 vs $5.83 per-hour), same-day confirmation of date/rollout by multiple outlets, price comparisons in The Decoder/TPS/BitsMinds/BibiGPT.
  • INTERPRETATION: price-vs-quality tradeoff reading (The Decoder/TPS), in-session-vs-delegated architecture split (The New Stack/BibiGPT), cost leadership playbook, "voice as next OS surface."
  • PREDICTION: OpenAI price response; Live API GA timeline; voice-agent UX standardization on narrated background work.
⌘

Lab: NO-LAB

≡

Research sources

Primary Sources (6)
Primary
"Gemini 3.8 Live powers Google Search Live" — Search Engine Land (Barry Schwartz)** Confirms Search Live is live on the Google app on Sep 15; Rajan Patel (VP Engineering, Search) quote: "New Gemini audio models just dropped — 3.8 Live is now powering real-time conversations in Search Live"; consumer UX (tap "Live" icon, voice Q&A, links on screen, transcript link, AI Mode history) — ** PRIMARY (quoting Google source) + INDEPENDENT (independent outlet) ---Date: ** September 15, 2026
Visit source ↗
Primary
Gemini 3.8 Live Extended Thinking — Model Docs — Google AI for Developers** API contract for `gemini-3.8-live-extended-thinking`; async/non-blocking ONLY (blocking returns hard error); `thinking_config` (low/medium/high; MINIMAL not supported); `interaction_status` field (IN_PROGRESS/IDLE); `turnComplete` no longer means idle; proactive audio permanently enabled — ** PRIMARY — critical new API semantics and migration contractDate: ** September 2026
Visit source ↗
Primary
Gemini 3.8 Live — Model Docs — Google AI for Developers** API contract for `gemini-3.8-live`; model capabilities (audio generation supported, function calling supported, thinking interleaved, search grounding supported, caching NOT supported, structured outputs NOT supported); `NON_BLOCKING` default; `send_client_content` with roles; video-by-default turn coverage — ** PRIMARY — developer-facing API specificationDate: ** September 2026
Visit source ↗
Primary
Gemini 3.8 Audio (Live, Live Extended Thinking) — Model Card — Google DeepMind** Technical model card; 131,072 input / 65,536 output tokens; context window 128K; Frontier Safety Framework assessment (no Tracked/Critical Capability Levels likely, based on 3.7 Flash results); knowledge cutoff January 2025; limitations (hallucinations, jailbreak resistance); input modalities (audio, images, video, text) — ** PRIMARY — authoritative technical and safety documentationDate: ** September 15, 2026
Visit source ↗
Primary
"Build real-time voice applications with Gemini 3.8 Live and 3.5 Transcribe" — blog.google (Google DeepMind)** Developer-focused post; confirms pricing ($0.005/min audio input, $0.018/min audio output); lists asynchronous function calling, visual context, 97+ languages, alphanumeric precision, incremental content updates; S2S leaderboard rank (#1 Extended Thinking, #2 base Live); integration partners (LangChain added vs consumer post) — ** PRIMARY — official developer pricing and capability breakdownDate: ** September 15, 2026
Visit source ↗
Primary
"Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking" — blog.google (Google)** Official announcement; model capabilities, benchmarks (S2S Quality Index 82.6, τ-Voice 68.6%, τ-Voice-banking 35.1%, Big Bench Audio 97.7%); rollout tiers (developers/enterprise/consumers); SynthID watermarking; partner list (Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, Vision Agents, Salesforce, Genspark, Lumeris) — ** PRIMARY — company announcement, source of all benchmark claims and capability descriptionsDate: ** September 15, 2026 (updated September 17, 2026)
Visit source ↗
Independent Sources (10)
Independent
"Google Launches Gemini 3.8 Live Models That Can Reason While They Talk" — TechRepublic** Confirms both models "generally available through the Gemini API and Google AI Studio"; deployment of same standard audio rates for Extended Thinking; documents interaction_status protocol in practical detail; useful developer-facing summary with comparison table — ** INDEPENDENT — independent developer-facing documentation verification ---Date: ** September 17, 2026
Visit source ↗
Independent
"Google Launches Gemini 3.8 Live Voice Models With Real-Time Vision and 97-Language Switching" — AI Buzz Wire** Second place Speech Agent Arena for base 3.8 Live; different consumer availability mapping (Search Live = 3.8 Live; Extended Thinking = Gemini Live + Workspace Docs Pro/Ultra + Gmail/Keep all subscribers); confirms Gemini 3.8 family debut "earlier this month with Flash and Flash Cyber" — ** INDEPENDENT — independent availability verificationDate: ** September 15, 2026
Visit source ↗
Independent
"Google Launches Gemini 3.8 Live, Undercutting OpenAI's GPT-Live-1 on Price by Up to 70%" — TPS** Price analysis: 10,000-hr/month scenario = $13,800 Gemini vs $30,000+ OpenAI GPT-Live-1; flags OpenAI may respond "with a price cut or a cheaper tier of GPT-Live-1 in the coming months"; frames as "pricing play more than a technical leapfrog" — ** INDEPENDENT — independent pricing economics, prediction of OpenAI responseDate: ** September 15, 2026
Visit source ↗
Independent
"OpenAI's voice model doesn't think. That's the point." — The New Stack / TLDRocket** Architecture split articulation: Google keeps speech+reasoning+tools in one session vs OpenAI's separate backend; "the benchmarks don't settle the fight, because the stacks aren't the same"; contextualizes interrupt/cleanup complexity for both stacks — ** INDEPENDENT — independent architectural analysis, frames why benchmarks are non-comparableDate: ** September 15, 2026
Visit source ↗
Independent
"Gemini 3.8 Live vs GPT-Live-1: 10 Checks for Realtime Voice and Meeting Summary" — BibiGPT Blog** Detailed architecture comparison: in-session reasoning (Google) vs delegated backend (OpenAI); confirms Google 3.8 Live accepts images/video in the Live API session (GPT-Live-1 API is audio + text only at launch); $0.005/$0.018 vs $0.05/min; architecture context from The New Stack (Sep 15, 2026) — ** INDEPENDENT — independent deep-dive architecture comparisonDate: ** September 16, 2026
Visit source ↗
Independent
"Google Ships Gemini 3.8 Live and Extended Thinking, Tops Speech-to-Speech Quality Index at 82.6" — datanorth.ai (Alexis Dufresne)** Independent verification of Artificial Analysis measurements: S2S Index 82.6 vs 81.5 (GPT-Live-1 Astra); Full Duplex Bench 91.9% vs 94.9%; τ-Voice 68.6% vs 67.9%; TTF audio 1.35 s vs 1.34 s; independent hourly cost figures ($3.50/hr vs $5.83/hr). Flags GPT-Live-1 row as single trial — "treat the OpenAI column as the noisier of the two." — ** INDEPENDENT — independent benchmark measurement, with important caveat about OpenAI data qualityDate: ** September 15, 2026
Visit source ↗
Independent
"Gemini 3.8 Live Extended Thinking powers Gemini Live, Gmail, & Keep" — 9to5Google (Abner Li)** Consumer-facing surface mapping: Extended Thinking → Gemini Live, Docs Live (Pro/Ultra), Gmail Live, Keep Live (all Google AI subscribers); base 3.8 Live → Search Live; confirms benchmarks cited — ** INDEPENDENT — clear consumer-availability mappingDate: ** September 15, 2026
Visit source ↗
Independent
"Gemini 3.8 Live Transforms Conversational AI" — TechTarget (Esther Shittu)** Enterprise analyst quotes — Bradley Shimmin (Futurum Group): "just a natural conversation with the ability to interrupt it in real time to inject [context]"; confirms Sep 15 launch; frames enterprise use cases (customer support, sales, multimodal agentic apps); "catching up to vendors such as Apple" framing — ** INDEPENDENT — independent analyst commentary on enterprise positioningDate: ** September 16, 2026
Visit source ↗
Independent
"Google launches Gemini 3.8 Live to take on OpenAI's GPT-Live-1 at a fraction of the cost" — The Decoder (Matthias Bastian)** Independent framing: Google 3.8 Live "ranked first" on S2S leaderboard; price comparison ($1.38/hr vs ≥$3.00/hr OpenAI); cost-vs-quality tradeoff read: "OpenAI's model should still deliver more natural conversations thanks to full duplex"; Google "once again optimized for price over quality" — ** INDEPENDENT — independent price/quality analysis with direct competitive comparisonDate: ** September 15, 2026
Visit source ↗
Independent
Artificial Analysis — Speech to Speech Models and Providers Analysis** Independent leaderboard showing Google models (including Gemini 3.8 Live family) alongside GPT-Live-1, Grok Voice Think Fast 2.0 and Qwen models; independent benchmark data — foundation for cost/quality comparisons cited across all secondary reporting — ** INDEPENDENT — independent third-party benchmark providerDate: ** Accessed September 2026
Visit source ↗
Secondary Sources (7)
Secondary
"Google Announces Gemini 3.8 Live and 3.8 Live Extended Thinking" — Thurrott.com** Straightforward secondary tech-blog coverage; confirms Sep 15 date; quotes Google (Ouyang and Jaganathan); "most advanced live dialogue models" framing — ** SECONDARY — date and framing corroboration ---Date: ** September 15, 2026
Visit source ↗
Secondary
"Gemini 3.8 Live Launches: Voice Agents Beat GPT-Live-1" — Web Pulse (wpnews.pro)** Specific pricing: $3 per million input tokens (~$0.005/min); $12 per million output tokens (~$0.018/min); reports session limits: "15 minutes for audio-only and 2 minutes for audio plus video — design session resumption logic"; documents WebSocket spec (16kHz PCM in, 24kHz out, 128K context) — ** SECONDARY — granular API/session specs (session limits are single-source, flagged UNVERIFIED in research/S09.md)Date: ** September 16, 2026
Visit source ↗
Secondary
"Gemini 3.8 Live Tops Voice AI and Undercuts GPT" — BitsMinds** Independent Artificial Analysis data: Extended Thinking 82.6 vs GPT-Live-1 81.5 vs Grok 81.3; cost breakdown: ET $3.50/hr, base Live $0.84/hr; quality gap is narrow but price gap is wide; "three capabilities are new rather than incremental" (mid-conversation language switch, async tools, visual input) — ** SECONDARY — independent cost-quality analysisDate: ** September 16, 2026
Visit source ↗
Secondary
"Google Ships Gemini 3.8 Live for Production Grade Voice Agents" — Rohan Paul** Architecture detail: "native speech-to-speech architecture replaces the cascade pipeline (ASR + LLM + TTS), cutting latency from 450–1,000 ms to sub-200 ms"; names Salesforce, Genspark, Lumeris, LiveKit, Pipecat, Agora, Vercel as integration partners; documents τ-Voice score as measuring "multi-step task" completion (airline, retail, telecom scenarios) — ** SECONDARY — context on τ-Voice benchmark and architecture positioningDate: ** September 16, 2026
Visit source ↗
Secondary
"Google ships Gemini 3.8 Live and a separate Extended Thinking voice model" — Ground Truth** Flags that "Extended Thinking is a different endpoint, not a hidden reasoning toggle"; documents the `interaction_status` wrinkle for clients; notes "Availability is staged rather than universal" — warns against "rewritten as blanket availability across every Gemini consumer surface" — ** SECONDARY — editorial caution on availability interpretationDate: ** September 16, 2026
Visit source ↗
Secondary
"Gemini 3.8 Live / 3.8 Live Extended Thinking Officially Announced" — Jetstream** Comprehensive independent summary of rollout tiers; documents Gemini 3.8 Flash and 3.8 Flash Cyber as prior Gemini 3.8 family releases (Sep 2, 2026); correct consumer availability listing (Workspace: Docs AI Pro/Ultra, Gmail/Keep all Google AI Plus/Pro/Ultra) — ** SECONDARY — summary confirmationDate: ** September 15, 2026
Visit source ↗
Secondary
"Gemini 3.8 Live vs Extended Thinking: Which Should You Use?" — WaveSpeed Blog** Implementation comparison: lower client complexity (3.8 Live) vs higher complexity (Extended Thinking); flags "the wider Gemini Live API remains in Preview"; recommends pinning model versions, logging events, rerunning canaries after changes; notes Extended Thinking = "separate endpoint, not a hidden reasoning toggle" — ** SECONDARY — implementation/migration guideDate: ** September 16, 2026
Visit source ↗
Unverified Sources (1)
Unverified
"Google Launches Gemini 3.8 Live Voice Models With Real-Time Vision" — tech-insider.org** Detailed Gemini 3.8 family timeline (Sep 2 Flash launch, Sep 15 Live launch); references TestingCatalog reporting that model names surfaced on GCP console quota/metrics page hours before public announcement — pre-launch leak detail; notes Gigazine report that Gemini 3.8 Flash pricing may double in 2027 — ** UNVERIFIED — references third-party sources (TestingCatalog, Gigazine) not directly accessed; included only for timeline context; GCP leak and Flash price-escalation claims not independently verified hereDate: ** September 15, 2026
Visit source ↗