OpenAI adds GPT-Live-1 real-time full-duplex VOICE model to API (discovery said "video" — CORRECTED)
On Thursday Sep 10, 2026, OpenAI announced and shipped, effective immediately, general availability of GPT-Live-1 in the API: The model: GPT-Live-1 is OpenAI's premier full-duplex real-time voice model ("Our premier model for natural, expressive voice conversations with smooth interruption handling" — model card). It reasons over incoming and outgoing audio together in a single model, avoiding the latency and brittle handoffs of chained STT→LLM→TTS pipelines. It can listen while speaking, respond to interruptions and short acknowledgements ("mhmm", "yeah"), and keep a conversation running while a backend model or agent performs reasoning, tool calls, web search, or coding. Delegation architecture (the defining feature): GPT-Live-1 handles conversation only; reasoning/tools live in a backend the developer chooses. Two modes: Responses delegation (OpenAI-managed — you configure a Responses model such as gpt-5.6-luna, gpt-5.6-terra or gpt-6-astra with tools like web_search or custom functions) and client delegation (your application builds the request and runs its own backend model, agent harness, or service; results are fed back via session.commentary.append / session.thinking.append). The launch post demonstrates pairing GPT-Live-1 with the Codex SDK for repository work mid-conversation. OpenAI's guidance is explicit: business logic, authorization checks, tool handlers and durable state stay in the application. Transports: WebRTC for browser/mobile voice sessions (microphone/speaker media tracks + data channel), WebSockets for server-side audio streams, sideband connections for server control/observation of a live session, and Telephony/SIP support for phone agents (restaurant reservations, customer support, order updates). New voice options: 12 new real-time voices (Quartz, Ripple, Vesper, Willow, Stone, Gleam, Meridian, Bossa, Tempo, Beacon, Delta, Cinder) spanning accents, dialects and languages, plus the existing default marin; custom voices via sales eligibility. Built-in speech plumbing: native ASR transcripts (session.input_transcript.delta) and response text (session.output_transcript.delta), strong alphanumeric understanding, keyword biasing, and native turn detection for developers who want explicit turn boundaries even though the model is not turn-based. Pricing & limits: voice layer $0.05/min billed per second (not rounded up; a 90-second session = $0.075); backend model/tool usage billed separately at normal rates; free API tier not supported; rate limits measured in concurrent sessions: 25 (Tier 1) → 50 (Tier 2) → 200 (Tier 3) → 300 (Tier 4) → 500 (Tier 5). Knowledge cutoff: July 31, 2025. Paired same-day launch: OpenAI also moved the Agents API into public beta the same day (Sep 10) — GPT-Live-1 is the voice that talks to a customer; the Agents API (managed Codex harness) is a backend it can hand the real work to. OpenAI additionally pointed enterprises at OpenAI Presence (introduced Jul 22) as the managed enterprise voice-agent product built on GPT-Live-1. Day-one ecosystem: Twilio announced a native GPTLiveProvider in Agent Connect connecting Twilio Voice Media Streams directly to GPT-Live-1 sessions (Sep 10). Customer testimonials in the launch post: Yelp (Host/Hatch call handling), Speak (Live Tutor Lessons — ~80% fewer interruptions during thinking pauses), Intercom Fin (natural call flow for support), Cognition (Devin voice collaboration).

Tailored emphasis while keeping the full article available.
⌘ Jump to architecture, developer details, and the hands-on route.
The essential information in 30 seconds
On Thursday Sep 10, 2026, OpenAI announced and shipped, effective immediately, general availability of GPT-Live-1 in the API:
- The model: GPT-Live-1 is OpenAI's premier full-duplex real-time voice model ("Our premier model for natural, expressive voice conversations with smooth interruption handling" — model card). It reasons over incoming and outgoing audio together in a single model, avoiding the latency and brittle handoffs of chained STT→LLM→TTS pipelines. It can listen while speaking, respond to interruptions and short acknowledgements ("mhmm", "yeah"), and keep a conversation running while a backend model or agent performs reasoning, tool calls, web search, or coding.
- Delegation architecture (the defining feature): GPT-Live-1 handles conversation only; reasoning/tools live in a backend the developer chooses. Two modes: Responses delegation (OpenAI-managed — you configure a Responses model such as
gpt-5.6-luna,gpt-5.6-terraorgpt-6-astrawith tools likeweb_searchor custom functions) and client delegation (your application builds the request and runs its own backend model, agent harness, or service; results are fed back viasession.commentary.append/session.thinking.append). The launch post demonstrates pairing GPT-Live-1 with the Codex SDK for repository work mid-conversation. OpenAI's guidance is explicit: business logic, authorization checks, tool handlers and durable state stay in the application. - Transports: WebRTC for browser/mobile voice sessions (microphone/speaker media tracks + data channel), WebSockets for server-side audio streams, sideband connections for server control/observation of a live session, and Telephony/SIP support for phone agents (restaurant reservations, customer support, order updates).
- New voice options: 12 new real-time voices (Quartz, Ripple, Vesper, Willow, Stone, Gleam, Meridian, Bossa, Tempo, Beacon, Delta, Cinder) spanning accents, dialects and languages, plus the existing default
marin; custom voices via sales eligibility. - Built-in speech plumbing: native ASR transcripts (
session.input_transcript.delta) and response text (session.output_transcript.delta), strong alphanumeric understanding, keyword biasing, and native turn detection for developers who want explicit turn boundaries even though the model is not turn-based. - Pricing & limits: voice layer $0.05/min billed per second (not rounded up; a 90-second session = $0.075); backend model/tool usage billed separately at normal rates; free API tier not supported; rate limits measured in concurrent sessions: 25 (Tier 1) → 50 (Tier 2) → 200 (Tier 3) → 300 (Tier 4) → 500 (Tier 5). Knowledge cutoff: July 31, 2025.
- Paired same-day launch: OpenAI also moved the Agents API into public beta the same day (Sep 10) — GPT-Live-1 is the voice that talks to a customer; the Agents API (managed Codex harness) is a backend it can hand the real work to. OpenAI additionally pointed enterprises at OpenAI Presence (introduced Jul 22) as the managed enterprise voice-agent product built on GPT-Live-1.
- Day-one ecosystem: Twilio announced a native GPTLiveProvider in Agent Connect connecting Twilio Voice Media Streams directly to GPT-Live-1 sessions (Sep 10). Customer testimonials in the launch post: Yelp (Host/Hatch call handling), Speak (Live Tutor Lessons — ~80% fewer interruptions during thinking pauses), Intercom Fin (natural call flow for support), Cognition (Devin voice collaboration).
Context: GPT-Live-1 first shipped as the ChatGPT Voice model on July 8, 2026 (paid plans; GPT-Live-1 mini for Free). The API release on Sep 10 is the developer-surface debut of the same model family, 64 days later. OpenAI claims, in its own evaluations, a +30 percentage-point improvement on Full Duplex Bench over GPT-Realtime-2.1 and a #1 Tau3 finish when paired with GPT-6 Astra at medium reasoning effort (both COMPANY CLAIM).
- Real-time voice is now a commodity API building block: $0.05/min per-second billing for ChatGPT-Voice-grade full-duplex conversation makes "an AI that can hold a phone conversation" a routine engineering cost, not a research project — the same commoditization pattern as text tokens in 2023–24.
- The voice/brain split is an architectural statement: GPT-Live-1 deliberately separates conversation from reasoning. OpenAI is telling developers to stop treating voice as a model feature and start treating it as a layer you can swap while keeping your tools, guardrails and state machine. This is a different design center than the Realtime API (one model does all) and a direct challenge to incumbent speech stacks (Twilio-style orchestration with STT/LLM/TTS chains).
- It lands inside the week's live-voice platform race: same window as Google's Gemini 3.8 Live (Sep 15) and xAI's Voice Agent API at $0.08/min (Sep 17, Galaxy Day 3). OpenAI's $0.05/min voice layer sets the price reference point; Google answers with assistant-default distribution, xAI undercuts on telephony-agent positioning.
- Pairing with the Agents API (same day) and Presence completes the enterprise story: OpenAI now sells (a) a managed enterprise voice agent (Presence), (b) the raw voice model (GPT-Live-1 API), and (c) a managed reasoning harness (Agents API beta) — a vertical stack for customer-facing voice automation that competes with telecom/contact-center incumbents (Twilio, Five9, NICE) even as Twilio integrates it.
- It corrects the record on "multimodal voice": the GPT-Live-1 API accepts no images and no video (audio+text only). Anyone building "see the world" agents must continue using the Realtime API (image input) or other multimodal models — a meaningful constraint for the week's "live-video agent" narrative that the discovery record incorrectly attributed to GPT-Live-1.
CONFIRMED
- What: On Thursday 2026-09-10, OpenAI made GPT-Live-1 — the full-duplex real-time voice model that has powered ChatGPT Voice since July 8, 2026 — generally available in the API through a new Live sessions surface (
v1/live/sessions), priced at $0.05 per minute for the front-end voice layer (billed per second), with backend model/tool usage billed separately. GPT-Live-1 listens and speaks simultaneously, handles interruptions and background noise, and delegates reasoning and tool use to a separate backend model or agent (OpenAI Responses model or customer's own backend) so the conversation continues while work happens in the background. - ⚠ LOUD FLAG — discovery framing correction (not a date mismatch): The discovery record titles this story "OpenAI adds GPT-Live-1 real-time video model to API" and describes GPT-Live-1 as a "real-time multimodal (video+voice+text) model." Both characterizations are factually incorrect. GPT-Live-1 is a voice model whose only modalities are audio and text; image and video are explicitly unsupported in the API. Primary-source proof:
- OpenAI's official model page (
https://developers.openai.com/api/docs/models/gpt-live-1): "GPT-Live 1 is a full-duplex voice model for real-time conversations… Input modalities: audio, text — Output modalities: audio, text — Unsupported modalities: image, video" and "Videosv1/videos— Not supported." - OpenAI's own launch post (Sep 10): "a powerful, natural voice model for building voice-enabled apps" — every capability listed (interruption handling, delegation, tone/pace/style, telephony) is voice-specific.
- OpenAI Help Center (ChatGPT Voice FAQ): "Live does not support video or screen sharing" — video/screen-sharing remains exclusive to the older Advanced Voice Mode in ChatGPT consumer apps.
- Independent confirmation: Coursiv's Sep 11 technical review ("Inputs and outputs: Audio and text; images and video are not supported"), AI Pricing Guru ("It does not accept images"), The Decoder ("The speech model can listen and talk at the same time"), Unite.AI, TechRepublic, GIGAZINE — all describe a voice/speech model, none mention video.
- The "video" framing appears to be a discovery-stage conflation, most plausibly with (a) the Realtime API's
gpt-realtime-2.1, which does accept image input, (b) GPT-6 Astra's multimodal/video-adjacent capabilities released Sep 3, or (c) Advanced Voice Mode's video/screen-sharing, which GPT-Live explicitly does not include. Recommended orchestrator action: rename the story to reflect reality (e.g., "OpenAI adds GPT-Live-1 real-time full-duplex voice model to API") and correct the what_changed field: "OpenAI made GPT-Live-1, its real-time full-duplex voice model (audio+text), available through the API to all developers." The event date, organization, category and in-window eligibility all stand.
- OpenAI's official model page (
- Event-date verification trail (primary sources): OpenAI launch post datelined "September 10, 2026" (
https://openai.com/index/introducing-gpt-live-1-in-the-api/); API changelog entry "Sep 10 — Feature · Model: gpt-live-1 · API: v1/live/sessions — GPT-Live 1 is now generally available in the API" (https://developers.openai.com/api/docs/changelog); OpenAI Developers X post "GPT-Live-1 is now available in the API…" timestamped Sep 10, 2026; official developer-forum announcement posted 2026-09-10T18:57:29Z. Independent anchors on the same day: The Decoder ("Sep 10, 2026"), Unite.AI ("Published September 10, 2026"), Twilio ("September 10, 2026"). Verdict: event date 2026-09-10 falls inside the configured window — no mismatch. The correction needed is to the story's substance (video), not its date. - Sourcing caveats to flag:
- The discovery record lists TechCrunch among independent sources; no TechCrunch article URL for either the July 8 launch or the Sep 10 API launch could be located/verified during this research session (aggregators confirm TechCrunch covered the July 8 consumer launch, but no direct URL was retrievable). Corroboration instead comes from The Decoder, Unite.AI, TechRepublic, Coursiv, Twilio, GIGAZINE, AI Pricing Guru, Progressive Robot and diyai.
- The Verge covered the consumer launch on July 8, 2026 (before the research window) — cited here only as independent evidence of what GPT-Live-1 is (a full-duplex voice model; no video/screen sharing at launch), not as in-window coverage.
- Benchmark figures cited both in the OpenAI post and third-party roundups are OpenAI's own evaluations (Full Duplex Bench, Tau3, Tau Banking Voice) — COMPANY CLAIM until independently replicated; Coursiv explicitly notes "none of the current numbers come from outside OpenAI."
- Evidence labels used: FACT (event date, API GA, pricing, endpoint, modalities, unsupported modalities, voices, rate limits, knowledge cutoff, July 8 ChatGPT lineage, same-day Agents API beta, Twilio/Presence ecosystem moves), COMPANY CLAIM (benchmarks, Speak's ~80% interruption reduction, Yelp/Speak/Intercom/Cognition testimonials as published by OpenAI), INDEPENDENT EVIDENCE (multi-outlet agreement on headline facts; The Decoder, Unite.AI, TechRepublic, Twilio technical framing; Coursiv/AI Pricing Guru spec tables; Microsoft Learn Foundry page; community-reported production behavior), INTERPRETATION (voice/backend split as architectural shift; commoditization of full-duplex voice; competitive context vs Gemini 3.8 Live and xAI Voice Agent API), PREDICTION (mini variant in API, Realtime deprecation pressure, DevDay Sep 29 follow-ups, price competition at ~$0.05/min).
What happened?
On Thursday Sep 10, 2026, OpenAI announced and shipped, effective immediately, general availability of GPT-Live-1 in the API:
- The model: GPT-Live-1 is OpenAI's premier full-duplex real-time voice model ("Our premier model for natural, expressive voice conversations with smooth interruption handling" — model card). It reasons over incoming and outgoing audio together in a single model, avoiding the latency and brittle handoffs of chained STT→LLM→TTS pipelines. It can listen while speaking, respond to interruptions and short acknowledgements ("mhmm", "yeah"), and keep a conversation running while a backend model or agent performs reasoning, tool calls, web search, or coding.
- Delegation architecture (the defining feature): GPT-Live-1 handles conversation only; reasoning/tools live in a backend the developer chooses. Two modes: Responses delegation (OpenAI-managed — you configure a Responses model such as
gpt-5.6-luna,gpt-5.6-terraorgpt-6-astrawith tools likeweb_searchor custom functions) and client delegation (your application builds the request and runs its own backend model, agent harness, or service; results are fed back viasession.commentary.append/session.thinking.append). The launch post demonstrates pairing GPT-Live-1 with the Codex SDK for repository work mid-conversation. OpenAI's guidance is explicit: business logic, authorization checks, tool handlers and durable state stay in the application. - Transports: WebRTC for browser/mobile voice sessions (microphone/speaker media tracks + data channel), WebSockets for server-side audio streams, sideband connections for server control/observation of a live session, and Telephony/SIP support for phone agents (restaurant reservations, customer support, order updates).
- New voice options: 12 new real-time voices (Quartz, Ripple, Vesper, Willow, Stone, Gleam, Meridian, Bossa, Tempo, Beacon, Delta, Cinder) spanning accents, dialects and languages, plus the existing default
marin; custom voices via sales eligibility. - Built-in speech plumbing: native ASR transcripts (
session.input_transcript.delta) and response text (session.output_transcript.delta), strong alphanumeric understanding, keyword biasing, and native turn detection for developers who want explicit turn boundaries even though the model is not turn-based. - Pricing & limits: voice layer $0.05/min billed per second (not rounded up; a 90-second session = $0.075); backend model/tool usage billed separately at normal rates; free API tier not supported; rate limits measured in concurrent sessions: 25 (Tier 1) → 50 (Tier 2) → 200 (Tier 3) → 300 (Tier 4) → 500 (Tier 5). Knowledge cutoff: July 31, 2025.
- Paired same-day launch: OpenAI also moved the Agents API into public beta the same day (Sep 10) — GPT-Live-1 is the voice that talks to a customer; the Agents API (managed Codex harness) is a backend it can hand the real work to. OpenAI additionally pointed enterprises at OpenAI Presence (introduced Jul 22) as the managed enterprise voice-agent product built on GPT-Live-1.
- Day-one ecosystem: Twilio announced a native GPTLiveProvider in Agent Connect connecting Twilio Voice Media Streams directly to GPT-Live-1 sessions (Sep 10). Customer testimonials in the launch post: Yelp (Host/Hatch call handling), Speak (Live Tutor Lessons — ~80% fewer interruptions during thinking pauses), Intercom Fin (natural call flow for support), Cognition (Devin voice collaboration).
Context: GPT-Live-1 first shipped as the ChatGPT Voice model on July 8, 2026 (paid plans; GPT-Live-1 mini for Free). The API release on Sep 10 is the developer-surface debut of the same model family, 64 days later. OpenAI claims, in its own evaluations, a +30 percentage-point improvement on Full Duplex Bench over GPT-Realtime-2.1 and a #1 Tau3 finish when paired with GPT-6 Astra at medium reasoning effort (both COMPANY CLAIM).
What changed?
- Before: Developers building real-time voice agents used the Realtime API (
v1/realtimewithgpt-realtime-2.1/2.1-mini, released Jul 6, 2026), a single all-in-one speech-to-speech model that handles audio, reasoning and tool selection inside one model. GPT-Live-1 existed only as a consumer ChatGPT feature (Jul 8) with plan-based hourly caps, no public model ID, no pricing, and a "notify me" signup for API access (as documented by multiple builders through mid-July). - Change: OpenAI made GPT-Live-1 generally available to all developers through the new
v1/live/sessionsendpoint at $0.05/min (voice layer, per second), with full-duplex conversation handled by a dedicated voice model and reasoning/tools delegated to a separately chosen backend (Responses delegation or client delegation). SynthID watermarking had already been added to supported GPT-Live audio on Jul 31, signaling API readiness. - After: Any developer with a paid OpenAI API key can now embed ChatGPT-Voice-grade full-duplex conversation in their own app, phone line or agent, choosing their own reasoning backend per task (cheap model for scheduling, frontier model for hard cases). The voice layer is commoditized at a legible price; the architecture splits "voice" from "brain," making each independently replaceable. Real-time voice moves from consumer novelty to a standard API building block — in the same week Google (Gemini 3.8 Live, Sep 15) and xAI (Voice Agent API at $0.08/min, Sep 17) shipped competing real-time voice surfaces.
Before → Change → After
| Phase | State |
|---|---|
| Before | Voice agents built on the Realtime API: gpt-realtime-2.1 (Jul 6, 2026) does speech+reasoning+tools in one model; image input supported; per-token audio pricing ($32/$64 per 1M audio in/out). GPT-Live-1 locked inside ChatGPT Voice (Jul 8): Go/Plus/Pro defaults, Free uses mini, plan caps up to 15h/day, no API access ("notify me" form), no public model ID. Video/screen-sharing only in Advanced Voice Mode. |
| Change | Sep 10, 2026: GPT-Live-1 GA in the API on v1/live/sessions at $0.05/min voice (per-second billing); full-duplex voice layer + delegated backend (Responses or client); WebRTC/WebSocket/Telephony; 12 new voices; native transcripts, keyword biasing, turn detection; no images/video; no free tier; 25–500 concurrent sessions by tier. Paired with Agents API public beta; Twilio GPTLiveProvider same day. |
| After | Developers can build production full-duplex phone/browser voice agents with a replaceable brain; voice-layer price ~$0.05/min invites volume use-cases (support, reservations, order status, tutoring); the frontend/backend split changes how voice apps are architected (prompt split, event-model rewrite, no manual turn control); competitive pressure on Realtime API migration; Google (Gemini 3.8 Live) and xAI (Voice Agent API $0.08/min) respond within the same week; OpenAI DevDay (Sep 29) expected to extend the surface (mini variant, video, more voices). |
How it works
⌘ For Builder- Single-model full duplex: GPT-Live-1 jointly models user input audio and generated output audio, so it can decide to keep talking, stop, acknowledge ("mhmm"), or pivot mid-stream without a voice-activity-detection state machine owned by the developer. This replaces the STT→LLM→TTS concatenation where each handoff adds latency and timing drift (OpenAI's framing, echoed by GIGAZINE's analysis).
- Session lifecycle (
v1/live/sessions): WebSocket clients connect towss://api.openai.com/v1/live/sessionsand sendsession.start(model, instructions, audio format, voice, delegation config); the server repliessession.started. WebRTC sessions are created withPOST /v1/live/sessions(server exchanges the browser's SDP offer for an answer; audio travels on media tracks, JSON events on a data channel). Sessions are fixed-config at startup: model, instructions, voice and audio format are immutable (instructions can be appended viasession.instructions.append; delegation settings are sparse-updatable within a mode). - Delegation flow: GPT-Live-1 decides when a request needs the backend; it emits
session.delegation.createdwith adelegation_id. In Responses delegation, it prepares a request to your configured Responses model (with tools such asweb_searchor functions), which runs and returns results throughresponse.eventenvelopes; the voice model then communicates the result. In client delegation, your application assembles the backend call itself and returns content viasession.commentary.append(content to speak, paraphrasable) orsession.thinking.append(quiet context). Interrupting speech does not cancel backend work — the application controls permissions, task state and which results reach the conversation. - Context management: instructions up to 16,384 tokens at startup; automatic conversation summarization/compaction — when context exceeds 90%, GPT-Live-1 starts a replacement voice engine within the same session seeded with original instructions + up to 8,192 tokens of history (recent messages + summary).
session.instructions.append/session.thinking.append/session.commentary.appendinject content with acknowledgedstart_ms/end_mswindows on the session timeline. - Transcripts:
session.input_transcript.deltaandsession.output_transcript.deltadeliver timed fragments for UI, logging, moderation and early-start work; the API also reports cumulativesession.usage.updated(~once per minute; final usage atsession.closed). - Audio formats: WebSocket PCM16 mono 24 kHz (
audio/pcm) or G.711 μ-law 8 kHz (audio/pcmu) for direct telephony bridging (e.g., Twilio Media Streams); WebRTC negotiates its own format. Session durations billed per second and not rounded up; a WebRTC creation bills 15s of initialization that is credited once the session starts (cost-optimization guide). - Data controls: GPT-Live-1 is eligible for Zero Data Retention (with limitations; stored sessions ignored under ZDR), processes in US or EU/EEA+Switzerland regions (2026 guide data per Progressive Robot), and session recording (store:true) is off by default, retained 30 days when enabled.
Why it matters
- Real-time voice is now a commodity API building block: $0.05/min per-second billing for ChatGPT-Voice-grade full-duplex conversation makes "an AI that can hold a phone conversation" a routine engineering cost, not a research project — the same commoditization pattern as text tokens in 2023–24.
- The voice/brain split is an architectural statement: GPT-Live-1 deliberately separates conversation from reasoning. OpenAI is telling developers to stop treating voice as a model feature and start treating it as a layer you can swap while keeping your tools, guardrails and state machine. This is a different design center than the Realtime API (one model does all) and a direct challenge to incumbent speech stacks (Twilio-style orchestration with STT/LLM/TTS chains).
- It lands inside the week's live-voice platform race: same window as Google's Gemini 3.8 Live (Sep 15) and xAI's Voice Agent API at $0.08/min (Sep 17, Galaxy Day 3). OpenAI's $0.05/min voice layer sets the price reference point; Google answers with assistant-default distribution, xAI undercuts on telephony-agent positioning.
- Pairing with the Agents API (same day) and Presence completes the enterprise story: OpenAI now sells (a) a managed enterprise voice agent (Presence), (b) the raw voice model (GPT-Live-1 API), and (c) a managed reasoning harness (Agents API beta) — a vertical stack for customer-facing voice automation that competes with telecom/contact-center incumbents (Twilio, Five9, NICE) even as Twilio integrates it.
- It corrects the record on "multimodal voice": the GPT-Live-1 API accepts no images and no video (audio+text only). Anyone building "see the world" agents must continue using the Realtime API (image input) or other multimodal models — a meaningful constraint for the week's "live-video agent" narrative that the discovery record incorrectly attributed to GPT-Live-1.
What became possible?
- Full-duplex phone agents in a weekend: reservations, order status, customer support and outbound calls where the caller can interrupt, change direction and talk over the agent — previously a multi-vendor integration project (STT+VAD+LLM+TTS+telephony), now a
session.startcall plus a Twilio Media Stream bridge (Twilio shipped the native GPTLiveProvider the same day). - Bring-your-own-brain voice: pair GPT-Live-1 with Luna for cheap high-volume scheduling, Terra for balanced work, Astra for hard reasoning, or a third-party/self-hosted model via client delegation — voice quality is decoupled from reasoning capability, so teams can swap backends without touching the voice layer.
- Voice-enabled coding and ambient agents: the launch post demonstrates GPT-Live-1 + Codex SDK (talk to a repo-reading agent), and the cost guide sketches ambient agents that close the voice session while a long task runs and resume on completion ("Resume conversation" button).
- Natural-language telephony UX: barge-in handling, "mhmm" backchannels, silence tolerance, and keyword-biased alphanumeric capture (order IDs, dates) — the interaction grammar of real phone calls rather than IVR menus.
- Language/voice localization: 12 accent/dialect/language voices at launch with more promised, letting products localize voice agents as a configuration change rather than a TTS vendor swap.
Implications
⌘ For BuilderTechnical
- Two live-voice API paradigms now coexist:
v1/realtime(single-model speech+reasoning+tools; image input supported) versusv1/live/sessions(dedicated voice layer + delegated backend; no image/video). Teams must pick deliberately; OpenAI's migration guide is explicit that moving Realtime→Live is not a model-name swap: prompts split (conversation style → session instructions; business rules/tool schemas → delegation instructions), tools move to the delegation schema, manual turn control is removed (audio streams continuously; the model decides when to speak), and audio/transcript events are renamed onto the session. - Event-model rewrite:
session.input_audio.append(unacknowledged),session.output_audio.delta(timed), transcript deltas,session.delegation.created,response.eventenvelopes wrapping nested Responses events,session.usage.updatedsnapshots — a materially different protocol surface than the Realtime API'sresponse.create/input_audio_buffermodel. Community developers report real migration friction (see §13). - Latency/turn-taking gains are real but vendor-measured: OpenAI reports 0.798s turn-taking latency (Full Duplex Bench v1) vs 1.41s for gpt-realtime-2.1 and +30pp on Full Duplex Bench v3 — submitted only from OpenAI's own evals; no independent benchmark existed as of Sep 11 (Coursiv).
- Concurrency-based rate limiting changes capacity planning: 25–500 concurrent sessions by tier means production voice apps plan around simultaneous calls, not tokens — a different operational model (and a cap that tier-1 teams will hit quickly in telephony rollouts).
- Context compaction is now a first-class session feature: automatic >90% summarization with a same-session engine replacement changes long-conversation assumptions versus Realtime.
Developer
- New SDK surface:
client.live.connect()/client.live.create()in official SDKs (Node, Python) with WebRTC and WebSocket flows; a WebRTC quickstart (browser+server) and a WebSocket quickstart (PCM16 24kHz stdin/stdout) are the canonical entry points. Python needsopenai[realtime]; WebRTC needs Node ≥ 22.6. - Architecture guidance is unusually opinionated: keep business logic, authorization, tool handlers and durable state in your application; never copy a Realtime prompt wholesale; treat the voice layer as replaceable. Client delegation is the escape hatch for custom routing, third-party backends, and result redaction before the voice model speaks.
- Cost modeling changes: voice = duration-based ($0.05/min), backend = token-based at its own rate. A 10-minute call is $0.50 voice + backend + tools. Per-second billing (no rounding) plus explicit
session.closediscipline (idle sessions bill continuously) and the 15s WebRTC-init credit are the key optimizations. Teams should estimate cost-per-successful-task, not per-minute. - Community-reported gotchas (EARLY RESEARCH — OpenAI forum, Sep 18, several builders):
- User-side audio transcription is not retrievable via the API in some deployments — the recorded transcript contains only assistant text, not the caller's words (the model card says transcripts are provided natively, but at least one team could not retrieve user speech; discrepancy unresolved — treat with caution).
response.done.outputarrives empty in delegation mode; function calls come individually viaresponse.output_item.doneinsideresponse.eventenvelopes.- Double
response.completedevents (inner delegation-level + outer session-level) can double-advance state machines. - No
session.update/response.cancel/input_audio_buffer.clearequivalents over the data channel as of Sep 18 — goodbye flows and interruption handling need new patterns; tool-dependentend_conversationflows often never fire because the voice model answers social closers directly. - Voice-model commentary can precede
session.delegation.created, breaking naive "voice spoke ⇒ delegation" turn detection.
- Migration economics: for teams already on Realtime, immediate benefits are per-second billing, better tool calling, interruptions and speed (per community testing); blocking issues are transcript retrieval and event-model rewrites. Several teams advise running both stacks during migration.
Enterprise
- Customer-facing voice automation is now buyable at API prices: contact-center-like inbound/outbound voice (reservations, order status, support triage) becomes an internal build on $0.05/min voice + a reasoning backend, or a procurement decision via OpenAI Presence (managed enterprise voice agents, launched Jul 22, limited GA via OpenAI Forward Deployed Engineers and systems integrators).
- Day-one enterprise validation: Yelp (Host/Hatch call handling), Intercom (Fin), Speak (tutoring), Cognition (Devin) all ship on GPT-Live-1; Twilio's native Agent Connect integration lowers last-mile telephony risk for the enterprise channel.
- Data-control posture matters for regulated sectors: GPT-Live-1 is ZDR-eligible with limitations and supports EU/EEA+Switzerland regional processing, but a delegated Agents API backend is US-only and not ZDR-eligible — a European voice-agent deployment loses end-to-end ZDR the moment it hands work to the Agents API. Stored sessions (store:true) persist 30 days with no public deletion endpoint. Compliance teams must map the voice layer and each backend separately.
- Concurrency ceilings cap telephony scale: 500 concurrent sessions at Tier 5 bounds simultaneous calls — call-center peak-load planning changes from token budgets to session budgets; overflow strategies (queuing, regional sharding) are needed at scale.
Strategic
- OpenAI is defining the "voice layer" as a product category — a thin, cheap, replaceable conversation front end over any reasoning backend. That framing pressures both ends of the market: speech incumbents (assembly-style pipelines, turn-based voice platforms) and frontier-model rivals who bundle voice with intelligence.
- Pricing marker set at $0.05/min: xAI's $0.08/min Voice Agent API (Sep 17) and Google's Gemini 3.8 Live (Sep 15) will be benchmarked against it; expect voice-layer price compression toward $0.01–0.03/min within quarters, following the token-price pattern.
- Ecosystem lock-in via delegation: optional today, "Responses delegation" makes OpenAI the orchestration layer between the voice and the brain; client delegation keeps customers open. The strategic prize is owning the glue (voice→backend protocol, presence, agents) rather than the voice model alone.
- Same-week narrative: in a week dominated by agent-safety disclosures (S01, S02, S14) and the pacing debate (S15), OpenAI shipping a production full-duplex phone voice (plus the Agents API) is the "deploy now" counter-narrative — agentic voice is being commoditized while the industry argues about controlling agents.
- Foundry distribution: Microsoft Foundry documents GPT-Live natively under Azure Foundry endpoints (learn.microsoft.com GPT-Live event reference), extending OpenAI's voice layer into enterprise Azure channels alongside the GPT-6 Astra GA there (S06).
Risks & limitations
- Untested-at-scale voice commerce: consumer ChatGPT Voice is proven at 150M+ users (per OpenAI's July claims), but enterprise phone-call reliability (noise, accents, call quality, DTMF, hold music) is new territory; early community reports of hangup/delegation unreliability in semi-rigid flows are a yellow flag for production telephony.
- Data/privacy edge cases: 30-day stored-session retention with no public deletion endpoint; ZDR gaps when delegation crosses to Agents API; user-speech transcript handling uncertainty (see §13) matters for call-recording consent regimes (two-party consent states, GDPR).
- Migration traps: teams porting Realtime code naively will hit empty
response.done.output, doubleresponse.completed, and missing session controls — production incidents during migration are likely. - Concurrency throttling: tier caps (25 at Tier 1) will surprise early telephony deployers; silent session-idle billing can inflate invoices; per-second metering makes idle-session hygiene a cost control.
- Voice impersonation/safety: expanded voice roster (accents/dialects/languages) plus streamable cloning-adjacent custom voices (sales-gated) re-raises voice-spoofing and social-engineering risk; OpenAI's SynthID watermarking (Jul 31) mitigates but does not eliminate it.
- Benchmark overclaiming: +30pp Full Duplex Bench, Tau3 #1 and 0.798s latency are OpenAI's own evals; blind replication is pending. Procurement teams should demand vendor-neutral evaluation (per Coursiv's warning).
- Framing risk to the newsroom: the discovery record's "video model" error, if propagated, would mislead readers about both GPT-Live-1's capabilities and the state of OpenAI's real-time video API; this report corrects it (see §1).
- No image or video input/output — audio+text only; the "see the world" real-time agent use case is NOT served by this model (contrary to discovery). Realtime API
gpt-realtime-2.1supports image input; GPT-Live does not. - Endpoint confinement:
v1/live/sessionsonly — not callable via Chat Completions, Responses, Realtime, Assistants or Batch; no fine-tuning, structured outputs, or predicted outputs; notemperature/top_pcontrols (fixed). - Voice fixed per session (immutable after startup; new session to change), output audio PCM16 24kHz mono (no stereo), and input transcript retrieval uncertainty for user audio (community reports conflict with model-card promises; unresolved as of Sep 18).
- No free tier — every evaluation costs money; no public mini variant in the API (GPT-Live-1 mini exists in ChatGPT consumer; model page for the API mini 404s as of Sep 11, per Progressive Robot).
- Knowledge cutoff July 31, 2025 — current-events answers depend entirely on backend delegation + web search.
- Regional/data constraints: EEA processing only via regional domains; UK listed for regional storage only (2026 data-controls documentation); ZDR not end-to-end when delegating to Agents API.
- All headline benchmarks are vendor-measured (OpenAI evals; backend-specific — Terra low vs Astra medium materially changes scores).
Open questions
- Will GPT-Live-1 mini (and Medium/High reasoning variants) ship in the API, at what price, and when? (DevDay Sep 29 is the near-term candidate.)
- Will the Realtime API line get a deprecation date now that GPT-Live-1 is GA — forcing migration for
gpt-realtime-2.1users? - Can user-side speech transcripts actually be retrieved via the API (discrepancy between model-card claims and community testing)? OpenAI's clarification would settle a production blocker.
- When will video and screen-sharing arrive for GPT-Live (OpenAI said "soon" at the July launch; still absent Sep 10)? Its arrival would finally justify the discovery record's "video" framing — retroactively wrong, but directionally predictive.
- Will OpenAI publish independent third-party full-duplex benchmarks (or open Full Duplex Bench/Tau3 methodology) to back the +30pp and #1 claims?
- How will voice-layer pricing evolve once Google (Gemini 3.8 Live) and xAI ($0.08/min) publish comparable benchmark numbers?
- Does OpenAI Presence become the enterprise distribution layer that sidelines self-serve GPT-Live-1 API deployments, or do both scale?
What should you do with this?
⌘ For BuilderCircle 1 (immediate, weeks):
- For voice/agent developers: exercise the WebRTC quickstart and WebSocket quickstart with a paid key; validate interruption handling, barge-in, and delegation (start with
gpt-5.6-luna+web_search); measure time-to-first-audio and per-second bills on 2–5 min sessions before committing to production. Because the free tier is unsupported, budget a small paid evaluation (<$5 covers ~30 minutes of voice + backend). - For teams on Realtime today: start migration planning now but do NOT port mechanically — audit every
response.done.outputread, event-counting state machine (doubleresponse.completed), and hangup flow against the new event model before switching; run both stacks in parallel for at least one release cycle. - Recommended action: correct the discovery-level story framing (video → full-duplex voice) in STORY_INDEX/PROGRESS as flagged in §1, then proceed with the corrected title.
Circle 2 (this quarter):
- Contact-center/telephony operators: pilot GPT-Live-1 behind Twilio Agent Connect for reservations/order-status triage; plan peak concurrency against tier caps (request tier upgrades early); put idle-session timeout and per-second metering controls in place.
- Compliance teams: map the data plane per component (voice layer ZDR-eligible with limitations; Agents API backend US-only, non-ZDR; 30-day stored sessions) before GDPR/CCPA-facing deployments; decide stored-session policy explicitly (store:true is off by default).
- Platform/enterprise architects: evaluate Foundry/Foundry-side GPT-Live support (Microsoft Learn reference) for Azure-aligned shops; compare against xAI Voice Agent API ($0.08/min) and Gemini Live API as multi-vendor options; define a vendor-neutral full-duplex evaluation harness (the industry lacks one — building it internally is a differentiator).
Circle 3 (12+ months):
- Speech/voice-platform incumbents (telecom, CCaaS, speech AI vendors): treat full-duplex voice-at-$0.05/min as a structural price shock; reposition toward integration, compliance, vertical workflows and proprietary data (the model alone will not sustain margins).
- Frontier labs and regulators: the voice/brain split creates a new audit surface (who is accountable for what the voicelayer says vs what the backend does); expect the disclosure/governance frameworks published this week (S02's misalignment framework; EU AI Act GPAI obligations effective Sep 15) to be tested by voice-agent deployments (GDPR audio-data processing, AI Act transparency for conversational systems).
- Buyers of "AI receptionists/voice agents": pricing pressure is coming; short-term contracts (≤12 months) avoid lock-in at today's rates.
- Vertical full-duplex phone agents (restaurants, clinics, salons, field service, order status): a $0.05/min voice layer plus a small backend makes per-call cost ~$0.10–0.30 — viable vs $1–3 human-assisted calls; consulting teams can productize "GPT-Live-1 + tooling" blueprints.
- Voice agent building SaaS: the migration pain (prompt split, event model, hangup flows, cost telemetry) is a services/software opportunity — observability, evals (spoken-task Pass@1), and Realtime→Live migration tooling are unfilled gaps.
- Language/localization play: 12 accent/dialect/language voices make regional voice products cheaper to localize; training/consulting on multilingual voice UX is a viable niche.
- Training content: "build a phone agent in 2 hours" workshops (WebRTC quickstart → Twilio bridge → delegation) are strong lead-gen material for agencies serving SMBs — the demo is genuinely impressive and cheap to run.
INSPECT (follow-along) is the highest-integrity option available in this environment; VERIFY is gated. Rationale:
- GPT-Live-1 does not support the free API tier (model page: "Unsupported usage tiers: Free"), so a live call requires a paid OpenAI API key with billing enabled and incurs real charges ($0.05/min voice + backend tokens). This research environment has no API credentials or billing relationship, so a live
session.startcannot honestly be executed here. - Recommended once a paid key exists: a VERIFY protocol — (1) POST
/v1/live/sessions(WebRTC) from a browser page on localhost with audio capture; (2) confirmsession.startedand time-to-first-audio for a greeting; (3) interrupt mid-speech and confirm continued natural turn-taking; (4) ask a current-events question withgpt-5.6-luna+web_searchResponses delegation and confirm the result arrives inside the conversation; (5) measure per-second billing on a 2-minute session and compare with the documented formula (90s = $0.075 voice). Full procedure and a documented code walkthrough (WebSocket quickstart JSON, WebRTC create call, delegation config, event list) are inlabs/S52.md.
What happens next?
- Near term (days–weeks): developer community pressure on OpenAI to (a) expose user-side transcripts, (b) add
session.cancel-like controls, (c) confirm Realtime deprecation timing; expect documentation updates and possible API tweaks before OpenAI DevDay on Sep 29 (Fort Mason, San Francisco) — the natural venue for GPT-Live-1 mini, video support announcements, and benchmark methodology releases. - This week's competitive clock: Google's Gemini 3.8 Live (Sep 15) and xAI's Voice Agent API (Sep 17) give buyers three full-duplex voice surfaces within 7 days; benchmark shootouts (Artificial Analysis-style) are imminent.
- Data/legal follow-ups: OpenAI published a GPT-Live system card (with a corrected/re-run safety evaluation noted Aug 4); voice-agent deployments will start showing up in EU AI Act GPAI transparency reporting and GDPR audio-processing audits.
- Product trajectory: expect voice-layer price compression, a mini-tier for low-cost telephony, video/screen-sharing parity with Advanced Voice Mode ("working to introduce these capabilities soon" per July launch materials), deeper Codex/Agents-API integration, and Foundry/enterprise adoption as the GPT-6 Astra GA (Sep 17, S06) normalizes frontier models on Azure.
Editorial takeaway
Report this story as "OpenAI opens its full-duplex real-time voice model to all developers — GPT-Live-1 hits the API at $0.05/minute" — NOT as a video model. The correction matters: GPT-Live-1 accepts audio and text only, and OpenAI is explicit that image and video are unsupported, so the "real-time video agent" angle belongs to other models (Realtime with image input, or future GPT-Live updates). The genuinely newsworthy substance is (1) ChatGPT-Voice-grade full-duplex conversation as a commodity API at a legible $0.05/min per-second price; (2) the architectural bet that voice and reasoning should be separable layers — a reframing of how voice agents are built, with real migration costs for Realtime users; (3) the same-week collision with Google's Gemini 3.8 Live and xAI's $0.08/min Voice Agent API, making this the week real-time voice became a platform war; and (4) the pairing with the Agents API beta and OpenAI Presence, completing OpenAI's enterprise voice stack. Flag the benchmark claims as vendor-measured, and surface the early community-reported production gaps (transcript retrieval, hangup reliability) honestly — they are the difference between a demo and a phone line.

Lab: INSPECT
⌘ For Builder- WebRTC for browser/mobile voice (media tracks carry audio; data channel carries JSON events). Flow: browser creates SDP offer → your server
POST https://api.openai.com/v1/live/sessionswith{ session, transport: { type: "webrtc", sdp } }using the project API key (never in browser code) → server returns SDP answer → browser applies it → wait forsession.startedon the data channel. - WebSocket for server-side audio: connect
wss://api.openai.com/v1/live/sessions, sendsession.startas first message, wait forsession.started.
{
"type": "session.start",
"event_id": "event_start",
"session": {
"model": "gpt-live-1",
"instructions": "Be concise. Delegate requests needing current information to the backend, which can search the web.",
"audio": { "format": { "type": "audio/pcm", "rate": 24000 }, "output": { "voice": "marin" } },
"delegation": {
"type": "responses",
"responses": {
"model": "gpt-5.6-luna",
"tools": [{ "type": "web_search" }],
"tool_choice": "auto"
}
}
}
}
Key callouts verified against docs: model/voice/audio format are immutable after startup; instructions up to 16,384 tokens at start (append later via session.instructions.append); omit delegation (or "type":"client") for client delegation where your app runs the backend and feeds results back via session.commentary.append / session.thinking.append.
session.started (resolved config + session ID) → stream 24 kHz PCM16 mono input via session.input_audio.append (unacknowledged) → receive session.output_audio.delta (timed base64 PCM) → session.input_transcript.delta / session.output_transcript.delta (user/assistant transcript fragments) → session.delegation.created when the voice model hands work to the backend → response.event envelopes carrying nested Responses events (response.output_item.done, response.completed — note: double-completion reported by community, see step 5) → session.usage.updated (~once/min) → session.close → final usage in session.closed.
Total = (billable voice seconds ÷ 60 × $0.05) + backend costs. A 90-second session = $0.075 voice. Per-second billing, no rounding up; WebRTC creation bills 15s credited at start; idle sessions keep billing — close explicitly.
A buyer following this quickstart should expect to rewrite Realtime-API habits: response.done.output arrives empty (use response.output_item.done), response.completed can fire twice (inner + outer), no session.cancel/input_audio_buffer.clear equivalents, user-side transcript retrieval reported unavailable in some deployments, and tool-dependent end_conversation flows may never fire because the voice model answers social closers directly.
Recommended follow-up VERIFY protocol (once a paid key is available)
- Post a WebRTC session from a localhost page with microphone capture; confirm
session.startedand time-to-first-audio. - Interrupt mid-speech; confirm natural turn-taking and a spoken acknowledgement.
- Ask a current-events question with
gpt-5.6-luna+web_searchResponses delegation; confirm the grounded result arrives inside the conversation. - Run a 2-minute session; compare the invoice to the documented formula and confirm per-second (not per-minute) billing.
- Attempt user-side transcript retrieval (
session.input_transcript.deltaaccumulation / stored sessionGET /v1/live/sessions/{id}/content) to settle the community-reported discrepancy.
Lab verdict: INSPECT completed. Live VERIFY is gated on a paid OpenAI API key + billing (free tier unsupported) — outside this environment's capabilities. The protocol above makes the live extension a 15-minute task for any developer with a key.