News Weekly
LV 10 XP
0% read
S52model-release
#52 Issue #1Confirmed

OpenAI adds GPT-Live-1 real-time full-duplex VOICE model to API (discovery said "video" — CORRECTED)

On Thursday Sep 10, 2026, OpenAI announced and shipped, effective immediately, general availability of GPT-Live-1 in the API: The model: GPT-Live-1 is OpenAI's premier full-duplex real-time voice model ("Our premier model for natural, expressive voice conversations with smooth interruption handling" — model card). It reasons over incoming and outgoing audio together in a single model, avoiding the latency and brittle handoffs of chained STT→LLM→TTS pipelines. It can listen while speaking, respond to interruptions and short acknowledgements ("mhmm", "yeah"), and keep a conversation running while a backend model or agent performs reasoning, tool calls, web search, or coding. Delegation architecture (the defining feature): GPT-Live-1 handles conversation only; reasoning/tools live in a backend the developer chooses. Two modes: Responses delegation (OpenAI-managed — you configure a Responses model such as gpt-5.6-luna, gpt-5.6-terra or gpt-6-astra with tools like web_search or custom functions) and client delegation (your application builds the request and runs its own backend model, agent harness, or service; results are fed back via session.commentary.append / session.thinking.append). The launch post demonstrates pairing GPT-Live-1 with the Codex SDK for repository work mid-conversation. OpenAI's guidance is explicit: business logic, authorization checks, tool handlers and durable state stay in the application. Transports: WebRTC for browser/mobile voice sessions (microphone/speaker media tracks + data channel), WebSockets for server-side audio streams, sideband connections for server control/observation of a live session, and Telephony/SIP support for phone agents (restaurant reservations, customer support, order updates). New voice options: 12 new real-time voices (Quartz, Ripple, Vesper, Willow, Stone, Gleam, Meridian, Bossa, Tempo, Beacon, Delta, Cinder) spanning accents, dialects and languages, plus the existing default marin; custom voices via sales eligibility. Built-in speech plumbing: native ASR transcripts (session.input_transcript.delta) and response text (session.output_transcript.delta), strong alphanumeric understanding, keyword biasing, and native turn detection for developers who want explicit turn boundaries even though the model is not turn-based. Pricing & limits: voice layer $0.05/min billed per second (not rounded up; a 90-second session = $0.075); backend model/tool usage billed separately at normal rates; free API tier not supported; rate limits measured in concurrent sessions: 25 (Tier 1) → 50 (Tier 2) → 200 (Tier 3) → 300 (Tier 4) → 500 (Tier 5). Knowledge cutoff: July 31, 2025. Paired same-day launch: OpenAI also moved the Agents API into public beta the same day (Sep 10) — GPT-Live-1 is the voice that talks to a customer; the Agents API (managed Codex harness) is a backend it can hand the real work to. OpenAI additionally pointed enterprises at OpenAI Presence (introduced Jul 22) as the managed enterprise voice-agent product built on GPT-Live-1. Day-one ecosystem: Twilio announced a native GPTLiveProvider in Agent Connect connecting Twilio Voice Media Streams directly to GPT-Live-1 sessions (Sep 10). Customer testimonials in the launch post: Yelp (Host/Hatch call handling), Speak (Live Tutor Lessons — ~80% fewer interruptions during thinking pauses), Intercom Fin (natural call flow for support), Cognition (Devin voice collaboration).

Two luminous sound ribbons cross without breaking, pausing or queueing, one interruption absorbed instantly while a warm acknowledgement ripple continues beneath.
How do you want to read this?

Tailored emphasis while keeping the full article available.

Best for you · Decision maker

▥ Enterprise and strategic impact, risks, and the actions to take.

At a glance

The essential information in 30 seconds

What happened

On Thursday Sep 10, 2026, OpenAI announced and shipped, effective immediately, general availability of GPT-Live-1 in the API:

  1. The model: GPT-Live-1 is OpenAI's premier full-duplex real-time voice model ("Our premier model for natural, expressive voice conversations with smooth interruption handling" — model card). It reasons over incoming and outgoing audio together in a single model, avoiding the latency and brittle handoffs of chained STT→LLM→TTS pipelines. It can listen while speaking, respond to interruptions and short acknowledgements ("mhmm", "yeah"), and keep a conversation running while a backend model or agent performs reasoning, tool calls, web search, or coding.
  2. Delegation architecture (the defining feature): GPT-Live-1 handles conversation only; reasoning/tools live in a backend the developer chooses. Two modes: Responses delegation (OpenAI-managed — you configure a Responses model such as gpt-5.6-luna, gpt-5.6-terra or gpt-6-astra with tools like web_search or custom functions) and client delegation (your application builds the request and runs its own backend model, agent harness, or service; results are fed back via session.commentary.append / session.thinking.append). The launch post demonstrates pairing GPT-Live-1 with the Codex SDK for repository work mid-conversation. OpenAI's guidance is explicit: business logic, authorization checks, tool handlers and durable state stay in the application.
  3. Transports: WebRTC for browser/mobile voice sessions (microphone/speaker media tracks + data channel), WebSockets for server-side audio streams, sideband connections for server control/observation of a live session, and Telephony/SIP support for phone agents (restaurant reservations, customer support, order updates).
  4. New voice options: 12 new real-time voices (Quartz, Ripple, Vesper, Willow, Stone, Gleam, Meridian, Bossa, Tempo, Beacon, Delta, Cinder) spanning accents, dialects and languages, plus the existing default marin; custom voices via sales eligibility.
  5. Built-in speech plumbing: native ASR transcripts (session.input_transcript.delta) and response text (session.output_transcript.delta), strong alphanumeric understanding, keyword biasing, and native turn detection for developers who want explicit turn boundaries even though the model is not turn-based.
  6. Pricing & limits: voice layer $0.05/min billed per second (not rounded up; a 90-second session = $0.075); backend model/tool usage billed separately at normal rates; free API tier not supported; rate limits measured in concurrent sessions: 25 (Tier 1) → 50 (Tier 2) → 200 (Tier 3) → 300 (Tier 4) → 500 (Tier 5). Knowledge cutoff: July 31, 2025.
  7. Paired same-day launch: OpenAI also moved the Agents API into public beta the same day (Sep 10) — GPT-Live-1 is the voice that talks to a customer; the Agents API (managed Codex harness) is a backend it can hand the real work to. OpenAI additionally pointed enterprises at OpenAI Presence (introduced Jul 22) as the managed enterprise voice-agent product built on GPT-Live-1.
  8. Day-one ecosystem: Twilio announced a native GPTLiveProvider in Agent Connect connecting Twilio Voice Media Streams directly to GPT-Live-1 sessions (Sep 10). Customer testimonials in the launch post: Yelp (Host/Hatch call handling), Speak (Live Tutor Lessons — ~80% fewer interruptions during thinking pauses), Intercom Fin (natural call flow for support), Cognition (Devin voice collaboration).

Context: GPT-Live-1 first shipped as the ChatGPT Voice model on July 8, 2026 (paid plans; GPT-Live-1 mini for Free). The API release on Sep 10 is the developer-surface debut of the same model family, 64 days later. OpenAI claims, in its own evaluations, a +30 percentage-point improvement on Full Duplex Bench over GPT-Realtime-2.1 and a #1 Tau3 finish when paired with GPT-6 Astra at medium reasoning effort (both COMPANY CLAIM).

Why it matters
  • Real-time voice is now a commodity API building block: $0.05/min per-second billing for ChatGPT-Voice-grade full-duplex conversation makes "an AI that can hold a phone conversation" a routine engineering cost, not a research project — the same commoditization pattern as text tokens in 2023–24.
  • The voice/brain split is an architectural statement: GPT-Live-1 deliberately separates conversation from reasoning. OpenAI is telling developers to stop treating voice as a model feature and start treating it as a layer you can swap while keeping your tools, guardrails and state machine. This is a different design center than the Realtime API (one model does all) and a direct challenge to incumbent speech stacks (Twilio-style orchestration with STT/LLM/TTS chains).
  • It lands inside the week's live-voice platform race: same window as Google's Gemini 3.8 Live (Sep 15) and xAI's Voice Agent API at $0.08/min (Sep 17, Galaxy Day 3). OpenAI's $0.05/min voice layer sets the price reference point; Google answers with assistant-default distribution, xAI undercuts on telephony-agent positioning.
  • Pairing with the Agents API (same day) and Presence completes the enterprise story: OpenAI now sells (a) a managed enterprise voice agent (Presence), (b) the raw voice model (GPT-Live-1 API), and (c) a managed reasoning harness (Agents API beta) — a vertical stack for customer-facing voice automation that competes with telecom/contact-center incumbents (Twilio, Five9, NICE) even as Twilio integrates it.
  • It corrects the record on "multimodal voice": the GPT-Live-1 API accepts no images and no video (audio+text only). Anyone building "see the world" agents must continue using the Realtime API (image input) or other multimodal models — a meaningful constraint for the week's "live-video agent" narrative that the discovery record incorrectly attributed to GPT-Live-1.
Evidence

CONFIRMED

23 sources · 84 min read
Story identity
  • What: On Thursday 2026-09-10, OpenAI made GPT-Live-1 — the full-duplex real-time voice model that has powered ChatGPT Voice since July 8, 2026 — generally available in the API through a new Live sessions surface (v1/live/sessions), priced at $0.05 per minute for the front-end voice layer (billed per second), with backend model/tool usage billed separately. GPT-Live-1 listens and speaks simultaneously, handles interruptions and background noise, and delegates reasoning and tool use to a separate backend model or agent (OpenAI Responses model or customer's own backend) so the conversation continues while work happens in the background.
  • ⚠ LOUD FLAG — discovery framing correction (not a date mismatch): The discovery record titles this story "OpenAI adds GPT-Live-1 real-time video model to API" and describes GPT-Live-1 as a "real-time multimodal (video+voice+text) model." Both characterizations are factually incorrect. GPT-Live-1 is a voice model whose only modalities are audio and text; image and video are explicitly unsupported in the API. Primary-source proof:
    1. OpenAI's official model page (https://developers.openai.com/api/docs/models/gpt-live-1): "GPT-Live 1 is a full-duplex voice model for real-time conversations… Input modalities: audio, text — Output modalities: audio, text — Unsupported modalities: image, video" and "Videos v1/videos — Not supported."
    2. OpenAI's own launch post (Sep 10): "a powerful, natural voice model for building voice-enabled apps" — every capability listed (interruption handling, delegation, tone/pace/style, telephony) is voice-specific.
    3. OpenAI Help Center (ChatGPT Voice FAQ): "Live does not support video or screen sharing" — video/screen-sharing remains exclusive to the older Advanced Voice Mode in ChatGPT consumer apps.
    4. Independent confirmation: Coursiv's Sep 11 technical review ("Inputs and outputs: Audio and text; images and video are not supported"), AI Pricing Guru ("It does not accept images"), The Decoder ("The speech model can listen and talk at the same time"), Unite.AI, TechRepublic, GIGAZINE — all describe a voice/speech model, none mention video.
    • The "video" framing appears to be a discovery-stage conflation, most plausibly with (a) the Realtime API's gpt-realtime-2.1, which does accept image input, (b) GPT-6 Astra's multimodal/video-adjacent capabilities released Sep 3, or (c) Advanced Voice Mode's video/screen-sharing, which GPT-Live explicitly does not include. Recommended orchestrator action: rename the story to reflect reality (e.g., "OpenAI adds GPT-Live-1 real-time full-duplex voice model to API") and correct the what_changed field: "OpenAI made GPT-Live-1, its real-time full-duplex voice model (audio+text), available through the API to all developers." The event date, organization, category and in-window eligibility all stand.
  • Event-date verification trail (primary sources): OpenAI launch post datelined "September 10, 2026" (https://openai.com/index/introducing-gpt-live-1-in-the-api/); API changelog entry "Sep 10 — Feature · Model: gpt-live-1 · API: v1/live/sessions — GPT-Live 1 is now generally available in the API" (https://developers.openai.com/api/docs/changelog); OpenAI Developers X post "GPT-Live-1 is now available in the API…" timestamped Sep 10, 2026; official developer-forum announcement posted 2026-09-10T18:57:29Z. Independent anchors on the same day: The Decoder ("Sep 10, 2026"), Unite.AI ("Published September 10, 2026"), Twilio ("September 10, 2026"). Verdict: event date 2026-09-10 falls inside the configured window — no mismatch. The correction needed is to the story's substance (video), not its date.
  • Sourcing caveats to flag:
    1. The discovery record lists TechCrunch among independent sources; no TechCrunch article URL for either the July 8 launch or the Sep 10 API launch could be located/verified during this research session (aggregators confirm TechCrunch covered the July 8 consumer launch, but no direct URL was retrievable). Corroboration instead comes from The Decoder, Unite.AI, TechRepublic, Coursiv, Twilio, GIGAZINE, AI Pricing Guru, Progressive Robot and diyai.
    2. The Verge covered the consumer launch on July 8, 2026 (before the research window) — cited here only as independent evidence of what GPT-Live-1 is (a full-duplex voice model; no video/screen sharing at launch), not as in-window coverage.
    3. Benchmark figures cited both in the OpenAI post and third-party roundups are OpenAI's own evaluations (Full Duplex Bench, Tau3, Tau Banking Voice) — COMPANY CLAIM until independently replicated; Coursiv explicitly notes "none of the current numbers come from outside OpenAI."
  • Evidence labels used: FACT (event date, API GA, pricing, endpoint, modalities, unsupported modalities, voices, rate limits, knowledge cutoff, July 8 ChatGPT lineage, same-day Agents API beta, Twilio/Presence ecosystem moves), COMPANY CLAIM (benchmarks, Speak's ~80% interruption reduction, Yelp/Speak/Intercom/Cognition testimonials as published by OpenAI), INDEPENDENT EVIDENCE (multi-outlet agreement on headline facts; The Decoder, Unite.AI, TechRepublic, Twilio technical framing; Coursiv/AI Pricing Guru spec tables; Microsoft Learn Foundry page; community-reported production behavior), INTERPRETATION (voice/backend split as architectural shift; commoditization of full-duplex voice; competitive context vs Gemini 3.8 Live and xAI Voice Agent API), PREDICTION (mini variant in API, Realtime deprecation pressure, DevDay Sep 29 follow-ups, price competition at ~$0.05/min).
✓

What happened?

On Thursday Sep 10, 2026, OpenAI announced and shipped, effective immediately, general availability of GPT-Live-1 in the API:

  1. The model: GPT-Live-1 is OpenAI's premier full-duplex real-time voice model ("Our premier model for natural, expressive voice conversations with smooth interruption handling" — model card). It reasons over incoming and outgoing audio together in a single model, avoiding the latency and brittle handoffs of chained STT→LLM→TTS pipelines. It can listen while speaking, respond to interruptions and short acknowledgements ("mhmm", "yeah"), and keep a conversation running while a backend model or agent performs reasoning, tool calls, web search, or coding.
  2. Delegation architecture (the defining feature): GPT-Live-1 handles conversation only; reasoning/tools live in a backend the developer chooses. Two modes: Responses delegation (OpenAI-managed — you configure a Responses model such as gpt-5.6-luna, gpt-5.6-terra or gpt-6-astra with tools like web_search or custom functions) and client delegation (your application builds the request and runs its own backend model, agent harness, or service; results are fed back via session.commentary.append / session.thinking.append). The launch post demonstrates pairing GPT-Live-1 with the Codex SDK for repository work mid-conversation. OpenAI's guidance is explicit: business logic, authorization checks, tool handlers and durable state stay in the application.
  3. Transports: WebRTC for browser/mobile voice sessions (microphone/speaker media tracks + data channel), WebSockets for server-side audio streams, sideband connections for server control/observation of a live session, and Telephony/SIP support for phone agents (restaurant reservations, customer support, order updates).
  4. New voice options: 12 new real-time voices (Quartz, Ripple, Vesper, Willow, Stone, Gleam, Meridian, Bossa, Tempo, Beacon, Delta, Cinder) spanning accents, dialects and languages, plus the existing default marin; custom voices via sales eligibility.
  5. Built-in speech plumbing: native ASR transcripts (session.input_transcript.delta) and response text (session.output_transcript.delta), strong alphanumeric understanding, keyword biasing, and native turn detection for developers who want explicit turn boundaries even though the model is not turn-based.
  6. Pricing & limits: voice layer $0.05/min billed per second (not rounded up; a 90-second session = $0.075); backend model/tool usage billed separately at normal rates; free API tier not supported; rate limits measured in concurrent sessions: 25 (Tier 1) → 50 (Tier 2) → 200 (Tier 3) → 300 (Tier 4) → 500 (Tier 5). Knowledge cutoff: July 31, 2025.
  7. Paired same-day launch: OpenAI also moved the Agents API into public beta the same day (Sep 10) — GPT-Live-1 is the voice that talks to a customer; the Agents API (managed Codex harness) is a backend it can hand the real work to. OpenAI additionally pointed enterprises at OpenAI Presence (introduced Jul 22) as the managed enterprise voice-agent product built on GPT-Live-1.
  8. Day-one ecosystem: Twilio announced a native GPTLiveProvider in Agent Connect connecting Twilio Voice Media Streams directly to GPT-Live-1 sessions (Sep 10). Customer testimonials in the launch post: Yelp (Host/Hatch call handling), Speak (Live Tutor Lessons — ~80% fewer interruptions during thinking pauses), Intercom Fin (natural call flow for support), Cognition (Devin voice collaboration).

Context: GPT-Live-1 first shipped as the ChatGPT Voice model on July 8, 2026 (paid plans; GPT-Live-1 mini for Free). The API release on Sep 10 is the developer-surface debut of the same model family, 64 days later. OpenAI claims, in its own evaluations, a +30 percentage-point improvement on Full Duplex Bench over GPT-Realtime-2.1 and a #1 Tau3 finish when paired with GPT-6 Astra at medium reasoning effort (both COMPANY CLAIM).

Δ

What changed?

  • Before: Developers building real-time voice agents used the Realtime API (v1/realtime with gpt-realtime-2.1 / 2.1-mini, released Jul 6, 2026), a single all-in-one speech-to-speech model that handles audio, reasoning and tool selection inside one model. GPT-Live-1 existed only as a consumer ChatGPT feature (Jul 8) with plan-based hourly caps, no public model ID, no pricing, and a "notify me" signup for API access (as documented by multiple builders through mid-July).
  • Change: OpenAI made GPT-Live-1 generally available to all developers through the new v1/live/sessions endpoint at $0.05/min (voice layer, per second), with full-duplex conversation handled by a dedicated voice model and reasoning/tools delegated to a separately chosen backend (Responses delegation or client delegation). SynthID watermarking had already been added to supported GPT-Live audio on Jul 31, signaling API readiness.
  • After: Any developer with a paid OpenAI API key can now embed ChatGPT-Voice-grade full-duplex conversation in their own app, phone line or agent, choosing their own reasoning backend per task (cheap model for scheduling, frontier model for hard cases). The voice layer is commoditized at a legible price; the architecture splits "voice" from "brain," making each independently replaceable. Real-time voice moves from consumer novelty to a standard API building block — in the same week Google (Gemini 3.8 Live, Sep 15) and xAI (Voice Agent API at $0.08/min, Sep 17) shipped competing real-time voice surfaces.
↔

Before → Change → After

PhaseState
BeforeVoice agents built on the Realtime API: gpt-realtime-2.1 (Jul 6, 2026) does speech+reasoning+tools in one model; image input supported; per-token audio pricing ($32/$64 per 1M audio in/out). GPT-Live-1 locked inside ChatGPT Voice (Jul 8): Go/Plus/Pro defaults, Free uses mini, plan caps up to 15h/day, no API access ("notify me" form), no public model ID. Video/screen-sharing only in Advanced Voice Mode.
ChangeSep 10, 2026: GPT-Live-1 GA in the API on v1/live/sessions at $0.05/min voice (per-second billing); full-duplex voice layer + delegated backend (Responses or client); WebRTC/WebSocket/Telephony; 12 new voices; native transcripts, keyword biasing, turn detection; no images/video; no free tier; 25–500 concurrent sessions by tier. Paired with Agents API public beta; Twilio GPTLiveProvider same day.
AfterDevelopers can build production full-duplex phone/browser voice agents with a replaceable brain; voice-layer price ~$0.05/min invites volume use-cases (support, reservations, order status, tutoring); the frontend/backend split changes how voice apps are architected (prompt split, event-model rewrite, no manual turn control); competitive pressure on Realtime API migration; Google (Gemini 3.8 Live) and xAI (Voice Agent API $0.08/min) respond within the same week; OpenAI DevDay (Sep 29) expected to extend the surface (mini variant, video, more voices).
⚙

How it works

  • Single-model full duplex: GPT-Live-1 jointly models user input audio and generated output audio, so it can decide to keep talking, stop, acknowledge ("mhmm"), or pivot mid-stream without a voice-activity-detection state machine owned by the developer. This replaces the STT→LLM→TTS concatenation where each handoff adds latency and timing drift (OpenAI's framing, echoed by GIGAZINE's analysis).
  • Session lifecycle (v1/live/sessions): WebSocket clients connect to wss://api.openai.com/v1/live/sessions and send session.start (model, instructions, audio format, voice, delegation config); the server replies session.started. WebRTC sessions are created with POST /v1/live/sessions (server exchanges the browser's SDP offer for an answer; audio travels on media tracks, JSON events on a data channel). Sessions are fixed-config at startup: model, instructions, voice and audio format are immutable (instructions can be appended via session.instructions.append; delegation settings are sparse-updatable within a mode).
  • Delegation flow: GPT-Live-1 decides when a request needs the backend; it emits session.delegation.created with a delegation_id. In Responses delegation, it prepares a request to your configured Responses model (with tools such as web_search or functions), which runs and returns results through response.event envelopes; the voice model then communicates the result. In client delegation, your application assembles the backend call itself and returns content via session.commentary.append (content to speak, paraphrasable) or session.thinking.append (quiet context). Interrupting speech does not cancel backend work — the application controls permissions, task state and which results reach the conversation.
  • Context management: instructions up to 16,384 tokens at startup; automatic conversation summarization/compaction — when context exceeds 90%, GPT-Live-1 starts a replacement voice engine within the same session seeded with original instructions + up to 8,192 tokens of history (recent messages + summary). session.instructions.append / session.thinking.append / session.commentary.append inject content with acknowledged start_ms/end_ms windows on the session timeline.
  • Transcripts: session.input_transcript.delta and session.output_transcript.delta deliver timed fragments for UI, logging, moderation and early-start work; the API also reports cumulative session.usage.updated (~once per minute; final usage at session.closed).
  • Audio formats: WebSocket PCM16 mono 24 kHz (audio/pcm) or G.711 μ-law 8 kHz (audio/pcmu) for direct telephony bridging (e.g., Twilio Media Streams); WebRTC negotiates its own format. Session durations billed per second and not rounded up; a WebRTC creation bills 15s of initialization that is credited once the session starts (cost-optimization guide).
  • Data controls: GPT-Live-1 is eligible for Zero Data Retention (with limitations; stored sessions ignored under ZDR), processes in US or EU/EEA+Switzerland regions (2026 guide data per Progressive Robot), and session recording (store:true) is off by default, retained 30 days when enabled.
!

Why it matters

▥ For Decision maker
  • Real-time voice is now a commodity API building block: $0.05/min per-second billing for ChatGPT-Voice-grade full-duplex conversation makes "an AI that can hold a phone conversation" a routine engineering cost, not a research project — the same commoditization pattern as text tokens in 2023–24.
  • The voice/brain split is an architectural statement: GPT-Live-1 deliberately separates conversation from reasoning. OpenAI is telling developers to stop treating voice as a model feature and start treating it as a layer you can swap while keeping your tools, guardrails and state machine. This is a different design center than the Realtime API (one model does all) and a direct challenge to incumbent speech stacks (Twilio-style orchestration with STT/LLM/TTS chains).
  • It lands inside the week's live-voice platform race: same window as Google's Gemini 3.8 Live (Sep 15) and xAI's Voice Agent API at $0.08/min (Sep 17, Galaxy Day 3). OpenAI's $0.05/min voice layer sets the price reference point; Google answers with assistant-default distribution, xAI undercuts on telephony-agent positioning.
  • Pairing with the Agents API (same day) and Presence completes the enterprise story: OpenAI now sells (a) a managed enterprise voice agent (Presence), (b) the raw voice model (GPT-Live-1 API), and (c) a managed reasoning harness (Agents API beta) — a vertical stack for customer-facing voice automation that competes with telecom/contact-center incumbents (Twilio, Five9, NICE) even as Twilio integrates it.
  • It corrects the record on "multimodal voice": the GPT-Live-1 API accepts no images and no video (audio+text only). Anyone building "see the world" agents must continue using the Realtime API (image input) or other multimodal models — a meaningful constraint for the week's "live-video agent" narrative that the discovery record incorrectly attributed to GPT-Live-1.
✦

What became possible?

  • Full-duplex phone agents in a weekend: reservations, order status, customer support and outbound calls where the caller can interrupt, change direction and talk over the agent — previously a multi-vendor integration project (STT+VAD+LLM+TTS+telephony), now a session.start call plus a Twilio Media Stream bridge (Twilio shipped the native GPTLiveProvider the same day).
  • Bring-your-own-brain voice: pair GPT-Live-1 with Luna for cheap high-volume scheduling, Terra for balanced work, Astra for hard reasoning, or a third-party/self-hosted model via client delegation — voice quality is decoupled from reasoning capability, so teams can swap backends without touching the voice layer.
  • Voice-enabled coding and ambient agents: the launch post demonstrates GPT-Live-1 + Codex SDK (talk to a repo-reading agent), and the cost guide sketches ambient agents that close the voice session while a long task runs and resume on completion ("Resume conversation" button).
  • Natural-language telephony UX: barge-in handling, "mhmm" backchannels, silence tolerance, and keyword-biased alphanumeric capture (order IDs, dates) — the interaction grammar of real phone calls rather than IVR menus.
  • Language/voice localization: 12 accent/dialect/language voices at launch with more promised, letting products localize voice agents as a configuration change rather than a TTS vendor swap.
◎

Implications

▥ For Decision maker

Technical

  • Two live-voice API paradigms now coexist: v1/realtime (single-model speech+reasoning+tools; image input supported) versus v1/live/sessions (dedicated voice layer + delegated backend; no image/video). Teams must pick deliberately; OpenAI's migration guide is explicit that moving Realtime→Live is not a model-name swap: prompts split (conversation style → session instructions; business rules/tool schemas → delegation instructions), tools move to the delegation schema, manual turn control is removed (audio streams continuously; the model decides when to speak), and audio/transcript events are renamed onto the session.
  • Event-model rewrite: session.input_audio.append (unacknowledged), session.output_audio.delta (timed), transcript deltas, session.delegation.created, response.event envelopes wrapping nested Responses events, session.usage.updated snapshots — a materially different protocol surface than the Realtime API's response.create/input_audio_buffer model. Community developers report real migration friction (see §13).
  • Latency/turn-taking gains are real but vendor-measured: OpenAI reports 0.798s turn-taking latency (Full Duplex Bench v1) vs 1.41s for gpt-realtime-2.1 and +30pp on Full Duplex Bench v3 — submitted only from OpenAI's own evals; no independent benchmark existed as of Sep 11 (Coursiv).
  • Concurrency-based rate limiting changes capacity planning: 25–500 concurrent sessions by tier means production voice apps plan around simultaneous calls, not tokens — a different operational model (and a cap that tier-1 teams will hit quickly in telephony rollouts).
  • Context compaction is now a first-class session feature: automatic >90% summarization with a same-session engine replacement changes long-conversation assumptions versus Realtime.

Developer

  • New SDK surface: client.live.connect() / client.live.create() in official SDKs (Node, Python) with WebRTC and WebSocket flows; a WebRTC quickstart (browser+server) and a WebSocket quickstart (PCM16 24kHz stdin/stdout) are the canonical entry points. Python needs openai[realtime]; WebRTC needs Node ≥ 22.6.
  • Architecture guidance is unusually opinionated: keep business logic, authorization, tool handlers and durable state in your application; never copy a Realtime prompt wholesale; treat the voice layer as replaceable. Client delegation is the escape hatch for custom routing, third-party backends, and result redaction before the voice model speaks.
  • Cost modeling changes: voice = duration-based ($0.05/min), backend = token-based at its own rate. A 10-minute call is $0.50 voice + backend + tools. Per-second billing (no rounding) plus explicit session.close discipline (idle sessions bill continuously) and the 15s WebRTC-init credit are the key optimizations. Teams should estimate cost-per-successful-task, not per-minute.
  • Community-reported gotchas (EARLY RESEARCH — OpenAI forum, Sep 18, several builders):
    • User-side audio transcription is not retrievable via the API in some deployments — the recorded transcript contains only assistant text, not the caller's words (the model card says transcripts are provided natively, but at least one team could not retrieve user speech; discrepancy unresolved — treat with caution).
    • response.done.output arrives empty in delegation mode; function calls come individually via response.output_item.done inside response.event envelopes.
    • Double response.completed events (inner delegation-level + outer session-level) can double-advance state machines.
    • No session.update / response.cancel / input_audio_buffer.clear equivalents over the data channel as of Sep 18 — goodbye flows and interruption handling need new patterns; tool-dependent end_conversation flows often never fire because the voice model answers social closers directly.
    • Voice-model commentary can precede session.delegation.created, breaking naive "voice spoke ⇒ delegation" turn detection.
  • Migration economics: for teams already on Realtime, immediate benefits are per-second billing, better tool calling, interruptions and speed (per community testing); blocking issues are transcript retrieval and event-model rewrites. Several teams advise running both stacks during migration.

Enterprise

  • Customer-facing voice automation is now buyable at API prices: contact-center-like inbound/outbound voice (reservations, order status, support triage) becomes an internal build on $0.05/min voice + a reasoning backend, or a procurement decision via OpenAI Presence (managed enterprise voice agents, launched Jul 22, limited GA via OpenAI Forward Deployed Engineers and systems integrators).
  • Day-one enterprise validation: Yelp (Host/Hatch call handling), Intercom (Fin), Speak (tutoring), Cognition (Devin) all ship on GPT-Live-1; Twilio's native Agent Connect integration lowers last-mile telephony risk for the enterprise channel.
  • Data-control posture matters for regulated sectors: GPT-Live-1 is ZDR-eligible with limitations and supports EU/EEA+Switzerland regional processing, but a delegated Agents API backend is US-only and not ZDR-eligible — a European voice-agent deployment loses end-to-end ZDR the moment it hands work to the Agents API. Stored sessions (store:true) persist 30 days with no public deletion endpoint. Compliance teams must map the voice layer and each backend separately.
  • Concurrency ceilings cap telephony scale: 500 concurrent sessions at Tier 5 bounds simultaneous calls — call-center peak-load planning changes from token budgets to session budgets; overflow strategies (queuing, regional sharding) are needed at scale.

Strategic

  • OpenAI is defining the "voice layer" as a product category — a thin, cheap, replaceable conversation front end over any reasoning backend. That framing pressures both ends of the market: speech incumbents (assembly-style pipelines, turn-based voice platforms) and frontier-model rivals who bundle voice with intelligence.
  • Pricing marker set at $0.05/min: xAI's $0.08/min Voice Agent API (Sep 17) and Google's Gemini 3.8 Live (Sep 15) will be benchmarked against it; expect voice-layer price compression toward $0.01–0.03/min within quarters, following the token-price pattern.
  • Ecosystem lock-in via delegation: optional today, "Responses delegation" makes OpenAI the orchestration layer between the voice and the brain; client delegation keeps customers open. The strategic prize is owning the glue (voice→backend protocol, presence, agents) rather than the voice model alone.
  • Same-week narrative: in a week dominated by agent-safety disclosures (S01, S02, S14) and the pacing debate (S15), OpenAI shipping a production full-duplex phone voice (plus the Agents API) is the "deploy now" counter-narrative — agentic voice is being commoditized while the industry argues about controlling agents.
  • Foundry distribution: Microsoft Foundry documents GPT-Live natively under Azure Foundry endpoints (learn.microsoft.com GPT-Live event reference), extending OpenAI's voice layer into enterprise Azure channels alongside the GPT-6 Astra GA there (S06).
⚠

Risks & limitations

▥ For Decision maker
Risks
  • Untested-at-scale voice commerce: consumer ChatGPT Voice is proven at 150M+ users (per OpenAI's July claims), but enterprise phone-call reliability (noise, accents, call quality, DTMF, hold music) is new territory; early community reports of hangup/delegation unreliability in semi-rigid flows are a yellow flag for production telephony.
  • Data/privacy edge cases: 30-day stored-session retention with no public deletion endpoint; ZDR gaps when delegation crosses to Agents API; user-speech transcript handling uncertainty (see §13) matters for call-recording consent regimes (two-party consent states, GDPR).
  • Migration traps: teams porting Realtime code naively will hit empty response.done.output, double response.completed, and missing session controls — production incidents during migration are likely.
  • Concurrency throttling: tier caps (25 at Tier 1) will surprise early telephony deployers; silent session-idle billing can inflate invoices; per-second metering makes idle-session hygiene a cost control.
  • Voice impersonation/safety: expanded voice roster (accents/dialects/languages) plus streamable cloning-adjacent custom voices (sales-gated) re-raises voice-spoofing and social-engineering risk; OpenAI's SynthID watermarking (Jul 31) mitigates but does not eliminate it.
  • Benchmark overclaiming: +30pp Full Duplex Bench, Tau3 #1 and 0.798s latency are OpenAI's own evals; blind replication is pending. Procurement teams should demand vendor-neutral evaluation (per Coursiv's warning).
  • Framing risk to the newsroom: the discovery record's "video model" error, if propagated, would mislead readers about both GPT-Live-1's capabilities and the state of OpenAI's real-time video API; this report corrects it (see §1).
Limitations
  • No image or video input/output — audio+text only; the "see the world" real-time agent use case is NOT served by this model (contrary to discovery). Realtime API gpt-realtime-2.1 supports image input; GPT-Live does not.
  • Endpoint confinement: v1/live/sessions only — not callable via Chat Completions, Responses, Realtime, Assistants or Batch; no fine-tuning, structured outputs, or predicted outputs; no temperature/top_p controls (fixed).
  • Voice fixed per session (immutable after startup; new session to change), output audio PCM16 24kHz mono (no stereo), and input transcript retrieval uncertainty for user audio (community reports conflict with model-card promises; unresolved as of Sep 18).
  • No free tier — every evaluation costs money; no public mini variant in the API (GPT-Live-1 mini exists in ChatGPT consumer; model page for the API mini 404s as of Sep 11, per Progressive Robot).
  • Knowledge cutoff July 31, 2025 — current-events answers depend entirely on backend delegation + web search.
  • Regional/data constraints: EEA processing only via regional domains; UK listed for regional storage only (2026 data-controls documentation); ZDR not end-to-end when delegating to Agents API.
  • All headline benchmarks are vendor-measured (OpenAI evals; backend-specific — Terra low vs Astra medium materially changes scores).
?

Open questions

▥ For Decision maker
  • Will GPT-Live-1 mini (and Medium/High reasoning variants) ship in the API, at what price, and when? (DevDay Sep 29 is the near-term candidate.)
  • Will the Realtime API line get a deprecation date now that GPT-Live-1 is GA — forcing migration for gpt-realtime-2.1 users?
  • Can user-side speech transcripts actually be retrieved via the API (discrepancy between model-card claims and community testing)? OpenAI's clarification would settle a production blocker.
  • When will video and screen-sharing arrive for GPT-Live (OpenAI said "soon" at the July launch; still absent Sep 10)? Its arrival would finally justify the discovery record's "video" framing — retroactively wrong, but directionally predictive.
  • Will OpenAI publish independent third-party full-duplex benchmarks (or open Full Duplex Bench/Tau3 methodology) to back the +30pp and #1 claims?
  • How will voice-layer pricing evolve once Google (Gemini 3.8 Live) and xAI ($0.08/min) publish comparable benchmark numbers?
  • Does OpenAI Presence become the enterprise distribution layer that sidelines self-serve GPT-Live-1 API deployments, or do both scale?
↗

What happens next?

  • Near term (days–weeks): developer community pressure on OpenAI to (a) expose user-side transcripts, (b) add session.cancel-like controls, (c) confirm Realtime deprecation timing; expect documentation updates and possible API tweaks before OpenAI DevDay on Sep 29 (Fort Mason, San Francisco) — the natural venue for GPT-Live-1 mini, video support announcements, and benchmark methodology releases.
  • This week's competitive clock: Google's Gemini 3.8 Live (Sep 15) and xAI's Voice Agent API (Sep 17) give buyers three full-duplex voice surfaces within 7 days; benchmark shootouts (Artificial Analysis-style) are imminent.
  • Data/legal follow-ups: OpenAI published a GPT-Live system card (with a corrected/re-run safety evaluation noted Aug 4); voice-agent deployments will start showing up in EU AI Act GPAI transparency reporting and GDPR audio-processing audits.
  • Product trajectory: expect voice-layer price compression, a mini-tier for low-cost telephony, video/screen-sharing parity with Advanced Voice Mode ("working to introduce these capabilities soon" per July launch materials), deeper Codex/Agents-API integration, and Foundry/enterprise adoption as the GPT-6 Astra GA (Sep 17, S06) normalizes frontier models on Azure.
★

Editorial takeaway

▥ For Decision maker

Report this story as "OpenAI opens its full-duplex real-time voice model to all developers — GPT-Live-1 hits the API at $0.05/minute" — NOT as a video model. The correction matters: GPT-Live-1 accepts audio and text only, and OpenAI is explicit that image and video are unsupported, so the "real-time video agent" angle belongs to other models (Realtime with image input, or future GPT-Live updates). The genuinely newsworthy substance is (1) ChatGPT-Voice-grade full-duplex conversation as a commodity API at a legible $0.05/min per-second price; (2) the architectural bet that voice and reasoning should be separable layers — a reframing of how voice agents are built, with real migration costs for Realtime users; (3) the same-week collision with Google's Gemini 3.8 Live and xAI's $0.08/min Voice Agent API, making this the week real-time voice became a platform war; and (4) the pairing with the Agents API beta and OpenAI Presence, completing OpenAI's enterprise voice stack. Flag the benchmark claims as vendor-measured, and surface the early community-reported production gaps (transcript retrieval, hangup reliability) honestly — they are the difference between a demo and a phone line.

A slim conversation-only front module passes work to a larger separate reasoning department via a narrow return channel, with business rules held in the application's own locked box.
⌘

Lab: INSPECT

Step 1 — Choose transport
  • WebRTC for browser/mobile voice (media tracks carry audio; data channel carries JSON events). Flow: browser creates SDP offer → your server POST https://api.openai.com/v1/live/sessions with { session, transport: { type: "webrtc", sdp } } using the project API key (never in browser code) → server returns SDP answer → browser applies it → wait for session.started on the data channel.
  • WebSocket for server-side audio: connect wss://api.openai.com/v1/live/sessions, send session.start as first message, wait for session.started.
Step 2 — Minimal session config (WebSocket, from the official quickstart)
{
  "type": "session.start",
  "event_id": "event_start",
  "session": {
    "model": "gpt-live-1",
    "instructions": "Be concise. Delegate requests needing current information to the backend, which can search the web.",
    "audio": { "format": { "type": "audio/pcm", "rate": 24000 }, "output": { "voice": "marin" } },
    "delegation": {
      "type": "responses",
      "responses": {
        "model": "gpt-5.6-luna",
        "tools": [{ "type": "web_search" }],
        "tool_choice": "auto"
      }
    }
  }
}

Key callouts verified against docs: model/voice/audio format are immutable after startup; instructions up to 16,384 tokens at start (append later via session.instructions.append); omit delegation (or "type":"client") for client delegation where your app runs the backend and feeds results back via session.commentary.append / session.thinking.append.

Step 3 — Session lifecycle events to observe (from the reference)

session.started (resolved config + session ID) → stream 24 kHz PCM16 mono input via session.input_audio.append (unacknowledged) → receive session.output_audio.delta (timed base64 PCM) → session.input_transcript.delta / session.output_transcript.delta (user/assistant transcript fragments) → session.delegation.created when the voice model hands work to the backend → response.event envelopes carrying nested Responses events (response.output_item.done, response.completed — note: double-completion reported by community, see step 5) → session.usage.updated (~once/min) → session.close → final usage in session.closed.

Step 4 — Cost math (verified against the cost-optimization guide)

Total = (billable voice seconds ÷ 60 × $0.05) + backend costs. A 90-second session = $0.075 voice. Per-second billing, no rounding up; WebRTC creation bills 15s credited at start; idle sessions keep billing — close explicitly.

Step 5 — Production-risk walkthrough (from community thread, Sep 18 — EARLY RESEARCH)

A buyer following this quickstart should expect to rewrite Realtime-API habits: response.done.output arrives empty (use response.output_item.done), response.completed can fire twice (inner + outer), no session.cancel/input_audio_buffer.clear equivalents, user-side transcript retrieval reported unavailable in some deployments, and tool-dependent end_conversation flows may never fire because the voice model answers social closers directly.

Recommended follow-up VERIFY protocol (once a paid key is available)

  1. Post a WebRTC session from a localhost page with microphone capture; confirm session.started and time-to-first-audio.
  2. Interrupt mid-speech; confirm natural turn-taking and a spoken acknowledgement.
  3. Ask a current-events question with gpt-5.6-luna + web_search Responses delegation; confirm the grounded result arrives inside the conversation.
  4. Run a 2-minute session; compare the invoice to the documented formula and confirm per-second (not per-minute) billing.
  5. Attempt user-side transcript retrieval (session.input_transcript.delta accumulation / stored session GET /v1/live/sessions/{id}/content) to settle the community-reported discrepancy.

Lab verdict: INSPECT completed. Live VERIFY is gated on a paid OpenAI API key + billing (free tier unsupported) — outside this environment's capabilities. The protocol above makes the live extension a 15-minute task for any developer with a key.

≡

Research sources

Primary Sources (11)
Primary
ChatGPT — Release Notes | OpenAI Help CenterGPT-Live-1 launched in ChatGPT Voice Jul 8, 2026 (paid users; mini for Free); "GPT-Live-1 does not support video or screen sharing at this time"; works with web search/memory/text+images within a ChatGPT chat; rolling plan caps (Go/Plus 3h, Pro $100 15h, Pro $200 unlimited with GPT-Live-1). — Primary (release notes) / FACT (July 8 consumer lineage; no video).Date: entries for Jul 8, 2026 (GPT-Live-1 introduction) and Sep 9/10, 2026 (usage-limit updates)
Visit source ↗
Primary
ChatGPT Voice | OpenAI Help Center (GPT-Live FAQ)Consumer-product proof that GPT-Live is a voice experience and "**does not support video or screen sharing**" (Advanced Voice Mode retains those); GPT-Live-1 on paid plans, mini on Free; Live works with text and images in the same chat (ChatGPT-side widget behavior, not API video input). — Primary (OpenAI documentation) / FACT (no video in GPT-Live — supports the §1 correction).Date: accessed 2026-09-19 (mirrors in-window state)
Visit source ↗
Primary
Introducing GPT-Live-1 in the API — OpenAI Developer Community (announcement thread)Official developer-forum announcement timestamp (Sep 10, 2026); feature list (interruption handling, voice customization, built-in transcripts, keyword biasing, turn detection, WebRTC/WebSocket/telephony); $0.05/min pricing; Responses vs client delegation; "interrupting speech does not automatically cancel backend work." — Primary (announcement) / FACT (date, feature list).Date: 2026-09-10T18:57:29Z
Visit source ↗
Primary
Cost optimization (voice latency and cost) — OpenAI APIBilling model (per-second, no rounding; 90-second session = $0.075 at $0.05/min); WebRTC creation bills 15s credited at start; idle-session discipline; backend billed separately; ambient-agent pattern (close voice session during long backend tasks, resume on completion); total-cost formula and per-session usage accounting (do not sum snapshots). — Primary (documentation) / FACT (cost model for §9/§19).Date: accessed 2026-09-19
Visit source ↗
Primary
WebRTC (voice) — OpenAI APIBrowser flow (SDP offer → server POST /v1/live/sessions with project API key → SDP answer → session.started on data channel); example with Responses delegation gpt-5.6-terra + hosted web search; ephemeral-key guidance for Realtime; key stays on trusted server. — Primary (documentation) / FACT (protocol surface for browser voice).Date: accessed 2026-09-19
Visit source ↗
Primary
WebSockets (voice) — OpenAI APIServer-side connection flow (wss://api.openai.com/v1/live/sessions; session.start → session.started); Node/Python examples with model "gpt-live-1", 24 kHz PCM16 audio format, voice "marin", Responses delegation with gpt-5.6-luna + web_search; sideband connection concept. — Primary (documentation) / FACT (code walkthrough basis for INSPECT lab).Date: accessed 2026-09-19
Visit source ↗
Primary
Managing GPT-Live sessions | OpenAI APISession-config details (model/voice/instructions immutability, 16,384-token instructions, session.instructions.append / thinking.append / commentary.append semantics with start_ms/end_ms windows); transcript delta events; automatic background summarization at >90% context with replacement voice engine (original instructions + up to 8,192 tokens history); session forking/stored sessions; usage reporting. — Primary (documentation) / FACT (session semantics).Date: accessed 2026-09-19
Visit source ↗
Primary
Getting started with GPT-Live | OpenAI APIArchitecture explanation (GPT-Live handles conversation; backend handles delegated work); full-duplex definition; Responses delegation vs client delegation; WebRTC/WebSocket/telephony transport split; quickstart procedure (browser microphone + server API key; session.started verification). — Primary (documentation) / FACT (protocol surface).Date: accessed 2026-09-19
Visit source ↗
Primary
Changelog — OpenAI APISecond in-window datestamp: "Sep 10 — Feature · Model: gpt-live-1 · API: v1/live/sessions — GPT-Live 1 is now generally available in the API. Build full-duplex voice conversations…"; $0.05/min per-second voice pricing; Responses vs client delegation; same-day Sep 10 entries (Agents API public beta; project API key expiration). — Primary / FACT (event date, GA status, pricing).Date: entry dated "Sep 10" (2026)
Visit source ↗
Primary
GPT-Live 1 Model | OpenAI API (model page)"GPT-Live 1 is a full-duplex voice model for real-time conversations"; input modalities audio+text; output modalities audio+text; **unsupported modalities: image, video**; knowledge cutoff Jul 31, 2025; $0.05/min billed per second (not rounded); endpoint matrix (only v1/live/sessions; Videos v1/videos NOT supported; Chat Completions/Responses/Realtime/Assistants/Batch/Fine-tuning all Not supported); supported features streaming + function_calling; unsupported structured_outputs/fine_tuning/predicted_outputs; rate limits in concurrent sessions 25/50/200/300/500 (Tiers 1–5); **free tier unsupported**. — Primary / FACT (definitive on modalities — audio+text only; video/image explicitly unsupported).Date: accessed 2026-09-19 (page current as of research)
Visit source ↗
Primary
Build more natural voice experiences with GPT-Live-1 in the API — OpenAI launch postTHE event-date anchor; GPT-Live-1 characterization as a full-duplex VOICE model ("natural voice model… listening and speaking at the same time"); delegation to backend models (Luna/Terra/Astra, third-party); interruption handling; tone/pace/style via system prompt; telephony support; 12 new voices (Quartz, Ripple, Vesper, Willow, Stone, Gleam, Meridian, Bossa, Tempo, Beacon, Delta, Cinder); $0.05/min voice layer; benchmark claims (+30pp Full Duplex Bench vs gpt-realtime-2.1; #1 Tau3 with GPT-6 Astra at medium reasoning); customer testimonials (Yelp/Speak/Intercom/Cognition); OpenAI Presence tie-in; Codex pairing example. **No mention of video — the primary proof that the discovery record's "video model" framing is wrong.** — Primary / FACT (event, pricing, capabilities), COMPANY CLAIM (benchmarks, testimonials).Date: 2026-09-10 (dateline "September 10, 2026")
Visit source ↗
Independent Sources (10)
Independent
OpenAI releases API for its voice conversation AI 'GPT-Live-1' — GIGAZINEIn-window international independent coverage: full-duplex method; separation of conversation layer from heavy-task layer (delegation to frontier models e.g. GPT-6 Astra); STT→LLM→TTS replacement rationale; telephony use cases; 12 voices; $0.05/min (~¥7.7/min); OpenAI Developers X post timestamped September 10, 2026; benchmark claims as OpenAI-reported. — Independent / CONFIRMED (event date, architecture summary), COMPANY CLAIM relay (benchmarks).Date: 2026-09-11 (03:21 UTC)
Visit source ↗
Independent
GPT-Live event API reference — Microsoft Learn (Microsoft Foundry)GPT-Live sessions available via Microsoft Foundry endpoints (/openai/v1/live); full event reference (session.start/started/updated, output_audio.delta, transcript deltas, delegation.created, response.event envelopes, usage.updated, closed) — independent-of-openai.com documentation of the same session protocol, evidencing Foundry/Azure distribution of the voice layer. — Independent (platform documentation) / CONFIRMED (protocol surface; Foundry availability).Date: accessed 2026-09-19 (current)
Visit source ↗
Independent
GPT-Live-1 API and Agents API Beta: Essential Risks to Avoid — Progressive RobotData-control surface specific to GPT-Live-1 (ZDR eligible with limitations; processing regions US and EU/EEA+Switzerland; UK regional storage only; store:true 30-day session retention with no public deletion endpoint; Agents API backend US-only and NOT ZDR-eligible); timeline (Jul 8 ChatGPT launch → Jul 22 Presence → Jul 31 SynthID → Aug 4 system card re-run → Sep 10 API + Agents API beta → Sep 29 DevDay); open questions (mini in API, Realtime deprecation, ZDR for agents). — Independent / CONFIRMED (data/controls facts), EARLY RESEARCH signals (open questions).Date: 2026-09-11
Visit source ↗
Independent
GPT-Live-1 API Launch — Voice Pricing & Cost Guide — AI Pricing GuruIndependent spec/price verification (Sep 11, 10:32 UTC): $0.05/min per-second; backend billed separately; no free tier; concurrent-session limits 25 (Tier 1) → 500 (Tier 5); **does not accept images**; Live vs Realtime comparison (GPT-Live-1: no image input; gpt-realtime-2.1: image input supported); telephony support; 12-voice expansion; custom voices via sales eligibility. — Independent / CONFIRMED (pricing, limits, no-image modality).Date: 2026-09-10 (published 17:35 UTC; verified/updated 2026-09-11T10:35Z)
Visit source ↗
Independent
Build Voice AI Experiences with Twilio and GPT-Live-1 in the OpenAI API — Twilio (Lenore Files et al.)Day-one ecosystem integration (native GPTLiveProvider in Twilio Agent Connect connecting Twilio Voice Media Streams to GPT-Live-1 with no custom WebSocket plumbing); telephony last-mile positioning (production voice AI over Twilio's carrier network); twin tutorial links. Independent (third-party vendor) confirmation of the Sep 10 launch date. — Independent / CONFIRMED (date, integration existence).Date: 2026-09-10
Visit source ↗
Independent
OpenAI GPT-Live-1 Lets AI Listen and Speak at the Same Time — TechRepublic (Joseph Ofonagoro)Later-window independent coverage: full-duplex voice for real-time agents; interruption/pause/direction-change handling; tool/backend delegation; $0.05/min with separate backend charges; positions voice as an interface layer for broader AI agents (enterprise workflows). — Independent / CONFIRMED (event, price, framing).Date: 2026-09-15 (in-window follow-up)
Visit source ↗
Independent
GPT-Live-1 API: Pricing, Full-Duplex Voice, Benchmarks — Coursiv BlogDeepest independent technical review: spec table (audio+text I/O; **images and video not supported**; Live sessions endpoint only; streaming + function calling; no structured outputs/fine-tuning/predicted outputs; rate limits 25→500 concurrent sessions; free tier unsupported; knowledge cutoff Jul 31 2025; 12 voices); Responses vs client delegation mechanics; migration guidance detail (prompt split, tools move to delegation schema, manual turn control removed); benchmark table (Full Duplex Bench v3 tool-calling Pass@1 87% vs 60% for gpt-realtime-2.1, Terra backend low; turn-taking latency 0.798s vs 1.41s; Tau Banking Voice 32% vs 12.4%) with the explicit caveat that **all figures are OpenAI's own evaluations** — "none of the current numbers come from outside OpenAI"; July 8 ChatGPT-launch lineage; watch items (mini variant, Realtime deprecation, real bills). — Independent / CONFIRMED (specs, pricing, no-video modality), COMPANY CLAIM relay (benchmarks, flagged as vendor-measured).Date: 2026-09-11
Visit source ↗
Independent
OpenAI's GPT-Live-1 Arrives in the API at $0.05 Per Minute — Unite.AI (Jonas Reeve)In-window independent confirmation (published September 10, 2026): full-duplex voice model at $0.05/min voice layer; delegation; silent context management/background-noise handling; telephony support; 12 voices; OpenAI Presence context (launched Jul 22); SynthID watermarking added Jul 31; Codex pairing excerpt; no mention of video (consistent with voice-only). — Independent / CONFIRMED (event, pricing, feature set).Date: 2026-09-10
Visit source ↗
Independent
OpenAI's GPT-Live-1 API lets developers build apps that talk and listen at the same time — The Decoder (Matthias Bastian)THE independent in-window confirmation of the Sep 10 API launch; "speech model… full-duplex"; $0.05/min ("not cheap"); backend pairings by task; 12 new voices; Yelp phone-reservation usage with better call handling (CTO Alex Levy); ASR transcripts + response text out of the box. — Independent / CONFIRMED (event, price, capability framing — speech, not video).Date: 2026-09-10
Visit source ↗
Independent
ChatGPT's upgraded voice mode is better at shutting up — The Verge (Emma Roth)Independent confirmation of GPT-Live-1's identity as a full-duplex VOICE model ("This is a full duplex model… can speak and listen at the same time" — OpenAI product lead Atty Eleti); "smartest voice model" (research lead Kundan Kumar); delegation to text models (GPT-5.5) for reasoning/search; no video/screen-sharing at launch; safety features. Pre-window but decisive independent evidence for the §1 correction. — Independent / CONFIRMED (model identity — voice, not video).Date: 2026-07-08 (pre-window; consumer launch)
Visit source ↗
Secondary Sources (2)
Secondary
GPT-Live-1 API Tutorial: Build a Voice Agent in 20 Min — AI Bytes (Shadman Ahmed)Third-party code walkthrough used for the INSPECT lab: client.live.connect() WebSocket flow; session.start with model gpt-live-1 / marin voice; 24 kHz PCM16 downsampling requirement (browser 48 kHz); G.711 μ-law 8 kHz for Twilio bridging; barge-in handling via transcript deltas + playback flush; cost note ($0.05/min per-second; idle sessions keep billing); voice roster consistency (alloy…marin/cedar); no general-purpose voice cloning at launch. — Secondary (practitioner tutorial) / supports §5, §9, §19 lab walkthrough; consistent with primary docs.Date: 2026-09-15
Visit source ↗
Secondary
Introducing GPT-Live-1 in the API — thread replies (OpenAI Developer Community)EARLY RESEARCH on production behavior from multiple builders a week post-launch: user-side transcript retrieval discrepancy (recorded transcript contains only assistant messages in reported deployments); response.done.output consistently empty in delegation mode (function calls arrive via response.output_item.done in response.event envelopes); double response.completed events; missing session.update/response.cancel/input_audio_buffer.clear equivalents; voice-model commentary before session.delegation.created; hangup/end_conversation flows unreliable in semi-rigid workflows; praise for tool calling, per-second billing and interruption handling vs gpt-realtime-2.1. Unverified individual reports — flagged as EARLY RESEARCH, not CONFIRMED. — Secondary (community forum, user-reported) / EARLY RESEARCH (limitations, migration risk).Date: 2026-09-18 (replies; in-window to research date)
Visit source ↗