Wednesday, September 16, 2026
OpenAI Launches GPT-Live-1 API at 5 Cents a Minute

OpenAI Launches GPT-Live-1 API at 5 Cents a Minute



OpenAI GPT-Live-1 API is now generally available, giving developers a voice layer that can listen and speak simultaneously while a separate backend model handles reasoning, tools and application logic. OpenAI priced the live voice session at $0.05 per minute, billed by the second.

 

The launch gives voice-agent builders three core capabilities:

  • Full-duplex speech with interruption handling
  • Delegation to a separate reasoning model
  • Browser, server and telephony connection options

 

OpenAI GPT-Live-1 API Separates Voice From Reasoning

OpenAI added GPT-Live-1 to its API on September 10, 2026, according to the company’s developer changelog. The model is accessed through the v1/live/sessions endpoint and is designed to keep a spoken exchange moving while another model or agent completes slower work in the background.

 

That separation is the main architectural change. GPT-Live-1 manages the timing-sensitive conversation: listening, speaking and responding to interruptions. A backend model can independently search data, call business systems, reason through a request or perform a multi-step task.

 

OpenAI documents two ways to connect those layers. Responses delegation lets the Live session send work to the Responses API, while client delegation allows the application to route tasks through its own server. In both cases, the developer chooses the backend model instead of tying the voice interface to one fixed reasoning engine.

 

The design could make voice products easier to update. A company can change its reasoning model, retrieval system or tool chain without rebuilding the conversational layer. It can also reserve a more capable model for complex requests while keeping routine dialogue responsive.

 

Full-Duplex Design Keeps Conversations Moving

Most voice assistants alternate between listening and speaking. GPT-Live-1 instead uses full-duplex audio, meaning it can continue listening while it talks. OpenAI says this allows more natural interruptions, acknowledgements and turn-taking rather than forcing users to wait for a rigid end-of-speech boundary.

 

The company’s Live API guide describes the model as a conversation layer that remains active while the backend looks up information or carries out a task. That could help call-center, scheduling and support systems avoid long silent pauses during database queries or workflow execution.

 

Independent analysis from DataNorth reported that OpenAI measured a 30-percentage-point improvement over GPT-Realtime-2.1 on its Full Duplex Bench. The benchmark result is a vendor-reported figure, however, and production performance will still depend on microphones, network conditions, prompt design and backend latency.

 

OpenAI’s prompting guidance recommends keeping instructions for the live speaker concise and putting detailed workflows in the backend. That division treats natural dialogue and task execution as separate engineering problems, with the application coordinating the two.

 

WebRTC, WebSockets and SIP Shape Deployments

Developers can connect live sessions through WebRTC, WebSockets or SIP. WebRTC suits browser and mobile experiences where low-latency microphone access matters. WebSockets give server applications direct control over streamed events, while SIP supports phone-based services and contact-center infrastructure.

 

A browser implementation still needs a trusted server to create sessions and protect the API key. OpenAI’s documentation also leaves permissions, confirmations, function execution and durable task state with the application. The voice model can carry the conversation, but it does not replace the authorization layer around sensitive actions.

 

That boundary is especially important for agents that can make purchases, change reservations or access personal records. Developers must validate arguments, restrict available tools and require confirmation where appropriate. A natural voice can make an automated system feel more capable than the permissions behind it actually are.

 

For telephone products, OpenAI’s SIP guide supports either direct SIP connectivity or a server-side audio bridge. Those options let companies integrate the model with existing phone workflows instead of forcing every deployment into a new consumer app.

 

Five-Cent Rate Does Not Include Backend Work

The $0.05-per-minute price covers the live voice session and is billed per second. Backend model usage and tools are charged separately. A deployed agent’s real cost therefore depends on call duration, the model selected for delegated work, retrieval volume and any external services used during the conversation.

 

For a ten-minute interaction, the live layer alone would cost $0.50 before backend processing. The modular design gives developers more control over that second bill: straightforward requests can go to a lower-cost model, while difficult tasks can be escalated to a stronger one.

 

The commercial question is whether improved conversational flow justifies the added voice charge. Customer-service operators will compare it with human handling time, abandoned-call rates and existing speech pipelines. Consumer developers will also need to manage usage carefully because always-on audio sessions can accumulate cost faster than short text exchanges.

 

GPT-Live-1 enters a market where the interface is becoming less important than orchestration. The launch gives OpenAI a dedicated real-time speaker while leaving reasoning, tools and governance under application control. The first meaningful tests will come from live deployments, where latency, interruption accuracy and safe execution matter more than a polished demo.

 

Related Coverage

THEFLGHT
author

THEFLGHT

Elevating narratives from the heart of London's intellectual epicentre.

0 Comments:

Leave a Reply

AI Agent Data Breach: Spain Discloses Its First Report
Google Launches Gemini 3.8 Live and Extended Thinking Voice Models
Meta Launches Meta One AI Subscriptions Across Instagram, WhatsApp and Facebook
Axelera Launches Europa AI Chip for Dell and Supermicro Systems
Einride Autonomous Truck Launches Into Daily Lidl Service in Germany
DeepSeek IPO Plan Taps Yan Wentao as First CFO