Build more natural voice experiences with GPT‑Live‑1 in the API
OpenAI 正在其 API 中推出 GPT-Live-1,为开发者提供强大的自然语音模型,用于构建语音应用和业务工作流程。GPT-Live-1 最初在 ChatGPT 中引入,能够同时听和说,并能将更深层次的推理和操作委托给与其配对的模型和工具。在评估中,GPT-Live-1 将 Full Duplex Bench 性能比 GPT-Realtime-2.1 提高了 30 个百分点,并在与 GPT-6 Astra 配对时在 Tau3 上排名第一。OpenAI Presence 也利用 GPT-Live-1 来支持实时语音交互,帮助企业部署可信赖的 AI 代理。
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
发布当时偏移:UTC+02026年9月10日 00:00 UTC
收录当时偏移:UTC+02026年9月12日 16:01 UTC
- 发布
- 2026年9月10日 00:00
- 收录
- 2026年9月12日 16:01
- 来源类型
- 官方发布
- 档位
- 当事方
- 信源状态
- 正常
档位是按信源手工设定的编辑判断,不是逐条打分。
讨论趋势
百分比基于采集到的讨论信号,不代表新增评论数或独立参与人数。曲线仅用于同一话题在不同时段的比较。
We’re launching GPT‑Live‑1 in the API, giving developers a powerful, natural voice model for building voice-enabled apps and business workflows. First introduced in ChatGPT , GPT‑Live‑1 is capable of listening and speaking at the same time, and, as seen with Codex and ChatGPT Work (opens in a new window) , can delegate deeper reasoning and actions to the models and tools it is paired with.
For the API release of GPT‑Live‑1, we’ve focused on new capabilities that let developers steer and customize voice experiences around their users, workflows, and goals. A core GPT‑Live‑1 strength, smooth interruption handling, is already delivering business impact: in early evaluations, Speak found that GPT‑Live‑1 gave learners more time to think before the language tutor responded, cutting interruptions by almost 80% versus previous turn-based systems.
Key strengths of GPT‑Live‑1 in the API:
- Interruption handling: Improves interruption handling via a single model that reasons over incoming and outgoing audio together, avoiding the latency and brittle handoffs of chained STT–LLM–TTS architectures.
- Reasoning & tool calling delegation: GPT‑Live‑1 can delegate reasoning and tool calls to a backend text model like GPT‑6 Astra or a third-party model.
- Tone, pace, and style: Lets developers shape an agent’s tone, pace, and conversational style through the system prompt.
- Silent context management & background noise: Better handles background noise and silence without interrupting the conversation or narrating every step out loud.
- Long-session reliability: Improves context retention and conversational quality across extended interactions.
- Telephony support: Enables deployment of full-duplex voice agents for phone calls, from restaurant reservations to customer support.
Try GPT-Live-1
Start a session and speak naturally. Interrupt, laugh, change your mind - try it at home or in a loud space like a coffee shop or city street.
See what it can do - **Talk over it—naturally.** Ask for help, then interrupt mid-response to change the question or add detail.
- Take it with you. Try a conversation while walking outside or with everyday background noise, and see how it stays with you.
- Make it playful. Laugh, hesitate, use short acknowledgments, or briefly talk to someone nearby—then continue the conversation.
This demo is time-limited. By using it, you agree to OpenAI's Terms and acknowledge our Privacy Policy .
Simplify your voice-agent architecture and reduce voice latency
Traditional voice agents stitch together speech-to-text, a reasoning model, and text-to-speech. Each handoff adds latency and creates more opportunities to lose timing, context, or the natural rhythm of a conversation. Developers are often the ones left coordinating those stages, including what happens when someone interrupts, pauses, or changes direction.
GPT‑Live‑1 handles listening and speaking in a single model, simplifying the voice layer. It can respond to interruptions and acknowledgements as they happen, while delegating deeper reasoning to the back end. This lets the conversation continue while work happens in the background.
Developers choose the models, tools, and agent harness behind the conversation. For example, they might pair GPT‑Live‑1 with a model like Luna for high-volume tasks like scheduling or order updates, and use a model like Astra for complex customer issues that require reasoning. That flexibility lets developers match reasoning depth, speed, and cost to each task.
GPT‑Live‑1 natively provides ASR transcripts and response text. It also offers strong alphanumeric understanding and supports keyword biasing. Although GPT‑Live‑1 is not a turn-based model, it natively supports turn detection, so developers can continue to build around explicit turn boundaries.
Measuring the full-duplex advantage
Across our evaluations, GPT‑Live‑1 improves Full Duplex Bench performance by 30 percentage points over GPT‑Realtime‑2.1, with large gains in turn-taking latency and interactive behavior. Paired with GPT‑6 Astra at medium reasoning effort, it also ranks #1 on Tau3, which measures frontier voice-agent intelligence on end-to-end tasks.
Evaluates spoken customer-service tasks in airline, retail, and telecom domains. Pass@1 measures task success; the headline gives each domain equal weight.
* GPT Live backend: Astra (medium).