OZZZER · AI NEWS3 of 3 free stories opened
← Back to AI News

AI Audio · 8 Jul 2026 · 02:00 CEST

Introducing GPT-Live

OpenAI · 8 Jul 2026 · 02:00 CESTRead original at OpenAI ↗
Share
LinkedInX
Introducing GPT-Live

Publisher preview · OZZZER analysis pending editorial review.

PUBLISHER ARTICLE PREVIEW

From the original article

A new generation of voice models for natural human-AI interaction, now powering ChatGPT Voice.

Update July 31, 2026: Supported audio generated with GPT‑Live through ChatGPT Voice and the OpenAI API now includes SynthID watermarking. Our public verification tool can now detect OpenAI provenance signals in supported audio files, and we’ve introduced API access for verification so developers and organizations can incorporate provenance checks into their own workflows. These updates build on the safety approach described below and our broader work to make OpenAI-generated content easier to identify.

We’re launching GPT‑Live, a new generation of voice models that make talking with AI feel much more like having a real conversation.

GPT‑Live is built on a full-duplex architecture, meaning it can listen and speak at the same time. During conversations, GPT‑Live can show it’s paying attention with phrases like “mhmm” or “yeah”, engage in quick back-and-forth, or just stay quiet when you need a moment to think. The result is a voice experience that is refreshingly easy to talk to.

GPT‑Live is also our smartest voice model yet. For questions that require web search, deeper reasoning, or more complex work, it delegates to our latest frontier model behind the scenes and brings the result back into the conversation when it’s ready. While it works, GPT‑Live can keep talking with you and maintain the flow of conversation.

At launch, GPT‑Live will use GPT‑5.5 in the background. As we release new frontier models, we’ll continuously update the model used by GPT‑Live.

These advances power a new ChatGPT Voice experience that is more intelligent and natural to use. Over time, we believe this research will also unlock the ability to use voice for increasingly complex, longer-running, and more agentic work.

We’re beginning to roll out two versions of GPT‑Live – GPT‑Live‑1 and GPT‑Live‑1 mini – to ChatGPT users globally today. We also plan to bring them to the API soon, and developers and enterprises can sign up to be notified using this form⁠.

Our vision is to enable truly natural human–AI interaction: a world where collaborating with AI feels as fluid and responsive as working with another person, while reasoning and complex task execution happen seamlessly in the background.

Older generations of voice AI systems brought us closer to that vision, but with important tradeoffs.

Cascaded voice systems rely on a series of models acting one after another to process each turn. The original ChatGPT Voice chained three models together: a speech-to-text model to transcribe your speech, a large language model to produce a response, and a text-to-speech model to convert it back into speech. This approach enabled us to talk to frontier AI models for the first time, but the complexity came at a cost: information could be lost across models, and responses were slow and stilted.

Turn-based voice models like ChatGPT Advanced Voice Mode processed and generated audio within a single model, reducing latency and making conversations smoother — but they still operated through discrete turns. The model had to wait for the user to stop speaking

Source

OpenAI · 8 Jul 2026 · 02:00 CEST

Open the original at OpenAI ↗