OpenAI’s GPT-Live Voice System Can Now Listen and Speak at the Same Time

OpenAI's GPT-Live Voice System Can Now Listen and Speak at the Same Time

OpenAI says its third-generation voice architecture removes the turn-detection step from the audio path, cutting session startup from six network round trips to one.

OpenAI has published a detailed engineering explanation of GPT-Live, its third-generation voice system for ChatGPT, saying the technology can listen and speak simultaneously because it is built on a full-duplex architecture. The company says this removes a long-standing limitation of earlier voice assistants, which had to stop speaking before they could process what a user said.

The announcement, made via an OpenAI post and repeated in a tweet from the company’s official account, frames GPT-Live as a realtime voice architecture designed to keep audio flowing without interruption. According to OpenAI, the system streams incoming audio directly into the voice model while simultaneously streaming outbound speech back to the user.

How the Architecture Works

The core change, according to OpenAI’s engineering write-up, is that the system removes the turn detector from the audio path entirely. In previous generations of voice AI, the model had to detect when a user had finished speaking before it could begin formulating and delivering a response. That detection step introduced delays and created the stilted, walkie-talkie-like rhythm that many users found frustrating.

GPT-Live keeps a dedicated fast path open between the client device and the voice model. Audio flows in and out continuously. The model makes interaction decisions many times per second — whether to speak, keep listening, pause, allow the user to interrupt, or call up a tool — rather than waiting for a clear signal that a turn has ended.

That last point matters. Tool use and deeper reasoning, which can be computationally heavier, happen asynchronously. They’re offloaded to a separate path so they don’t block or interrupt the live audio loop. OpenAI says GPT-5.5 handles the deeper reasoning work, while the live audio path stays lean and fast.

The Startup Time Reduction

One of the more concrete claims in OpenAI’s announcement is the reduction in voice-session startup time. According to the company, the new architecture cuts the number of network round trips required to begin a voice session from six down to one. In practice, that means a conversation should begin faster after a user opens the voice interface.

Six round trips to one. That’s a significant drop in the handshake overhead, and it’s the kind of engineering detail that tends to matter more in real-world use than it does on a spec sheet.

OpenAI says the architecture is designed for continuous interaction rather than the traditional request-response model, where each exchange is a discrete event with a clear start and end. The company frames GPT-Live as a persistent media loop, more like a phone call than a search query.

What OpenAI Has and Hasn’t Said

OpenAI’s published explanation is largely its own account of how the system works. The supplied sources do not include independent academic verification of the performance claims, nor any regulatory assessment. The company has not published third-party benchmarks comparing GPT-Live’s latency or interruption handling against competing voice systems from Google, Apple, or Amazon.

The claim that the model can make interaction decisions “many times per second” is OpenAI’s own characterisation; the company has not specified a precise figure. The six-to-one round-trip reduction is stated as fact in OpenAI’s engineering post, but has not been independently audited.

There’s also no clarity yet on which ChatGPT subscription tiers will have access to GPT-Live, or on what timeline the rollout will reach users outside the United States.

Where This Fits in the Voice AI Market

Voice AI has been a contested space for several years. Apple’s Siri, Google Assistant, and Amazon Alexa all built their early architectures around the same turn-based model that OpenAI says it has now moved away from. More recently, Google’s Gemini Live and Meta’s voice features have also attempted to address the interruption and latency problems that made earlier voice assistants feel unnatural.

OpenAI’s decision to publish a detailed engineering explanation — rather than just a product announcement — suggests the company wants to distinguish GPT-Live on technical grounds, not just on marketing. But without independent testing, the claims remain OpenAI’s own.

What This Means for Kent Residents

For people in Kent who use ChatGPT’s voice mode — whether on a smartphone, tablet, or laptop — GPT-Live is the update most likely to change the day-to-day feel of those conversations, reducing the awkward pauses and interruptions that have been a common complaint. There’s no indication of a Kent-specific rollout or any change to local public services; this is a consumer software update with UK-wide relevance. Residents who use ChatGPT on free or paid plans should check OpenAI’s announcements for confirmation of when GPT-Live becomes available in the UK.

Source: @OpenAI

OpenAI's GPT-Live Voice System Can Now Listen and Speak at the Same Time Quiz

5 questions