Voice AI can now listen and speak at the same time, but industry executives say it still lacks the speed, accuracy and context needed for a broad breakthrough. PolyAI chief technology officer Shawn Wen said faster reasoning is the next challenge after full-duplex conversation, because delays make an exchange feel unnatural.

The harder problem is what happens after speech recognition. Otter marketing chief Alex Gay said an incorrect transcript can poison summaries and any automated action built on top of them. PolyAI similarly sees missed keywords as a context failure, not merely a cosmetic transcription error.

That distinction matters most in customer service and meetings. A natural-sounding voice may keep a caller engaged for the first few turns, but confidence disappears if the system misunderstands the request or acts on the wrong information. Otter is also exploring digital representatives for meetings, where tone and emotional expression would affect whether an exchange feels like a discussion rather than a question-and-answer bot.

Both companies also emphasized disclosure. People should know when they are speaking with an AI or being recorded, including meetings where no visible bot joins the call. Better voices therefore solve only one layer; reliable understanding, low latency and transparent operation still determine whether the system can be trusted.