Tech
EN AZ
These execs think voice AI hasn’t reached its ChatGPT moment yet

These execs think voice AI hasn’t reached its ChatGPT moment yet

techcrunch.com 11.10.2026 16:00 5 views
Voice AI's often misses important points for its context layer, and causes the whole pipeline to break

The theory of voice being the next big interface has picked up strong momentum, with investors pouring billions of dollars into voice AI startups working on areas ranging from model makers to enterprise customer service providers, and from meeting note-takers to AI-powered dictation. Every week there is a new model or a tool release that claims to sound human and converse like one. However, in reality, that might not be the case.

Enterprise voice AI platform PolyAI’s CTO Shawn Wen thinks that despite the release of full-duplex models — which can speak while listening to you — voice AI doesn’t have its “ChatGPT moment” yet. The next challenge is to make reasoning very fast, so that the models can fetch answers quickly and the conversation feels natural,” he told me on stage at the HumanX conference last month. He also said that AI agents in customer service should not sound robotic and should give callers enough confidence that they can solve problems.

Alex Gay, CMO for meeting notetaker Otter, opined that speaker identification, intent capture, and typing that up with organizational knowledge is a key step for enabling automation. The company is also working on digital twins that might represent people in meetings. For that technology, he said it’s paramount that the output voice gives the same emotive expressions of talking to a human in a meeting.

If you aren’t able to have that with an avatar, then it’s just a q and a chatbot,” Gay noted. While voice AI models have improved, AI assistants often don’t understand users, or your meeting notetaker shows the wrong transcript or a summary. Wen thinks that ASR (Automatic Speech Recognition) models often miss important keywords, and that creates an issue in capturing the whole context.

Otter’s Gay agreed with this, adding that the company keeps working on improving transcription. He also said that language is one area where voice models need to improve. It was just the layer that we could start to drive some of the productivity gains on the back of.

But if your original transcription didn’t have the accuracy that you needed, all the follow-up actions that you have become flawed. And the minute that starts to take action, that is wrong. You lose trust in the platform.

Extract — continue reading at the source.

Read full story