AI outfits are betting that talking to chatbots will replace typing, despite the weirdness of muttering at software in public.
According to the Financial Times, OpenAI and Google have been fiddling with their voice kit to make chatbots sound less like a bored satnav and more like a colleague who remembers the point.
The new systems are a break from old voice assistants, which turned speech into text, generated a text answer and then read it aloud. Newer models can chew through speech directly, cutting delays, wooden delivery and the mangled context that made the earlier stuff feel a bit naff.
OpenAI ChatGPT voice product lead Atty Eleti said: “People love using their voice to ask questions. It is just good for that humanlike back-and-forth conversation. Text is great too, but typing is a lot of effort.” The idea is that longer-running AI jobs will feel less painful when users can just say what they want.
The snag is that voice tech still has old baggage, including lag, background noise and the joyless spectacle of people barking commands at machines on trains. Google said its live voice use doubled in the year to April, while OpenAI claims more than 150 million people use voice in ChatGPT each week.
Microsoft responsible AI chief product officer Sarah Bird said voice was “really starting to rise again with these long-running agents where it just feels very natural to give them tasks via voice”. Bird said she had seen colleagues on GitHub sending voice commands to their phones to code, adding: “I think it’s going to be a new world soon on this.”
OpenAI said its older voice tech suffered from a “lag in intelligence”. Until ChatGPT Live arrived last week, its first voice model in two years, the outfit had been mocked for being slow and muddling languages or accents.
Google voice experiences lead Leland Rechis said: “People are just starting to speak more because they know the model understands, and they’re now expecting they can do more [and] they want those devices to do more.”
Rechis said Google had gone for a “far more spontaneous conversational style” after users warmed to AI once large language models were “behind the microphone”.
Coding is one early use, with fans claiming speech helps users explain intent, constraints and uncertainty better than a clipped prompt.
OpenAI international strategy and operations staffer Sandro Gianella wrote: “When I type, I often compress the problem too early and jump straight to the task. When I speak, I tend to explain more of the why: what I am trying to achieve, what my intention is, which constraints matter, where I am uncertain, and what a good outcome would look like.”
The voice land grab has pushed Silicon Valley start-ups into microphones, office soundproofing, black-tube gadgets, mouthguards and other kit to stop AI from mistaking office chatter for instructions.
Pocket and Plaud are trying handheld dictation devices, while OpenAI is expected to punt out an ambient voice assistant with a camera, microphone and speaker early next year, according to people familiar with the project.







