OpenAI and Google DeepMind Launch Advanced AI Voice Tools
OpenAI introduced voice-based agentic features for ChatGPT while Google DeepMind released two high-performance text-to-speech models to compete for the AI voice infrastructure market.
AI leaders OpenAI and Google DeepMind launched significant voice technology updates on Wednesday to capture the growing market for AI agents. OpenAI integrated its GPT-Live voice model with agentic AI capabilities across its mobile app and website, allowing users to delegate multi-step tasks such as booking appointments, analyzing spending, and reorganizing calendars. These updates enable Plus and Pro subscribers to trigger complex workflows, including drafting documents and summarizing Slack messages, via voice commands.
Product lead Atty Eleti stated that the company views voice as the primary future modality for AI interaction. The rollout includes a new Work tab for mobile users to build websites and access financial data, as well as the ability to resume mobile conversations on a desktop. This expansion precedes the OpenAI DevDay event and aligns with reports of a planned smart speaker launch in 2027.
Simultaneously, Google DeepMind released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. The flagship model ranked first on Artificial Analysis's Pronunciation Robustness Benchmark, which measures the reliable handling of technical terms and proper nouns. By offering both a premium and a cost-effective Lite version, Alphabet Inc. aims to compete with OpenAI and ElevenLabs for enterprise API adoption and the underlying voice layer of AI agents.