Sarvam AI and Navana.ai Launch New Indian Voice Models
Sarvam AI and Navana.ai released new speech recognition and text-to-speech models designed to handle the linguistic diversity and security needs of the Indian market.
Two Indian artificial intelligence companies launched specialized voice models on September 24 and 25, 2026, to address the complexities of multilingual speech in India.
Sarvam AI released Saaras V4, an automatic speech recognition model supporting English and 22 Indian languages. The model uses a 3-billion-parameter hybrid state-space language model to manage noisy audio, dialect variations, and code-mixed conversations. It provides five output modes—verbatim transcription, normalized text, code-mixed text, transliteration, and translation—and supports near-real-time streaming with a time to first token under 150 milliseconds. The company reported a language identification error rate of 5.22% across its supported Indian languages.
Simultaneously, Navana.ai unveiled Bodhi TTS, a text-to-speech model supporting 10 Indian languages and over 50 voices. Designed for regulated enterprises such as banks, Bodhi TTS allows for on-premises deployment to ensure data security. Navana.ai priced the service at ₹12 per 10,000 characters, claiming this is significantly lower than competitors. The company also opened invitation-only access to Bodhi STT, a speech recognition model supporting 10 languages and 40 dialects.