Categories
Technology

Sarvam AI unveils Saaras V4 for multilingual speech

New speech model supports 22 Indian languages and mixed-language audio

Bengaluru-based artificial intelligence startup Sarvam AI has launched Saaras V4, its latest automatic speech recognition model designed to handle multilingual and real-world audio more accurately.

The speech-to-text model supports all 22 scheduled Indian languages along with English and is built to work with noisy recordings, code-mixed conversations, different accents and multiple speakers. It is aimed at applications such as voice assistants, customer-service calls, meetings, interviews and other voice-based AI tools.

One of the key additions in Saaras V4 is its ability to process multiple speakers in a single pass. Instead of using separate systems for speaker identification and transcription, the model combines the two tasks. This can help businesses analyse call recordings, meetings and interviews where identifying who said what is important.

The model also offers five output formats from the same audio. Developers can generate a verbatim transcript, normalised text, code-mixed text, transliteration or translation without relying on separate post-processing systems.

This is particularly relevant in India, where people frequently switch between English and regional languages in the same conversation. A user speaking in a mixture of Hindi and English, for example, can be transcribed in a code-mixed format, while the same audio can also be converted into another output format.

Saaras V4 can automatically identify the language being spoken and transcribe it in the appropriate native script. Sarvam says the model has been designed to retain strong performance across Indian languages, including languages with fewer available digital speech resources.

The company has also added support for Global English, extending the model beyond Indian English. Sarvam says Saaras V4 recorded the lowest average word error rate across seven English speech benchmarks covering areas such as Indian English, international accents, meetings, financial conversations and media. These are company-reported benchmark results rather than an independent assessment.

For Indian languages, the model was tested on the Vistaar benchmark across 10 languages using both standard Word Error Rate and LLM-WER. Word Error Rate measures transcription mistakes, while LLM-WER also considers whether an error changes the meaning of a sentence.

Sarvam reported a language-identification error rate of 5.22% across 22 Indian languages on verified IndicVoices data, and 2.9% across the 10 most widely spoken languages. The model recorded a 16.03% Word Error Rate in the L5 keyword-prompting setting on the IndicContextEval benchmark.

Another feature is keyterm prompting. Developers can provide names, places, brands, acronyms or technical terms that the system should pay particular attention to while transcribing audio. The current API documentation allows up to 50 such terms for Saaras V4 on supported REST and batch requests.

The model is also designed for low-latency applications. Sarvam says Saaras V4 supports streaming with time to first token below 150 milliseconds, making it suitable for applications where users expect a quick response while speaking. It can also process long-form audio.

The technology behind the model combines an audio encoder with a 3-billion-parameter hybrid state-space language model developed by Sarvam. The company says the language model was trained from scratch in-house.

Developers can access Saaras V4 through Sarvam’s API using Python and Node.js SDKs. The model is also integrated with platforms including Vercel AI SDK, LiveKit Agents and Pipecat Agents. Its speech-to-text API supports transcription, translation, transliteration, verbatim and code-mixed output modes.

Saaras V4 is part of Sarvam’s wider push to build AI systems suited to India’s multilingual environment. The company has been developing models across language, speech, vision and AI infrastructure, with a focus on applications that can work across India’s diverse languages and communication patterns.

The latest speech model could have applications across sectors where voice data is central to business operations. Banks and financial companies could use it to analyse customer calls, while healthcare providers could use speech recognition for documentation. Government services, education, media and customer-support platforms are other potential use cases.

Sarvam’s API documentation lists Saaras V4 as the latest model while Saaras V3 remains the default recommended model in some API workflows. V4 is available across REST, batch and WebSocket speech-to-text services, with newer updates also extending keyterm prompting to streaming applications.

With Saaras V4, Sarvam is positioning speech recognition as a core layer for India’s growing voice-AI ecosystem. Its emphasis on Indian languages, code-mixed speech, speaker identification, real-time processing and multiple output formats is aimed at making voice applications easier to build for India’s varied linguistic environment.

 

Leave a Reply

Your email address will not be published. Required fields are marked *