Sarvam Releases Saaras V4 and Vision 2.1 to Hear and Read India's 22 Languages
Sarvam AI has released Saaras V4, a fast speech-to-text model for 22 Indian languages and English, and Vision 2.1, a document model that reads Indic handwriting, forms and tables, claiming gains over larger global models.

Bengaluru’s Sarvam AI is expanding its Indian-language speech and document tools with Saaras V4 and Sarvam Vision 2.1. Sarvam’s own pages date the speech-model announcement to 24 August 2026 and Vision 2.1 to 24 September 2026; later coverage brought the two releases together. One turns spoken Indian languages into text; the other reads printed and handwritten Indian scripts.
Both cover all 22 scheduled Indian languages plus English. The Saaras V4 decoder has roughly 3 billion parameters. Both come with benchmark claims that put them ahead of much larger global models on Indian-language tasks, although those numbers are largely company-reported.
Saaras V4: listening the way India speaks
Speech recognition for Indian languages is harder than for English. The problem is not only the number of languages. Speakers routinely switch between them mid-sentence, accents vary widely by region, and much real-world audio comes from noisy streets or low-quality phone lines.
Saaras V4 is built for that. Coverage from Business Standard, WION, MarkTechPost and others describes these features:
- Five output modes: plain transcription, translation, transliteration (writing one language in another's script), verbatim output, and a code-mixed mode for speech that blends languages.
- Architecture: an audio encoder paired with a 3-billion-parameter hybrid state-space language model as the decoder, trained from scratch. State-space models are an alternative to the transformer design behind most large language models, and are generally more efficient on long sequences such as audio.
- Streaming latency under 150 milliseconds, fast enough for live voice agents.
- Keyterm prompting: developers can give the model a list of names, product terms or acronyms it should expect, which helps with the brand names and jargon that trip up general-purpose models.
- Speaker diarization is available in the batch API.
Sarvam reports a language-identification error rate of 5.22% across 22 languages and 2.9% across the top 10, measured on verified IndicVoices data.
Sarvam reports contextual-prompting results on IndicContextEval. Its current page gives 16.03% word error rate in the L5 keyword-prompting setting. These are company-reported evaluation results, not a guarantee for every recording.
The model is available through Sarvam's API, with Python and Node.js SDKs and integrations for voice-agent frameworks such as LiveKit Agents and Pipecat.
Vision 2.1: reading India's documents
Much of India's paperwork still lives on paper. Land records, government forms, old newspapers and handwritten notes are full of Indian scripts that standard optical character recognition (OCR) handles poorly.
Sarvam Vision 2.1 is a vision-language model built for that kind of document intelligence. Reports from Zee Business, Storyboard18, Analytics India Magazine and others say it can:
- read multi-page tables that span several pages;
- extract fields from forms;
- read handwritten Indic text;
- do all this with lower hallucination and lower inference cost than the earlier Sarvam Vision released in February.
Sarvam also released a new benchmark alongside the model, the Sarvam Indic OCR Bench. It holds 6,909 samples: 6,609 across 22 Indian languages and 300 in English. They are drawn from newspapers, brochures, textbooks and historical texts dating from 1800 to the present.
On that benchmark, Vision 2.1 scores 87.39% overall word accuracy. The per-language breakdown is revealing. Konkani reaches about 97%, but Santhali manages only about 54%. That spread is the honest story of Indian-language AI: some scripts perform much better on this benchmark, while languages with little digitised material still lag badly.
Why this matters
Voice is how most Indians will use AI. Hundreds of millions of Indians are more comfortable speaking than typing, and many are more comfortable in their own language than in English. Customer service lines, banking, agriculture advisories and government services all run through voice. A fast, cheap, accurate speech model that understands code-mixed Hindi-English or Tamil-English is basic infrastructure for Indian AI products.
Documents are the other bottleneck. Digitising India's paper records, from court files to land titles, has been slowed by poor Indic OCR. A model that reads handwriting and complex tables in 22 languages could speed up that work considerably.
Deployment cost matters. Sarvam says inference-stack optimisations make Vision 2.1 cheaper to serve. Actual cost depends on workload and deployment.
It advances the sovereign AI agenda. Sarvam is one of the companies selected under the IndiaAI Mission to build indigenous foundation models. Specialised models like Saaras and Vision, tuned for Indian languages and trained by an Indian company, are the practical side of that mission: the tools developers actually build with.
The caveats
Benchmarks deserve scrutiny. Most of the comparisons against Gemini and GPT-4o come from Sarvam itself, and some rely on a benchmark Sarvam created. Independent evaluation on real-world audio and documents will be the real test. The Santhali result is a reminder that coverage of 22 languages does not mean equal quality across them.
Still, the direction is clear. Indian-language AI is moving from demos to production-grade tools that are fast, affordable and built for how India actually talks and writes.
Sources
Core announcement or reporting
Image: Sarvam AI. Official Sarvam AI wordmark; model-developer identity. Original source.