Sarvam Vision 2.1 targets forms, tables and Indic handwriting

By Aditya Rudraksh Sehgal••7 min read
Representational illustration of multilingual paper documents beside a laptop showing abstract document blocks
Representational illustration created for Dalimss News.

BENGALURU: Sarvam AI has pushed a new document model into its stack. Vision 2.1, described in company materials dated 24 September 2026, is aimed at forms, multi-page tables and handwritten text across English and 22 Indian languages - the same Indic breadth Sarvam has been building into its OCR line since the original Sarvam Vision release in February 2026.

The public API page for document digitisation already markets structured extraction for PDFs, PNG and JPG inputs, HTML and Markdown outputs, and async jobs for batch work. List pricing on that page starts at Rs 1.5 per page.

What 2.1 adds on paper

Sarvam says Vision 2.1 is stronger on complicated layouts: tables that span pages, forms that must map into fixed fields, and handwriting in Indian scripts that older OCR pipelines smear. The company reports 87.3 on olmOCR-Bench and 87.39 on its own Indic benchmark, calling both leading scores in its brief.

Training, according to Sarvam, mixed real and synthetic forms, web-sourced filled templates, and handwritten samples gathered from video and other sources, followed by supervised fine-tuning and reinforcement learning with verifiable rewards. The architecture still pairs a vision-language model with layout parsing and reading-order helpers rather than asking one model to guess an entire page alone.

Earlier Vision materials already stressed table parsing, chart reading and in-the-wild scene text. Version 2.1's pitch is reliability: fewer invented fields, more consistent outputs, and cheaper inference than Sarvam first planned for the API.

Who actually buys this

Banks, insurers, hospitals and government scanners sit on mountains of paper that never arrive as clean UTF-8. A model that keeps cell structure in a nested table, or reads a handwritten Marathi form without collapsing matras, shortens that backlog. Sarvam's developer path - OpenAI-compatible APIs, Python and Node SDKs, a free tier - matches that buyer's habit of wiring OCR into an existing workflow rather than buying another desktop suite.

Buyers should still run their own pages. Global OCR benches skew English; Indic accuracy varies sharply by script, scan quality and handwriting style. Sarvam's own earlier notes flagged edge cases on low-resource languages and occasional mistranslation while describing scenes. Vision 2.1's published scores are company-reported.

The dated update on 24 September 2026, the Rs 1.5-per-page API sticker, and the 22-language claim are the concrete public facts. For Indian enterprises digitising files that never left a metal cupboard, that is the product on offer.

Pricing at Rs 1.5 per page makes spreadsheet arithmetic easy for operations teams. A lakh pages is a known line item; surprise token bills are not. Async jobs matter when overnight batches hit fifty thousand scans from a district archive.

Sarvam's insistence on knowledge extraction - charts, nested tables, reading order - separates Vision 2.1 from engines that only dump raw text. Enterprises fail KYC and claims workflows when a merged cell loses its header, not when a single Latin character is wrong.

The 24 September update also cuts the distance between a February research blog and a purchasable API. Teams that piloted the first Vision stack now have a dated reason to re-benchmark forms in Hindi, Tamil and Odia before renewing contracts.

Sources and reporting

Sarvam AI product and API materials for document digitisation / Sarvam Vision; company update dated 24 September 2026 for Vision 2.1 describing improved form/table/handwriting handling, olmOCR-Bench score 87.3 and Indic benchmark 87.39, and lower inference cost relative to earlier plans.

Related Stories