The Hallway Track
Product Launches

From Scratch to SOTA: Training a 3B State-Space Vision Model — Krishna Prasad Srinivasan, Sarvam

Aravind Srinivas · AI Engineer · Sep 23, 2026 · Product Launches

Sarvam Vision, a 3B state-space VLM built entirely in India, beats models 100x larger at Indic document AI.

“India is largely absent from the machine-readable world.”— Aravind Srinivas

Sarvam unveiled Sarvam Vision, India's first sovereign vision-language model: a 3-billion-parameter state-space (non-transformer) model trained end-to-end in India for English and 22 official Indian languages. It uses block-by-block OCR with a document-complexity assessment layer, runs on a single GPU, and claims state-of-the-art document AI results that surpass models 100 times its size. It matters because it targets the vast, undigitized Indic-language document space that mainstream models built on Common Crawl (where Indic languages are under 1%) largely ignore.

Sarvam vision-language-model OCR Indic-languages state-space-model

Watch / read the original source →