▶ Watch ↗AI Engineer World's Fair 202621:09
Bio, Work & Ideas
Conference affiliation: Head of Vision Models · Sarvam · 2026
Krishna Prasad Srinivasan works on multilingual document intelligence and vision-language models. His AI Engineer World’s Fair 2026 biography identified him as Head of Vision Models at Sarvam, responsible for the vision vertical’s models, research, and product. He led a lean team that trained Sarvam Vision, described in that biography as India’s first sovereign vision-language model: a three-billion-parameter state-space model that topped global OCR benchmarks at launch and led the Indic OCR benchmark across 22 languages.
In his presentation, Srinivasan explains the team’s approach to extracting coherent knowledge from documents in English and 22 Indian languages. It combines block-level optical character recognition with layout and reading-order harnesses, using a state-space backbone to reduce long-sequence inference costs while accepting some loss of recall. He describes a four-stage curriculum spanning text pretraining, image-text pretraining, supervised OCR fine-tuning, and reinforcement learning. Synthetic and real-document data pipelines support languages with limited labeled data, while machine-verifiable rewards test character accuracy, table structure, equations, and grammar. He also presents human-in-the-loop document workflows with confidence scores, block-level grounding, and agentic proofreading, and reports deployments to digitize more than 35 million pages.
Previously, Srinivasan was Tech Lead for AI at Microsoft Research, where he built multilingual education copilots and developed Indic translation models that outperformed commercial systems. Earlier, as a researcher at Harvard, he engineered a novel OCR architecture using contrastive learning that outperformed industry benchmarks on complex multilingual documents.
▶ Watch ↗AI Engineer World's Fair 202621:09