← All organizations

Document intelligence and OCR

Datalab

Datalab trains document intelligence models that turn PDFs, images, spreadsheets and slides into structured data for AI research and enterprise workflows. Chandra converts documents into Markdown, HTML and JSON while preserving layout information. Its platform lets users extract fields with schemas, customize output with natural-language rules, separate combined documents and evaluate results. Applications include preparing AI training corpora and processing scientific papers, clinical-trial documents and insurance claims.

Founded in Brooklyn in 2024 by Vik Paruchuri and Sandy Kwon, Datalab had Paruchuri as CEO and Kwon as COO in 2025. Its engineering work includes Marker, a conversion pipeline that combines embedded document text with layout analysis and selective OCR through Surya. This approach can reuse readable PDF text while applying vision models to scanned or garbled content. Chandra extends the product family to layout-preserving recognition of handwriting, forms, tables and mathematics across more than 90 languages.

Datalab offers managed cloud, private-cloud and air-gapped deployments, allowing document processing within a customer's own network. Its managed service charges per page after a monthly free allowance, without a subscription or minimum. As of August 2026, the company reported that thousands of teams used Marker, Surya and Chandra.

www.datalab.to

1 talk

Newest first

1 speaker at AIE

Affiliations reflect their AIE appearances, not necessarily current employment.

Messages from the stage

Simpler tooling, fewer handoffs

Paruchuri advocated end-to-end ownership supported by AI-assisted tooling, reusable components, and a simple server-rendered stack, reducing the need for organizational specialization.

Affiliations reflect each recorded session, not necessarily current employment.

Company sources · checked 2026-08-28