← All organizations

AI data development and evaluation

Snorkel AI

Snorkel AI develops expert training data, evaluation systems and runnable environments for frontier AI labs and enterprise teams. Snorkel Flow labels and manages training data programmatically; Snorkel Evaluate helps teams create benchmark datasets, build specialized evaluators and identify error patterns. Its Expert Data-as-a-Service supplies datasets for evaluation and post-training, while Snorkel Data Series packages datasets with difficulty tiers, rubrics and evaluation slices. The company also builds specialized agents for enterprise workflows, using task-specific checks to assess their performance.

Founded in 2019 out of the Stanford AI Lab, Snorkel AI’s co-founders are CEO Alexander Ratner, Christopher Ré, Paroma Varma, Braden Hancock and Henry Ehrenberg. Its research roots include the 2017 Snorkel system for weak supervision: users write labeling functions that express heuristics, and the system statistically denoises their potentially inaccurate, correlated outputs to create training data. That emphasis on defining and measuring data quality extends to collaborative benchmarks such as Senior SWE-bench, which tests coding agents on feature implementation, runtime debugging and adherence to codebase conventions.

In 2025, the company reported production users including BNY, Wayfair, Chubb and the U.S. Air Force, and work with seven of the top ten U.S. banks. It raised a $100 million Series D led by Addition that year at a $1.3 billion valuation, bringing total funding to $237 million.

snorkel.ai

Start here

  1. The Art & Science of Benchmarking Agents

    Start with Vincent Chen's framework to learn how scientific requirements and practical constraints guide benchmark selection and design.

    Vincent ChenAI Engineer Europe 2026

  2. From Agent Traces to Agent Simulations — Rustem Feyzkhanov, Snorkel AI

    Learn how Rustem Feyzkhanov connects production observability to repeatable experiments for long-horizon agent workflows.

    Rustem FeyzkhanovAI Engineer World's Fair 2026

  3. Stop Making Models Bigger, Make Them Behave — Kobie Crawford, Snorkel

    Use Kobie Crawford's financial-query comparison to understand why disciplined schema inspection matters when assessing tool use across model sizes.

    Kobie CrawfordAI Engineer Europe 2026

  4. Task Fidelity Scaling Laws — Kobie Crawford, Snorkel AI

    Snorkel AI developer advocate Kobie Crawford explains how task fidelity influences agent evaluation and reinforcement-learning outcomes.

    Kobie CrawfordAI Engineer Europe 2026

Messages from the stage

Recreate work, then verify outcomes

Rustem Feyzkhanov describes containerized environments that reconstruct tools, data, and user interactions. Oracle solutions, final-state verifiers, and LLM judges support regression checks and comparisons of model cost and latency.

Task quality shapes learning

Kobie Crawford distinguishes meaningful model failures from failures caused by faulty task specifications. His financial-analysis demonstration separately examines how expert-curated data and reinforcement learning can improve a small model's SQL tool use.

Benchmarks need rigor and adoption

Vincent Chen pairs expert validation, adversarial quality control, and meaningful headroom with practical questions of researcher adoption, policy adherence, and organizational context.

Affiliations reflect each recorded session, not necessarily current employment.

Company sources · checked 2026-08-28