← All organizations

AI agent observability and evaluation

Braintrust

Braintrust builds an observability and evaluation platform for engineering and product teams developing AI applications and agents. Teams can inspect prompts, responses and tool calls, track latency and cost, compare models, and score outputs with code, language models or human reviewers. Production traces become versioned evaluation datasets for testing changes against real failures. Topics discovers patterns across production traces, while Loop generates prompts, scorers and datasets to help teams improve their agents.

Introduced in 2023, Braintrust was founded by CEO Ankur Goyal, who previously founded Impira and led Figma’s AI platform after Figma acquired it. Its engineering work includes Brainstore, a database built for agent traces that combines inverted indexes and columnar structures with object storage. Its phrase search uses filters based on consecutive three-word sequences to skip irrelevant storage segments, helping developers locate specific text in large trace collections. Hybrid deployment lets enterprises run the data plane on their own infrastructure.

Customers include Notion, Replit, Cloudflare, Ramp and Dropbox. In February 2026, Braintrust raised an $80 million Series B led by ICONIQ, with returning investors including Andreessen Horowitz and Greylock, at a reported $800 million valuation.

www.braintrust.dev

13 talks

Newest first

8 speakers at AIE

Affiliations reflect their AIE appearances, not necessarily current employment.

Start here

  1. Evals 101 — Doug Guthrie, Braintrust

    Start with Doug Guthrie for the components of an evaluation, baseline quality measurement, and SDK workflows, production logging, and online monitoring.

    Doug GuthrieAI Engineer World's Fair 2025

  2. Why building eval platforms is hard

    Phil Hetzel explains the data and systems engineering behind prompt comparisons, cross-functional experiments, and feedback between production observations and offline tests.

    Phil HetzelAI Engineer Europe 2026

  3. How Zapier Builds AI Products and Features With the Help of Braintrust

    The joint presentation by Olmo Maldonado and Ankur Goyal covers synthetic evaluation data, provider load testing, and regression detection, and discusses model accuracy–latency tradeoffs in Zapier’s AI products.

    Ankur Goyal · Olmo MaldonadoAI Engineer World's Fair 2024

  4. Why should anyone care about Evals?

    Braintrust founding engineer Manu Goyal introduces the conference’s Evals track by explaining why stronger model metrics alone cannot justify deploying AI systems.

    Manu GoyalAI Engineer World's Fair 2025

Messages from the stage

Evaluation as an experimentation loop

Manu Goyal presents evaluations as a way to test changes safely and quickly. Ankur Goyal connects their usefulness to rapid model adoption and coordinated improvements to datasets, prompts, and scoring.

Turning expert judgment into scoring

Phil Hetzel describes documenting experts’ rationales, translating them into scoring functions, and validating automated judges. The workshop led by Carlos Esteban and Doug examines confidence in LLM-as-a-judge results alongside subject-matter-expert review.

Evaluating the whole agent

Ameya Bhatawdekar identifies failure surfaces in node contracts, orchestration, and tool calling, arguing for assessment across distributions of outcomes. Phil Hetzel’s observability session adds trace search and clustering as ways to investigate agent behavior beyond uptime and latency.

Affiliations reflect each recorded session, not necessarily current employment.

Company sources · checked 2026-08-27