Building Metrics That Actually Work — David Karam, Pi Labs
AI Engineer World's Fair 2025 · 40:28
AI application evaluation
Pi Labs developed tools to help AI application teams define and measure output quality. Its Pi Studio generated evaluation rubrics from product requirements, system prompts, or descriptions of an application. Teams could upload outputs as CSV or JSON, or paste them directly, then apply a rubric to test prompts, compare model versions, and validate applications before deployment.
Founded by David Karam and Achint Srivastava, Pi Labs focused on applying research breakthroughs to the application layer. Pi Studio used preference pairs and user feedback to calibrate evaluation criteria, translating relative judgments into application-specific metrics. As teams collected more examples and feedback, its agent refined those metrics. Evaluations could run through the studio or integrations with Arize AI, Braintrust, and Promptfoo.
Pi Labs was acquired within Microsoft in 2025, with Microsoft’s acquisition history identifying the transaction as a LinkedIn acquisition. The founders and their team joined Microsoft, marking the company’s transition from an independent AI evaluation startup.
AI Engineer World's Fair 2025 · 40:28
AI Engineer World's Fair 2025 · 20:22
Affiliations reflect their AIE appearances, not necessarily current employment.
Start with the workshop to learn how an evaluation copilot, spreadsheets, an SDK, and Google Colab fit into practical scoring workflows.
David KaramAI Engineer World's Fair 2025
Choose this talk to learn how a quality-engineering loop uses real queries and observed failures to guide incremental RAG improvements.
David KaramAI Engineer World's Fair 2025
The evaluation workshop combines natural-language questions with generated Python checks and synthetic examples, then calibrates scores against data.
The RAG session compares BM25, vector retrieval, custom embeddings, and cross-encoder reranking. It also considers domain-specific signals and multiple retrieval backends for ambiguous search intent.
Affiliations reflect each recorded session, not necessarily current employment.