← All organizations

AI application evaluation

Pi Labs

Pi Labs developed tools to help AI application teams define and measure output quality. Its Pi Studio generated evaluation rubrics from product requirements, system prompts, or descriptions of an application. Teams could upload outputs as CSV or JSON, or paste them directly, then apply a rubric to test prompts, compare model versions, and validate applications before deployment.

Founded by David Karam and Achint Srivastava, Pi Labs focused on applying research breakthroughs to the application layer. Pi Studio used preference pairs and user feedback to calibrate evaluation criteria, translating relative judgments into application-specific metrics. As teams collected more examples and feedback, its agent refined those metrics. Evaluations could run through the studio or integrations with Arize AI, Braintrust, and Promptfoo.

Pi Labs was acquired within Microsoft in 2025, with Microsoft’s acquisition history identifying the transaction as a LinkedIn acquisition. The founders and their team joined Microsoft, marking the company’s transition from an independent AI evaluation startup.

withpi.ai

2 talks

Newest first

1 speaker at AIE

Affiliations reflect their AIE appearances, not necessarily current employment.

Start here

  1. Building Metrics That Actually Work — David Karam, Pi Labs

    Start with the workshop to learn how an evaluation copilot, spreadsheets, an SDK, and Google Colab fit into practical scoring workflows.

    David KaramAI Engineer World's Fair 2025

  2. Layering every technique in RAG, one query at a time

    Choose this talk to learn how a quality-engineering loop uses real queries and observed failures to guide incremental RAG improvements.

    David KaramAI Engineer World's Fair 2025

Messages from the stage

Build and calibrate application-specific scores

The evaluation workshop combines natural-language questions with generated Python checks and synthetic examples, then calibrates scores against data.

Match retrieval techniques to relevance and cost

The RAG session compares BM25, vector retrieval, custom embeddings, and cross-encoder reranking. It also considers domain-specific signals and multiple retrieval backends for ambiguous search intent.

Affiliations reflect each recorded session, not necessarily current employment.

Company sources · checked 2026-08-28