Evaluating AI Search: A Practical Framework for Augmented AI Systems
Julia Neagu · Deanna Emery · Maitar Asher
AI Engineer World's Fair 2025 · 20:33
AI agent evaluation and reinforcement learning
Quotient developed software for enterprises to monitor, evaluate and improve AI agents in production. Its platform analyzed full agent traces to identify hallucinations, reasoning failures and incorrect tool use. It automatically grouped those findings into structured evaluation datasets and reward signals, giving teams material to monitor and fine-tune agents using failures encountered in real workflows.
Founded in 2023 by Julia Neagu and Freddie Vargus, Quotient grew out of their work on GitHub Copilot, where they held data science and machine learning engineering roles. Its technical approach connected failure diagnosis with continual learning: observations from production became inputs for evaluation and training. Customers included Wayfair and Fortune 500 businesses.
Databricks acquired Quotient in 2026, bringing its team into the company and announcing plans to embed its evaluation capabilities across Genie, Genie Code and Agent Bricks. The acquisition connected Quotient’s technology to products for conversational data analysis, data engineering and enterprise agent development. In March 2026, Neagu confirmed that the entire team had moved to Databricks and said the standalone platform would close, with development continuing within Databricks.
Julia Neagu · Deanna Emery · Maitar Asher
AI Engineer World's Fair 2025 · 20:33
AI Engineer World's Fair 2024 · 18:14
Affiliations reflect their AIE appearances, not necessarily current employment.
Start here for a comparison of naive and advanced RAG pipelines, alongside discussion of missing source information, scalability, and security.
Atita Arora · Deanna EmeryAI Engineer World's Fair 2024
Continue here for a concrete evaluation workflow using LangGraph for dataset generation and LangSmith for experiment tracking.
Julia Neagu · Deanna Emery · Maitar AsherAI Engineer World's Fair 2025
The joint AI search session contrasts static benchmarks with dynamically generated, multi-source datasets. It also examines reference-free answer completeness metrics alongside hallucination and observability concerns.
The joint RAG session connects domain-specific evaluation datasets and context-relevance metrics to iterative experimentation. Query rewriting, re-ranking, and context sizing are among the pipeline choices discussed.
Affiliations reflect each recorded session, not necessarily current employment. The AI search session is a joint discussion with Tavily; the RAG session is a joint discussion with Qdrant.