Taming Rogue AI Agents with Observability-Driven Evaluation
AI Engineer World's Fair 2025 · 16:15
AI Evaluation and Observability
Galileo builds AI evaluation and observability software for developers and enterprise teams working on agents and retrieval-augmented generation applications. Teams can assemble evaluation datasets, capture expert annotations, test response quality, and diagnose failures in agent behavior. Its platform connects offline testing with production guardrails, using evaluation scores to control agent actions, tool access, and escalation paths.
Founded in 2021 by Vikram Chatterji, CEO, Atindriyo Sanyal, CPO, and Yash Sheth, CTO, Galileo develops specialized evaluation models and methods. Its Luna family uses small language models trained for evaluation tasks, including hallucination detection, security threats, and privacy checks. Luna compares response segments against context segments to identify unsupported content. Its ChainPoll method repeatedly asks an LLM to judge a response, producing both a hallucination assessment and an explanation.
Cisco completed its acquisition of Galileo Technologies on May 22, 2026, with plans to strengthen its Splunk Observability portfolio. Galileo offers a free developer tier, paid plans that scale with trace volume, and enterprise deployment through hosted services, private clouds, or on-premises infrastructure. In October 2024, the company reported that its enterprise-customer count had quadrupled since the start of that year and that it had added six Fortune 50 companies.
AI Engineer World's Fair 2025 · 16:15
Affiliations reflect their AIE appearances, not necessarily current employment.
Bennett describes specialized evaluation models and continuous human feedback as ways to keep automated assessment accountable when diagnosing low evaluation metrics.
Affiliations reflect each recorded session, not necessarily current employment.