AI Agent Observability and Evaluation
HoneyHive
HoneyHive builds an AI agent observability and evaluation platform for engineers, enterprise platform teams, and domain experts. Users can inspect tool calls and agent trajectories, score behavior with LLMs, code, or human feedback, and turn production failures into regression datasets. Experiments compare agent versions, while monitoring connects evaluations to live traffic. The platform supports managed SaaS, hybrid deployment with traces and evaluations in the customer’s cloud, and fully self-hosted operation.
Founded in 2022 by Mohak Sharma, CEO, and Dhruv Singh, HoneyHive connects development testing with production monitoring. Its engineering work addresses inconsistent agent telemetry: mappings across OpenTelemetry GenAI, OpenInference, and OpenLLMetry let teams work with a normalized view across frameworks. A 2026 study examined 73,000 discovered field schemas from three production customers, retaining roughly 17,000 meaningful fields after cleaning. It documented how changing field paths can break evaluators and alerts, motivating more stable telemetry interfaces.
As of August 2026, HoneyHive reported supporting dozens of AI applications at Commonwealth Bank (CBA), whose production agents serve 17M+ consumers and 55,000 internal users. In 2025, it announced general availability and $7.4M in cumulative funding, comprising a $5.5M seed round led by Insight Partners and a $1.9M pre-seed led by Zero Prime Ventures.
1 talk
Newest first1 speaker at AIE
Affiliations reflect their AIE appearances, not necessarily current employment.
Messages from the stage
Measure agreement with human judgments
Sharma recommends checking automated evaluators against human judgments using F1 scores or correlation metrics, addressing the risk that uncalibrated LLM judges produce misleading scores.
Affiliations reflect each recorded session, not necessarily current employment.
Company sources · checked 2026-08-28
- HoneyHive - Agent Observability & Evaluation Platform
- HoneyHive Blog
- Announcing HoneyHive
- HoneyHive, when AI Eval is mission critical
- HoneyHive raises $7.4M to bring evals and observability to AI agents
- Introducing HoneyHive v2
- Standardizing AI Observability Before It Breaks: A Study on 73,000 Agent Schemas
