← All organizations

AI Agent Observability and Evaluation

HoneyHive

HoneyHive builds an AI agent observability and evaluation platform for engineers, enterprise platform teams, and domain experts. Users can inspect tool calls and agent trajectories, score behavior with LLMs, code, or human feedback, and turn production failures into regression datasets. Experiments compare agent versions, while monitoring connects evaluations to live traffic. The platform supports managed SaaS, hybrid deployment with traces and evaluations in the customer’s cloud, and fully self-hosted operation.

Founded in 2022 by Mohak Sharma, CEO, and Dhruv Singh, HoneyHive connects development testing with production monitoring. Its engineering work addresses inconsistent agent telemetry: mappings across OpenTelemetry GenAI, OpenInference, and OpenLLMetry let teams work with a normalized view across frameworks. A 2026 study examined 73,000 discovered field schemas from three production customers, retaining roughly 17,000 meaningful fields after cleaning. It documented how changing field paths can break evaluators and alerts, motivating more stable telemetry interfaces.

As of August 2026, HoneyHive reported supporting dozens of AI applications at Commonwealth Bank (CBA), whose production agents serve 17M+ consumers and 55,000 internal users. In 2025, it announced general availability and $7.4M in cumulative funding, comprising a $5.5M seed round led by Insight Partners and a $1.9M pre-seed led by Zero Prime Ventures.

www.honeyhive.ai

1 talk

Newest first

1 speaker at AIE

Affiliations reflect their AIE appearances, not necessarily current employment.

Messages from the stage

Measure agreement with human judgments

Sharma recommends checking automated evaluators against human judgments using F1 scores or correlation metrics, addressing the risk that uncalibrated LLM judges produce misleading scores.

Affiliations reflect each recorded session, not necessarily current employment.

Company sources · checked 2026-08-28