AI Agent Monitoring and Observability
Raindrop
Raindrop builds monitoring and debugging software for engineers running AI agents. It captures messages, tool calls, retries and errors, then surfaces silent failures such as hallucinations, loops and broken tools. Teams can investigate issues through its triage agent in Slack or the web and use Experiments to test model, prompt or tool changes against production traffic. Workshop, its free, open-source local debugger, lets developers inspect traces, replay calls and give coding agents access through MCP to generate evaluations from actual failures.
Raindrop’s founders are CEO Zubin Koticha, COO Alexis Gauba and CTO Ben Hylak. Its earlier product, Dawn, focused on analytics for AI products. Today, its Signals behavior classifiers combine code that selects relevant trace context with task-specific models and semantic reasoning. This approach can detect failures spread across tool calls and responses, rather than examining each step in isolation. The system concentrates reasoning in classifier construction and samples production classifications to identify drift and retune.
In August 2026, the company reported that its classification infrastructure evaluated over 20 billion traces per month, with a median classification time of 100 milliseconds. Customers include Vercel, Speak, Clay and Framer. Raindrop announced $15 million in funding in 2025 from Lightspeed and other investors, including Figma Ventures, Vercel Ventures and YC.
3 talks
Newest firstEverything You Need To Know About Agent Observability
Danny Gollapalli · Ben Hylak · Zubin Koticha
AI Engineer Europe 2026 · 50:25
3 speakers at AIE
Affiliations reflect their AIE appearances, not necessarily current employment.
Start here
- Designing Agents (The Floor Is the Frontier)
Start here for Hylak's argument that raising an agent's minimum reliability matters more than demonstrating impressive peak capability.
Ben HylakAI Engineer World's Fair 2026
- Everything You Need To Know About Agent Observability
Learn how monitoring can capture refusals, task failures, user frustration, and moderation concerns, with audience questions on integrations and trace volume.
Danny Gollapalli · Ben Hylak · Zubin KotichaAI Engineer Europe 2026
- Building AI Products That Actually Work
For a product perspective, the joint discussion connects continuous iteration with observed behavior; Sid Bendre also introduces Oleve's consumer-product perspective and Trellis framework.
Ben Hylak · Sid BendreAI Engineer World's Fair 2025
Messages from the stage
User behavior as failure evidence
In the joint product discussion, Hylak questions overreliance on evaluations and language-model judges, highlighting explicit and implicit user feedback as signals of production problems.
Testing whether repairs hold
Hylak contrasts brittle tool-specific evaluations with issue detection, fix verification, simulation, and local code-aware testing. User oversight and harmful autonomous actions extend the discussion beyond task completion.
From detection to trace analysis
Zubin Koticha and Danny Gollapalli discuss regex and classifier-based detection, release experiments, and statistical relevance. The workshop moves into instrumentation, self-diagnostics, trace analysis, and tool failures.
Affiliations reflect each recorded session, not necessarily current employment. Building AI Products That Actually Work is a joint discussion with Sid Bendre of Oleve.
Company sources · checked 2026-08-28
- Raindrop | AI Agent Monitoring & Observability
- Blog – Raindrop AI
- Raindrop: Sentry for AI Agents | Y Combinator
- Raindrop Raises $15M from Lightspeed to Monitor the Agent Era
- Introducing Raindrop Workshop
- Introducing Raindrop 2.0: Self-Healing Agents
- rd-signal-2: Frontier Classification at Production Scale


