← All organizations

AI Agent Monitoring and Observability

Raindrop

Raindrop builds monitoring and debugging software for engineers running AI agents. It captures messages, tool calls, retries and errors, then surfaces silent failures such as hallucinations, loops and broken tools. Teams can investigate issues through its triage agent in Slack or the web and use Experiments to test model, prompt or tool changes against production traffic. Workshop, its free, open-source local debugger, lets developers inspect traces, replay calls and give coding agents access through MCP to generate evaluations from actual failures.

Raindrop’s founders are CEO Zubin Koticha, COO Alexis Gauba and CTO Ben Hylak. Its earlier product, Dawn, focused on analytics for AI products. Today, its Signals behavior classifiers combine code that selects relevant trace context with task-specific models and semantic reasoning. This approach can detect failures spread across tool calls and responses, rather than examining each step in isolation. The system concentrates reasoning in classifier construction and samples production classifications to identify drift and retune.

In August 2026, the company reported that its classification infrastructure evaluated over 20 billion traces per month, with a median classification time of 100 milliseconds. Customers include Vercel, Speak, Clay and Framer. Raindrop announced $15 million in funding in 2025 from Lightspeed and other investors, including Figma Ventures, Vercel Ventures and YC.

www.raindrop.ai

3 talks

Newest first

3 speakers at AIE

Affiliations reflect their AIE appearances, not necessarily current employment.

Start here

  1. Designing Agents (The Floor Is the Frontier)

    Start here for Hylak's argument that raising an agent's minimum reliability matters more than demonstrating impressive peak capability.

    Ben HylakAI Engineer World's Fair 2026

  2. Everything You Need To Know About Agent Observability

    Learn how monitoring can capture refusals, task failures, user frustration, and moderation concerns, with audience questions on integrations and trace volume.

    Danny Gollapalli · Ben Hylak · Zubin KotichaAI Engineer Europe 2026

  3. Building AI Products That Actually Work

    For a product perspective, the joint discussion connects continuous iteration with observed behavior; Sid Bendre also introduces Oleve's consumer-product perspective and Trellis framework.

    Ben Hylak · Sid BendreAI Engineer World's Fair 2025

Messages from the stage

User behavior as failure evidence

In the joint product discussion, Hylak questions overreliance on evaluations and language-model judges, highlighting explicit and implicit user feedback as signals of production problems.

Testing whether repairs hold

Hylak contrasts brittle tool-specific evaluations with issue detection, fix verification, simulation, and local code-aware testing. User oversight and harmful autonomous actions extend the discussion beyond task completion.

From detection to trace analysis

Zubin Koticha and Danny Gollapalli discuss regex and classifier-based detection, release experiments, and statistical relevance. The workshop moves into instrumentation, self-diagnostics, trace analysis, and tool failures.

Affiliations reflect each recorded session, not necessarily current employment. Building AI Products That Actually Work is a joint discussion with Sid Bendre of Oleve.

Company sources · checked 2026-08-28