← All organizations

AI Observability and Evaluation

Traceloop

Traceloop builds observability and evaluation tools for teams developing LLM applications and AI agents. Its platform traces production calls, captures prompts, responses and latency, and checks output quality for faithfulness and relevance. Developers can train custom evaluators on their own examples, compare models and prompts, and run evaluations during development or on live traffic. Deployment options include cloud, on-premises and air-gapped environments.

Founded in 2023 by CEO Nir Gazit and CTO Gal Kleinman, Traceloop grew from their experience building software and machine-learning infrastructure at Google and Fiverr. Its engineering foundation is OpenLLMetry, an open-source framework built on OpenTelemetry that instruments LLM applications; Hub offers an OpenTelemetry-based gateway approach. In 2025, the company reported more than 500,000 OpenLLMetry downloads per month, with Cisco, Dynatrace and IBM using the framework. Miro used the commercial platform to monitor production performance and experiment with models.

ServiceNow completed its acquisition of Traceloop by May 2026, in a deal with a reported estimated value of $60–80 million; official terms were undisclosed. Traceloop technology supplies runtime AI-agent observability within ServiceNow’s AI Control Tower, connecting application monitoring with enterprise AI governance. In its acquisition announcement, Traceloop committed to keeping OpenLLMetry open source and stated its intention to continue selling Traceloop functionality.

www.traceloop.com

2 talks

Newest first

1 speaker at AIE

Affiliations reflect their AIE appearances, not necessarily current employment.

Start here

  1. OpenLLMetry is all you need

    Start here to understand how OpenTelemetry logs, metrics, traces, SDKs, and collectors fit together before examining Pinecone query and vector tracing.

    Nir GazitAI Engineer Summit 2025

  2. Prompt Engineering is Dead

    Choose this talk for a concrete walkthrough of building evaluation datasets and using an agent to regenerate prompts for a Chroma-backed RAG chatbot.

    Nir GazitAI Engineer World's Fair 2025

Messages from the stage

Portable tracing with data controls

The OpenLLMetry session covers instrumentation across model providers, vector databases, and frameworks. Gazit also explains collector-based filtering of sensitive data and portability across existing observability platforms.

Test criteria for prompt revision

Gazit argues for replacing manual prompt tuning with automated optimization. His chatbot example combines a ground-truth-based LLM judge with checks against expected facts to assess responses as prompts change.

Affiliations reflect each recorded session, not necessarily current employment.

Company sources · checked 2026-08-28