← All organizations

AI observability and evaluation

Arize

Arize builds observability and evaluation tools that help AI engineers debug and improve agents, chatbots and other AI applications. Phoenix lets developers inspect model calls, retrieval and tool use, score outputs, refine prompts and compare application changes on shared datasets. It runs locally or can be self-hosted. Arize AX provides managed infrastructure, production observability and online evaluations, while its Alyx assistant helps teams run evaluations and investigate failures.

Arize was founded in 2020 by Jason Lopatecki, CEO, and Aparna Dhinakaran, CPO. Its team created OpenInference, conventions and instrumentation built around OpenTelemetry that capture model invocations alongside retrieval and external tool calls. Supported by both Phoenix and AX, OpenInference also works with other OpenTelemetry-compatible backends, allowing teams to use the instrumentation beyond Arize’s products.

Arize reported over 2 million monthly Phoenix downloads in 2025 and, as of August 2026, 1 billion evaluations per year. It raised a $70 million Series C in 2025. On August 13, 2026, Dynatrace announced a definitive acquisition agreement valuing the transaction at $915 million, subject to customary adjustments. Closing remained subject to regulatory review and customary conditions; both founders were to join Dynatrace at closing, with Lopatecki continuing to lead Arize.

arize.com

14 talks

Newest first

8 speakers at AIE

Affiliations reflect their AIE appearances, not necessarily current employment.

Start here

  1. Ship Real Agents: Hands-On Evals for Agentic Applications

    Start with Laurie Voss's workshop to learn how to choose evaluators, separate capability tests from regression tests, and guard against agents gaming the measurements.

    Laurie VossAI Engineer Europe 2026

  2. The Unreasonable Effectiveness of Prompt Learning

    Follow a coding-agent experiment that turns unit-test results and judge explanations into revised CLAUDE.md or Cline rules without changing model weights.

    Aparna DhinakaranAI Engineer Code 2025

  3. Your Agent Is an Infinite Canvas

    Learn the distinctions among hosted MCP tools, interactive MCP Apps, and WebMCP browser tools, including the deployment and browser restrictions that shape their use.

    Rachel Lee Nabors (RL Nabors)AI Engineer Europe 2026

  4. LLM Observability, Evaluation, Experimentation Platform — Dat Ngo, Arize AI

    Arize AI's Dat Ngo explains how OpenTelemetry traces and spans expose nondeterministic agent behavior, regressions, and incorrectly ordered tool calls.

    Dat NgoAI Engineer Europe 2026

Messages from the stage

Evaluate the application and its execution

The talks distinguish application-specific tests from general model benchmarks, then examine tool arguments, execution order, and conversational behavior. Dat Ngo also discusses calibrating LLM judges against trusted datasets and using cheaper deterministic checks.

Turn failure evidence into reviewed improvements

Aparna Dhinakaran presents prompt learning that updates persistent instructions using explanations of failures. Jason Lopatecki's Signal session extends the improvement workflow to investigating production traces and preparing repository fixes for human review.

Design around context, compute, and interaction

Sally-Ann DeLucia examines unreliable summarization and retrieval from long-term memory. Rachel Lee Nabors (RL Nabors) discusses evaluating selected on-device replacements for cloud calls and exposing structured browser actions through agent-accessible interfaces.

Affiliations reflect each recorded session, not necessarily current employment.

Company sources · checked 2026-08-27