From Signal to PR: Anatomy of a Self-Improving Agent
AI Engineer World's Fair 2026 · 20:36
AI observability and evaluation
Arize builds observability and evaluation tools that help AI engineers debug and improve agents, chatbots and other AI applications. Phoenix lets developers inspect model calls, retrieval and tool use, score outputs, refine prompts and compare application changes on shared datasets. It runs locally or can be self-hosted. Arize AX provides managed infrastructure, production observability and online evaluations, while its Alyx assistant helps teams run evaluations and investigate failures.
Arize was founded in 2020 by Jason Lopatecki, CEO, and Aparna Dhinakaran, CPO. Its team created OpenInference, conventions and instrumentation built around OpenTelemetry that capture model invocations alongside retrieval and external tool calls. Supported by both Phoenix and AX, OpenInference also works with other OpenTelemetry-compatible backends, allowing teams to use the instrumentation beyond Arize’s products.
Arize reported over 2 million monthly Phoenix downloads in 2025 and, as of August 2026, 1 billion evaluations per year. It raised a $70 million Series C in 2025. On August 13, 2026, Dynatrace announced a definitive acquisition agreement valuing the transaction at $915 million, subject to customary adjustments. Closing remained subject to regulatory review and customary conditions; both founders were to join Dynatrace at closing, with Lopatecki continuing to lead Arize.
AI Engineer World's Fair 2026 · 20:36
AI Engineer World's Fair 2026 · 30:52
AI Engineer World's Fair 2026 · 6:06
AI Engineer Europe 2026 · 16:17
AI Engineer Europe 2026 · 16:32
AI Engineer Europe 2026 · 2:04:18
AI Engineer World's Fair 2025 · 14:25
AI Engineer World's Fair 2025 · 24:46
AI Engineer World's Fair 2025 · 1:26:16
Aparna Dhinkaran · Aparna Dhinakaran
AI Engineer Summit 2025 · 15:28
Affiliations reflect their AIE appearances, not necessarily current employment.
Start with Laurie Voss's workshop to learn how to choose evaluators, separate capability tests from regression tests, and guard against agents gaming the measurements.
Laurie VossAI Engineer Europe 2026
Follow a coding-agent experiment that turns unit-test results and judge explanations into revised CLAUDE.md or Cline rules without changing model weights.
Aparna DhinakaranAI Engineer Code 2025
Learn the distinctions among hosted MCP tools, interactive MCP Apps, and WebMCP browser tools, including the deployment and browser restrictions that shape their use.
Rachel Lee Nabors (RL Nabors)AI Engineer Europe 2026
Arize AI's Dat Ngo explains how OpenTelemetry traces and spans expose nondeterministic agent behavior, regressions, and incorrectly ordered tool calls.
Dat NgoAI Engineer Europe 2026
The talks distinguish application-specific tests from general model benchmarks, then examine tool arguments, execution order, and conversational behavior. Dat Ngo also discusses calibrating LLM judges against trusted datasets and using cheaper deterministic checks.
Aparna Dhinakaran presents prompt learning that updates persistent instructions using explanations of failures. Jason Lopatecki's Signal session extends the improvement workflow to investigating production traces and preparing repository fixes for human review.
Sally-Ann DeLucia examines unreliable summarization and retrieval from long-term memory. Rachel Lee Nabors (RL Nabors) discusses evaluating selected on-device replacements for cloud calls and exposing structured browser actions through agent-accessible interfaces.
Affiliations reflect each recorded session, not necessarily current employment.