AI Engineer World's Fair 2026
From Vibes to Production: Evaluating and Shipping AI Agents That Work 101 — Laurie Voss, Arize AI
Read the talk
From Vibes to Production: Evaluating and Shipping AI Agents That Work 101
Laurie Voss builds a financial research agent, discovers failures hidden behind plausible reports, and turns those failures into focused evaluators, better prompts, controlled experiments, and a production feedback loop.
From a talk by Laurie Voss
At a glance
Ideas worth remembering
Read traces before automating evaluation. Observed failures and domain requirements determine what the grader should measure.
Give a research-based judge the relevant research context. In this demo, correctness rejected all thirteen reports; context-based faithfulness classified six as faithful and seven as unfaithful.
Use focused evaluators for separate quality dimensions, and distinguish shipping guardrails from aspirational metrics.
Calibrate judges against carefully labeled examples, investigate disagreement, and reserve held-out cases to test generalization.
Save failing and passing cases, use explanations to guide repairs, and compare changed agents with the same inputs and evaluators while accounting for remaining execution variability.
Production evaluations turn new failures into regression cases. Give coding agents the product requirements and recurring failure themes so repairs improve behavior rather than merely fit the tests.
A plausible answer is only the beginning of a test
An AI feature answers three queries convincingly, so the team ships it. Then an unfamiliar input breaks it. Laurie Voss, Head of Developer Relations at Arize AI, opens this workshop with the gap between that familiar development habit and a repeatable test suite. “Three times is not a test suite.” The problem becomes especially visible when a prompt change improves tone while also making the application invent product features: a local improvement can change behavior across many other inputs.
The workshop uses a Colab notebook, the Claude Agent SDK, and Arize AX. AX collects execution traces, stores evaluation results, and supports production monitoring. The setup needs an Anthropic API key for generation and judging, an Arize API key to authenticate telemetry, and an Arize space ID to select the workspace. AX and the open-source Phoenix project share core observability ideas, but this exercise uses AX.
The first useful distinction is between recording behavior and judging it. A trace records an execution as a tree of spans. An LLM invocation, a tool call, and an enclosing agent turn can each be spans; a parent span contains the steps it coordinates. Each span records inputs, outputs, timing, token counts, and other metadata. “Traces tell you what happened. Evals tell you whether it was any good.”
An evaluation needs a definition of acceptable behavior that survives changes in wording. Two reports can answer the same question correctly without sharing an expected output string. That makes whole-response string assertions a poor general test, while leaving plenty of room for deterministic checks on specific properties. A useful suite combines three approaches:
- Code evaluations: Check properties such as valid JSON, required fields, length limits, or forbidden phrases. They are reproducible and inexpensive, but overly specific checks can reject valid variations.
- LLM judges: Apply a rubric to meaning, including factual accuracy, faithfulness to sources, or appropriate tone. They cost money and can make inconsistent or incorrect judgments.
- Human review: Use domain expertise to identify unfamiliar failures and calibrate automated judges. It takes time, and fatigue makes human labels fallible too.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Agents can fail between steps—or succeed by an unexpected route
An agent adds decisions between model calls: which tool to use, what arguments to supply, whether the returned information is relevant, and when to stop. Multi-agent systems add routing and handoffs. A plausible final answer can conceal a mistake much earlier in the chain. Voss illustrates this with a Tesla query that retrieves information about the inventor rather than the car company. Search succeeds, writing succeeds, and formatting succeeds; the report still answers the wrong question because the retrieved entity never matched the user's intent.
The reverse problem is a grader rejecting a valid solution. In Voss's account of τ²-bench, an agent was expected to be unable to reschedule an economy flight. The supplied policies allowed it to upgrade the ticket to first class, where rescheduling was permitted. The agent found that route and failed the expected-result test. This is why an evaluator should usually check the requested outcome and applicable constraints rather than insist on the designer's imagined sequence of tool calls.
Evaluation suites also change purpose as an agent improves. A capability evaluation asks whether the system can perform a task it has not yet mastered; a low score supplies a hill to climb. Once that behavior works reliably, the same task becomes a regression evaluation, protecting an ability the product already depends on. A growing suite records both the next improvement target and the behaviors that must survive it.
Results need to support debugging as well as comparison. The workshop uses a score and a readable label, often a binary zero or one. LLM judges also return explanations. A travel response that fails because it omits cost estimates gives a much more useful repair target than a bare failure score. If the same omission appears across fifty traces, the repeated explanation points toward a systematic problem. Those explanations can later guide a coding agent's changes.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Instrument a two-turn financial analyst before judging it
Instrumentation separates collection from analysis. The OpenInference instrumentation package captures Claude Agent SDK activity. The registration call configures OpenTelemetry export to AX using the space ID, API key, and project name. OpenInference adds AI-specific attributes such as prompt and completion text, model identity, token counts, and invoked tools. The SDK instrumenter then captures model and tool calls as spans. A separate Arize client reads the collected data back into the notebook.
The notebook sets batch=False so spans arrive promptly while participants inspect a run. Batching is useful for efficient production export, but waiting for a batch would make this interactive exercise harder to follow. Voss also introduces Arize Skills, which can help coding agents instrument applications, run evaluations, and create datasets; the manual setup teaches what those automated operations actually do.
The financial analyst accepts a stock ticker and a focus area. Its first turn researches current information through web tools; its second compiles that research into a report. Claude Haiku supplies a fast, inexpensive worker, and the deliberately basic research and writing prompts leave room for mistakes. An enclosing OpenTelemetry span groups both turns into one execution, so the UI can show their relationship rather than two disconnected top-level calls.
The initial request asks for Tesla's financial performance and growth outlook. The agent searches, reads the results, decides whether it has enough information, and searches again if it does not. The recorded example makes four web searches before writing. That stopping decision is part of the agent's behavior: the application does not prescribe a fixed number of searches. Repeating the request can change the searches, sources, and report.
The trace immediately reveals a surprise. Asked to “write a report,” the agent attempts to create a Markdown file. In this configured notebook execution, the write fails, yet the agent still returns output. The Tesla report looks promising—with financial tables, growth drivers, risks, price targets, and a conclusion—but the trace exposes unnecessary failed work beneath it. Observability makes the distinction visible: receiving a report does not establish that every tool operation succeeded.
These same spans become evaluation inputs. The request describes what the agent was supposed to do, the returned report is the object being graded, and intermediate outputs can supply context. Capturing those pieces now makes it possible to ask more precise questions later, including whether the writer stayed within the research it actually received.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Read the reports, then decide what success requires
The notebook adds twelve test queries to the initial Tesla run, producing thirteen complete executions. Diversity comes from changing the task as well as the ticker: Apple versus Microsoft introduces comparison; Amazon profitability versus AWS profitability changes the analytical focus; a Coca-Cola dividend-yield question asks for different reasoning from a growth-stock report. The point is to exercise the kinds of distinctions users will ask the agent to handle.
Before writing an evaluator, read a dozen or more executions end to end. Ask what was requested, what was returned, and what specifically failed. Otherwise, automation tends to measure whatever is easy to count. For this analyst, the success criteria are concrete: reference the correct ticker, include real recent financial data, offer actionable recommendations, and distinguish forward-looking analysis from historical summary. Product managers, QA, support staff, and other domain experts help define that bar.
Synthetic queries provide a starting point before real traffic exists. Vary wording, specificity, and complexity: a formal Tesla research request and “Yo, is Tesla a buy right now” may express similar intent through very different inputs. Domain experts should review the set because generated queries tend to favor obvious phrasings. Include nonexistent tickers, multipart questions, and attempts to push the agent outside its intended behavior. Replace imagined cases with actual production traffic as it becomes available.
Reading the thirteen outputs exposes a failure that a polished sample concealed. Most reports contain 3,000–7,000 characters inline, including summaries, tables, and recommendations. Three instead return a short summary or a claim that the report was saved to disk. In the Microsoft example, repeated searches precede an attempted file write, and the final response points to a report that does not exist. The user receives neither the full analysis nor a usable artifact. This is a delivery failure, even if substantial research and writing happened internally.
Error analysis has two useful passes:
- Open coding: Write down the observed problem in ordinary language—vague recommendation, missing report, wrong ticker, unsupported number—without forcing it into a category immediately.
- Axial coding: Group similar observations afterward and investigate causes. Bad search results suggest a retrieval problem; appropriate data followed by an unsupported conclusion suggests a reasoning problem; answering outside the product's remit suggests a scope problem.
The cause matters because it determines the repair. A retrieval change and a writing-prompt change address different defects.
Voss's displayed category-frequency table uses random demo codes, so its rankings do not establish which failures dominate these reports. The practical prioritization rule still applies: weigh severity alongside frequency. A rare dangerous response can deserve attention before a frequent minor omission. Evaluation coverage should also come in layers. Code checks, semantic judges, and human review leave different gaps—the Swiss cheese analogy ends with the appropriately informal instruction, “stack the cheese.”
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A cheap ticker check catches a subtle loss of focus
The first automated check asks whether the requested stock ticker appears in the report. Its Python logic identifies uppercase words, removes known non-ticker acronyms, and checks the remaining candidates against the requested ticker. The evaluator accepts query and report and returns a score and label. This is a narrow requirement that does not need another model.
AX can run the function through code or its UI. In the UI, a representative trace supplies the mapping from recorded fields to evaluator parameters: select the request for query and the returned text for report. Scope matters because a full-report check needs the relevant execution output, while a check on one tool invocation needs that span's output. Offline results can be logged back to AX as annotations, making failing runs filterable alongside their traces.
Twelve reports pass and one fails. The failure is the Amazon query focused on AWS performance and profitability. Research produces material about AWS, and the writing turn stays entirely within that subsidiary-level discussion. The final report never mentions Amazon's stock ticker. A tiny deterministic check exposes the observable consequence of that narrowing: the requested company identity disappears from the answer. It does not establish that the AWS facts are wrong or that the whole analysis is poor; it identifies one unmet requirement.
Deterministic grading can do more than inspect strings. A code evaluator can check structure or compare a value with a database or API result. The important property is reproducible grading for the same evaluation inputs. For most outcome checks, avoid requiring a particular tool order: the ticker test cares that the answer retains the requested ticker, regardless of how the agent found its information. Intermediate checks still have a role when tool choice, arguments, or stopping behavior are themselves requirements.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Give the judge the research that the writer had
Semantic grading separates three ingredients: the judge model, the rubric, and the examples. Keeping them separate allows the same rubric to be tested with another model, or a revised rubric to be tested against the same examples. The workshop chooses Claude Sonnet as a more capable judge for Haiku's reports and first runs a built-in correctness evaluator.
Every report receives zero: correctness rejects all thirteen. The explanations reveal an information mismatch. The financial analyst researched live financial information, while this judge configuration grades the request and answer using the model's own knowledge without the collected research. Voss attributes the rejections to the judge treating recent information as beyond what it knows. The zero score therefore fails to distinguish useful reports from bad ones.
Faithfulness changes both the information supplied and the question asked. Extract the first turn's research output from the trace, attach it as a context column to the corresponding report, and evaluate the request, report, and context together. Now the judge can ask whether the writer's claims are supported by the material it received. What changes in this data flow when the judge gets context? The diagram shows the same research output feeding both writing and evaluation.
The resulting split is six faithful reports and seven unfaithful ones. Those are the judge's assessments of support in the supplied research, not independent verification of every financial fact or source. Unlike the blanket rejection, the split supplies a usable capability target: improve how the writer stays grounded. Correctness and faithfulness answer different questions, and this live-research application needs the judge to see the research context.
The requested financial analysis.
The faithfulness evaluator receives the first turn's research alongside the final report, so it can assess whether the report stays within that material.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Turn “actionable” into observable criteria
A grounded summary can still leave the reader without a decision. This analyst is supposed to offer recommendations—buy, sell, or hold—and forward-looking analysis, so it needs a domain-specific actionability evaluator. “Helpful and accurate” does not tell the judge what to look for. The rubric must describe evidence a reader can actually identify in the report.
The custom rubric separates several jobs:
- Task context: Tell the judge that it is evaluating a financial report for useful financial analysis. The relevant context matters more than an elaborate expert persona.
- Pass and fail criteria: Require specific recommendations and distinguish forward-looking analysis from a historical recap. Derive criteria from observed failures rather than an imagined ideal response.
- Tagged inputs: Delimit the user query, financial report, and examples with XML tags so their roles are clear.
- Configured output choices: Set
actionableandnot actionable, with scores one and zero, in evaluator configuration. AX uses a tool to collect the classification, so the rubric need not also demand a particular free-text output string.
Examples make the distinction concrete. The actionable example includes percentage revenue growth, a PE ratio, a quantified risk, and a specific recommendation. The non-actionable example can be factually acceptable yet offer only the vague suggestion that investors consider various factors. Voss favors examples because they clarify borderline judgments, while acknowledging the tradeoff: rigid examples can encourage overfitting and reduce variation in what the evaluator accepts.
Binary labels keep the decision easier to specify. A one-to-five scale requires separate rules explaining the difference between adjacent scores; a pass/fail criterion already describes a meaningful distinction. Voss also recommends asking for an explanation before the score, both to help the judge apply its criteria and to make disagreements inspectable.
The evaluator can run in the notebook or continuously against incoming traces. Its configuration combines a name, model, template, input mappings, and label choices. The resulting annotations sit beside ticker and faithfulness scores. Filtering for eval.actionability.label equal to not actionable isolates the cases to investigate and later rerun with changed prompts.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Separate diagnostic scores from shipping decisions
An evaluator prompt is another application component. Small wording changes can alter its judgments, so version the rubric and test it on examples with known labels. The “God evaluator”—one prompt grading accuracy, tone, completeness, policy compliance, and formatting—makes calibration difficult. A failure leaves the team unsure what to repair. Separate dimensions preserve a useful diagnosis: ticker identity, faithfulness, and actionability each name a different concern.
Scores also need different operational consequences. A fabricated stock price is a guardrail failure that can block shipping. A preference for complementary investment recommendations is an aspirational metric. Treating them identically either weakens an essential requirement or makes an optional improvement unnecessarily disruptive. In production, that distinction determines whether a regression should trigger an urgent alert or appear in a periodic quality report.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Check the judge's homework with human labels
An actionability judge is a classifier: it reads a report and predicts one of two labels. Human annotations provide a comparison set. In AX, a domain expert can read a span, select a Human Actionable label, and move to the next report without writing code. Exporting those spans returns the annotations as columns, allowing the notebook to compare human and model judgments on the same rows.
A trustworthy comparison set needs specific labeling rules, solvable tasks, and reference outputs that pass the graders. Test behavior in both directions: an agent should search when outside information is needed and avoid searching when it is unnecessary. Otherwise, a suite that rewards only tool use can teach the wrong habit. Voss suggests using 75% of labeled examples to develop the judge and holding out 25% to test whether revised criteria generalize.
The demo comparison agrees on six of thirteen reports, or 46%, because Voss deliberately assigned random human labels to create disagreements. That number is not a measurement of judge accuracy. With genuine annotations, disagreement prompts a close reading of the report, rubric, and judge's explanation. Either label can be wrong, or the rule can be ambiguous. Tightening “includes forward-looking analysis” to “includes forward-looking analysis with specific recommendations or guidance” resolves one such ambiguity: discussing the future alone does not make a report actionable.
Treat not actionable as the positive class when interpreting error-detection metrics:
- Precision: Of the reports the judge flags as not actionable, how many really fail that criterion?
- Recall: Of all reports that really are not actionable, how many does the judge catch?
Voss generally favors recall when missing a defect is more costly than investigating a false alarm, while recognizing domains where precision matters more. Small samples make both estimates unstable; they can guide early iteration without establishing shipping confidence.
Several biases can make the judge reward presentation rather than quality:
- Position bias: In comparisons, the order of candidate answers can influence preference.
- Length bias: Extra text can receive a better score even when it adds filler.
- Confidence bias: Assertive wording can make unsupported claims look credible.
- Self-preference bias: A model may favor outputs resembling its own generation. Voss uses different worker and judge models and recommends considering different providers as well.
Human judgment is a fallible reference too. Some disagreement is expected, especially on subjective criteria, and should lead to investigation rather than automatic distrust. The useful test is whether failures make sense: a failing trace should reveal what the agent got wrong and why. If the response satisfies the intended requirement but the grader rejects it, revisit the evaluator—the same lesson as the unexpectedly valid flight-rescheduling route.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Use explanations to propose a repair, then rerun the same cases
A prompt edit is only a proposed fix until the agent runs again. AX turns filtered failing traces into a dataset—in the workshop, AIEWF Financial Demo Fails—so iteration can concentrate on known weaknesses. Save passing examples too. The failure set makes focused experiments cheaper; the passing set checks that the repair preserves working behavior. As real traffic arrives, both sets should grow from observed cases.
The repair input combines requirements with failing traces and their explanations. A missing-recommendation explanation suggests making recommendations explicit in the writing prompt. A risk presented without evidence suggests requiring supporting evidence. The workshop uses Claude to produce revised research and writing prompts from this feedback, then substitutes those prompts into the otherwise unchanged agent. The changes have concrete reasons, but the prompt-improvement step remains a fallible LLM operation.
An experiment combines a dataset, a task, and evaluators. The task is a Python function that accepts an example and returns an output. It can run the whole agent, one component, or an API call. Here it reruns the complete financial analyst, then applies the same evaluator to the new reports. Each result retains the output and score, making the before-and-after difference inspectable at the example level.
Previously non-actionable cases become actionable, and Voss reports a displayed result of 100% after the prompt change. This is a rehearsed workshop result on a small selected dataset, not a demonstrated production success rate. Holding inputs and evaluators fixed removes variation from case selection, but live search and agent decisions still vary between runs. The comparison supports investigating the prompt improvement; it does not isolate the prompt from every source of randomness.
The useful observable change is that the same selected requests now produce reports accepted by the actionability rubric. The causal development is straightforward: identify failing outputs, read explanations, translate recurring omissions into stronger instructions, rerun the agent, and apply the same criteria. Real projects repeat this process rather than expect one perfect edit. The score tells the team whether to keep investigating; the explanations tell it where to look next.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Fix the data before reaching for model settings
Voss offers 12–20 examples as a directional workshop-scale signal and 200–400 as a practical target for shipping decisions. These are starting ranges rather than guarantees: the task and consequences of missed failures still determine what confidence is needed. More examples also have diminishing returns; the workshop notes the familiar sampling relationship that halving a margin of error requires roughly four times the sample size.
The proposed order of investment follows the source of the failure. Fix data quality first: wrong sources or stale knowledge cannot be reliably repaired by better wording. Next improve prompts with examples, explicit instructions, and constraints. Then consider a more capable model, accepting additional latency and cost. Temperature and top_p are easy to adjust, but Voss places them last because they usually address less than a clear data or task-specification fix.
Eval-driven development moves the definition of done ahead of implementation. If a refund agent must verify identity before processing a refund, write the evaluation for that requirement first and build until it passes. Here, sequence is consequential because verification must precede the action; it is different from prescribing an arbitrary research-tool order. People closest to product requirements can describe the intended behavior in natural language, which the team then turns into a testable rubric.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Production traffic supplies the next regression case
Passing a development experiment does not finish the work. Online evaluations apply the same ticker, faithfulness, and actionability checks to new production traces. Choose scope to match the question: a span-level evaluator can inspect a particular model or tool call; a trace-level evaluator can inspect the full execution tree. Sample expensive LLM evaluations—for example, 1% or 10% of traffic—to obtain a directional signal without paying to judge every run.
AX also includes Alyx, an agent that can turn a plain-language evaluation request into rubric and UI configuration. That helps domain experts contribute without implementing the plumbing themselves. Automation does not replace defining success: it moves the description into an evaluator that still needs the calibration and testing developed earlier.
How does a newly discovered production failure become protection for the next release? Online scores feed monitors, monitors prompt investigation, and the team saves relevant failing traces as regression examples. After a repair, an experiment checks the candidate before shipping it. The cycle below makes that return path visible: production creates new cases, while accumulated cases constrain subsequent changes.
The dataset becomes increasingly specific to the application: its users, edge cases, failure modes, and quality bar. Every caught failure can add a test that did not exist before. That compounding value requires continued work on evaluators and monitoring as new traffic reveals behaviors the original synthetic queries never anticipated.
New user traffic produces execution traces.
Labels direct monitoring; traces and explanations support diagnosis; saved cases let experiments check a proposed repair before the loop returns to production.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Give a coding agent requirements and recurring failure themes
The workshop edits two prompts, but real agent behavior can depend on prompts across many files, retrieval settings, and tool definitions. A repository-aware coding agent can work across those components. Export a batch of failing traces with their explanations and use them as repair context; Arize Skills can handle the AX operations needed to retrieve that data.
Two instructions keep the proposed repairs aimed at the product:
- Supply requirements as well as failures: Asking only for passing scores invites shortcuts, including fitting directly to test cases. Specify the behavior users need, with evaluations measuring whether the implementation provides it.
- Find themes across traces: Ten failures with one cause deserve a shared repair. Chasing each case separately can produce unnecessary code that solves examples while leaving the recurring problem intact.
Experiments remain the check on those changes. The coding agent proposes fixes from explanations; it does not get to make its own success claim sufficient. New production traffic returns to the same evaluations, revealing both remaining defects and unfamiliar ones. This closes the improvement loop across execution, diagnosis, implementation, and measurement.
The ending recommends starting smaller than the full workflow: instrument one agent, read real traces, and add one inexpensive deterministic check for an important requirement. Then add the semantic evaluator the application actually needs, inspect its results, and expand from recurring failures. Voss's closing judgment is that fifteen minutes reading real outputs can teach more than an hour building dashboards. The useful habit is routine attention to what the system did—and a growing set of tests that catches the next regression before users do.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
Related talks
- Beating RL With Reflection: GEPA and Optimize Anything
Extends the workshop's explanation-guided prompt repair into reflective search over prompts, agent programs, and repository skills, including learning an evaluator from human annotations.
Read the complete timestamped transcript
- 0:12
Hello, everyone. Uh, I have a packed room for a two-hour workshop, which is amazing. Welcome to AI Engineer. I know no one's officially welcomed you, but, uh, welcome to AI Engineer. Uh, this venue is amazing. This crowd is really quite surprising, and I have, uh, a ton of content for you today. Uh, my name is Laurie Voss. I am head of developer relations at Arize AI. Uh, some of you may remember me from when I used to
- 0:43
co-found npm Inc. Uh, so you may remember me from, like, JavaScript days. Uh, but these days, I talk about AI, uh, and how to test it, uh, and how to make it work. So what are we gonna be covering today? Uh, we've got a really good stretch of time together, uh, so we're gonna cover a lot of ground, and I'm gonna start with fundamentals. That's why this is called a 101 session. Uh, we're gonna start with what evals are, why you need them, uh, and why agents make evaluation harder than simple LLM
- 1:12
applications. Uh, then we're gonna set up tracing in Arize AX, uh, which is how you capture the raw data you need to run evals in the first place. Uh, we're gonna build a simple AI agent, uh, with the Claude Agent SDK. We're going to run it, and we're gonna look at, uh, the traces it produces. Can I get a show of hands if you've already built an agent? All right. Great. I am going to spend no time at all on how the agent works because I assume everybody knows how an agent works, and you're here to learn how to
- 1:42
evaluate them. Uh, once we've got some data, uh, we're gonna do something that a lot of tutorials skip, which is we're gonna actually look at the data. Uh, we're gonna read our traces. We're gonna categorize what went wrong, uh, and we're gonna figure out what to measure before we write a single eval. Then we're going to write three kinds of evals. There are code evals, which are simple deterministic checks, uh, very similar to unit tests. Then we're gonna use built-in LLM evals, uh, like faithfulness, uh, where a second
- 2:12
LLM judges the, uh, output of your first one, uh, and we're going to do a custom eval, uh, LLM as a judge from scratch. Uh, and we're also going to test whether our judges are judging correctly, a process called meta-evaluation. Then we're gonna finish with data sets and experiments, which is how you iterate on your agent, uh, and actually measure whether it's getting better, uh, including we're going to use Claude Code to automatically improve our app in response to feedback from our evals.
- 2:43
So first things first, uh, get your notebook. Uh, I messed up earlier and didn't have permission set correctly, so if you scanned this earlier and it didn't work, it works now. Uh, grab this notebook and make a copy for yourself. That is where we're gonna be living all day.
- 3:04
And I'm gonna wait for that to happen because otherwise I'm gonna have to keep swapping back to this slide. All right.
- 3:14
Most of those are down. No, no, I'm still seeing people taking pictures.
- 3:19
All right, cool. We're also going to need a certain amount of AI infrastructure. I'm going to be using Anthropic and Claude Code today. Uh, Arize AX works with OpenAI. It works with every other, uh, lab provider, but I had to pick one in order to do a workshop about it, so I've used Claude. So this is going-- this notebook's going to expect you to have, uh, an Anthropic API key, which you can get from console.anthropic.com. If not, you can, if you are very clever, do a fast swap to
- 3:49
OpenAI. Um, you're also going to need an Arize account. Go to arize.com and start a free trial. Um, from there you will need an Arize space ID and an Arize API key, both of which you can get from the Settings page. Uh, my lovely assistant, Dat, is here. Uh, that's him in the corner. If you run into any problems getting your API keys or anything like that, uh, just flag him down and he will run over to you and
- 4:19
help you figure it out, which is, uh, incredibly nice of him. Thank you, Dat. Um,
- 4:29
yes. So, uh, who, who had all that stuff already? Anybody already got an Arize AX account?
- 4:38
All right. I'm gonna let you hang out for this-- I'm gonna let you set that one up for a little while.
- 4:48
The way you do it is you go to arize.com
- 4:58
and you hit... Oops. It's sent me to the docs 'cause I always live in the docs. You go to arize.com, thank you, uh, and you hit the big pink Get Started button.
- 5:12
And then you sign up from here.
- 5:16
Do not accidentally sign up for Phoenix. Phoenix is our lovely open source project which also has, uh, a hosted service underneath Phoenix. Uh, sometimes when you Google Arize, you get sent to Phoenix instead. Don't sign up for Phoenix. You're doing AX today.
- 5:34
All right. While you are signing up, uh, I'm gonna say a little bit of more, bit more about what AX is. It is an observability and evaluations platform. It captures traces from your AI application. It stores and runs your evaluations, and it watches your app in production and alerts you when something goes wrong. Uh, by default, we host it for you. That is why we're using it today, because you don't have to, like, run some software to make it run. Uh, so there's no infrastructure to manage. Uh, though if you are from an enterprise that is super precious about your
- 6:04
data and you're like, "No, we need to run it ourselves," you can absolutely do that if you want to. Um, and like I said, uh, if you've used Arize Phoenix, which is our open source tool, a lot of this will feel familiar. Uh, they have the same core ideas. One is, uh, built for open source accessibility, and the other one is built for, uh, enterprise with production and teams and things like that.
- 6:30
So how are we doing on signing up for keys right now?
- 6:35
Wi-Fi sucks.
- 6:36
The Wi-Fi sucks. That is predictable really, isn't it?
- 6:49
All right. Is anybody, like, desperately in need of assistance right now?
- 6:55
The main issue is Wi-Fi that is not quality. Like, where, where there's hotspot, we have Wi-Fi that's works, but this Wi-Fi is not helping.
- 7:05
All right.
- 7:09
The stream labs is not working, so everyone needs to switch to their phones.
- 7:12
Okay. Well, the advice is to switch to your phones as a hotspot. I apologize for the Wi-Fi. I'm really worried about my two PM workshop where I have a Git repo that I know is two hundred megabytes large, and I'm like, "No one is gonna get the repo 'cause it's too big."
- 7:29
Clone it now.
- 7:29
Yeah, start cloning that repo now. Anyway, uh, let's talk about these, these fundamentals. There's a lot of theory in this workshop, which gives you time to, uh, catch up. Um, here is the mental model. Evals are testing for AI, and traces are logs for EI-- AI. Why we had to come up with two completely new names for those things, I don't really know. Uh, you write tests for your code. Given this input, I expect this output. Uh, evals do the same thing, but for AI outputs
- 7:59
where the outputs are never exactly the same twice. Just as logs record what your server did at runtime, uh, traces record what your AI did. So every agent call, every tool call, every LLM invocation with the inputs and outputs at each step. The building blocks of a trace are called spans. Uh, in this little screenshot, you can see a trace in Arize AX. It's made up of a bunch of spans. It's got this tree structure. Uh, and each span represents one step in the
- 8:28
execution. An LLM call is a span. A tool call is a span. A full agent turn is a span that contains other spans inside of it. Each span records its input, its output, uh, timing, token counts, and a whole bunch of metadata, uh, everything that you need to understand what happened at that step. Traces tell you what happened. Evals tell you whether it was any good. That is the model for-- mental model for today. This is not rocket science.
- 8:54
But let me talk about why we even need evals, uh, and that is the vibes problem. Uh, AI features usually get shipped by building an AI feature, running a few queries, saying, "Does this look right? Yeah, it looks good," and then you ship it. That is by far the most popular method of shipping AI applications today. Uh, and the problem is that then it fails on inputs that you didn't test. Uh, and the thing is, it probably looks good on your twe- test queries, but the problem is that you ran it three times. Three times is
- 9:24
not a test suite. Uh, the usual fix for, uh, flaky software is unit tests, and they don't work here because the output is non-deterministic. Uh, the same prompt produces different texts every run, so two different responses might both be correct, so there's no expected string that you can assert against. So teams fall back on human review. They watch it run, it looks fine, and they ship it. Uh, that doesn't scale, it doesn't catch regressions, and most importantly, it doesn't run in CI.
- 9:54
Here's some things you can't do without evals. Uh, if you change your system prompt to fix a tone issue, then the tone would get better, uh, but now your, uh, bot is hallucinating product features. Without evals, you wouldn't catch that, uh, until a user reports it. With a faithfulness eval, you would have seen the spike before it shipped, uh, and every prompt potentially affects every kind of input that users send. So you can be changing the prompt for one thing and have a completely unexpected effect at some totally different part of your application because the prompt is
- 10:24
just a huge block of text that you're sending to your AI. Uh, you can't manually test every combination, so without evals, you're playing Whac-A-Mole. You fix one problem, and you break something somewhere else. Evals instead give you a number that you can track and compare and act on. You also can't switch models without evals, and this matters more and more because this field moves so fast. Uh, new models drop every few months. Uh, without eval, switching models means weeks of manual testing. Is it better?
- 10:54
Is it worse? Did something get h-- Did something break? You don't know. With evals, you run the suite, you compare the scores, uh, and you know within hours, and that is the difference between flying by the seat of your pants and flying by instruments. And this is not theoretical. Uh, real teams, Descript, Bolt, Anthropic's own Claude Code, all follow the same arc of shipping on vibes and then discovering that vibes don't scale, uh, and then, uh, building evals. And that is the same arc that we're gonna follow in today's
- 11:24
workshop.
- 11:27
As I mentioned earlier, there are two broad types of evals. Code evals are deterministic functions. They run in milliseconds, and they cost nothing, and they give you an immediate, unambiguous answer like, did the output parse as JSON? Is the output less than five hundred tokens? Did it mention the thing that I asked about? The big advantage of code evals is that they are fast and cheap and totally reproducible. They get the same, they get the same input, and they give the same answer every time. That makes them very similar to old-style unit tests. The downside is that, like
- 11:56
old-style unit tests, they can be brittle. Uh, we have some strategies that can help avoid that, and I'm gonna be covering those later. Um, because two perfectly correct answers might be worded completely differently, and you have to adjust your expectations of unit tests to be able to handle that. The second type of eval is an LLM as a judge eval. These use a second LLM to grade your outputs against a rubric. Rubric being yet another word that we made up, uh, that just means rules. It is your set of rules that you put in
- 12:26
your prompt about what counts as good. Um, LLM evals are much more flexible. Uh, they can handle the kinds of questions that code evals can't answer. Uh, is this response, uh, factually accurate? Did it stay faithful to the source material? Is the tone right, uh, for a customer service context? The strength of LLM judges, uh, is that they understand the meaning, uh, not just the strings that are in your output. But LLM judges have trade-offs. The weakness is that they cost money to run, and they are
- 12:56
non-deterministic themselves. Uh, so they can be wrong, which means that you need to calibrate them against human judgment, and we're gonna do that later today. Um, there's also a third type of eval, which is human evaluation. Uh, a domain expert can read your output and grade it, uh, and that is the gold standard for quality. Uh, but it is slow and expensive, and humans get tired, which is why we can't use it all the time. Uh, fun fact that's not very fun, human annotators miss up to fifty percent of
- 13:26
defects due to fatigue. So they're not perfect either. So when you build an LLM as a judge, and you're like, "Well, it doesn't always catch my errors," remember that if you were judging with a human, the human would also not always catch your errors. Uh, but these are complementary approaches. Most real applications are gonna use a, uh, a collection of code evals, LLM as a judge, and human evaluation. So when do you use which? Use code evals when the answer is deterministic. So format validation,
- 13:56
length limits, forbidden phrases, uh, required fields. Use LLM judges when you need semantic understanding. A correctness eval asks, did it answer the question accurately? A faithfulness eval asks, did it stick to the source documents without hallucinating? Uh, and you should keep humans in the loop for failure modes that you haven't seen before, uh, and to verify that your LLM judges are actually judging correctly. LLM judges can be wrong. Uh, judging them against human judgment is how you know whether or not you can trust them.
- 14:27
All of that is true of any application that uses LLMs, uh, but agents make things even harder in some ways that are, uh, worth calling out. You can think of it as a ladder of complexity. A single LLM call is relatively simple. You get an input, you get an output. When you're talking about an agent, an agent is making a series of tool calls and a series of decisions, which means that at every single step in the, uh, chain, it is-- the chances that it has made a mistake are getting higher, and so the chances that it will go off the rails, uh, are going to
- 14:57
get, uh, greater. So you have to judge the intermediate steps as well. You have to, uh, check whether your agent is using the right tool. Did it pass the right arguments? Did it know when to stop? Getting stuck in loops is a real failure case for agents. Uh, and then multi-agent systems, if you are building those, add yet another level of complexity. They add handoffs. One agent decides to pass work to another. Did the triage agent route correctly? Did the specialist agent handle the, the handoff gracefully?
- 15:27
Each layer adds new ways that things go wrong, uh, and non-determinism means that you need to have more evals because errors can cascade.
- 15:38
Uh, here's what an example of an, uh, cascading failure looks like. If you're-- you ask a query about Tesla, meaning the car company, your agent searches the web, finds a whole bunch of information about an eighteenth-century inventor, passes that back to your agent, which writes a beautiful, well-formatted report about an eighteenth-century inventor and passes that to your boss. Nothing that it did there was wrong. It-- you told it to search the web for Tesla. It searched the web for Tesla. It got information about Tesla, and it wrote a report about Tesla. Where, where did it go wrong? Uh,
- 16:09
this is worse than an obvious failure because you have to be reading your inputs and outputs for that to work. You can't-- There was no unit test that would have caught that.
- 16:20
Uh, but agents can also do the opposite. They can get things right in a way that your tests weren't expecting and get graded wrong as a result. This happened to Anthropic when they ran a benchmark called τ²-bench, uh, which simulates multi-turn customer service tasks. Uh, they asked the agent to book flights, uh, and asked it to do something that they thought was impossible, which was reschedule an economy class flight. Uh, but it turned out that the policies that they gave the agent allowed the agent, uh,
- 16:49
to first upgrade the flight to first class. First class flights can be rescheduled, which, uh, economy flights cannot. So the agent found a way to reschedule an economy flight, uh, and failed the test because it was supposed to be not able to do that, and it was able to do that. Uh, so this is an example of how agents being flexible, agents finding things that you didn't think of, uh, needs to be accounted for in your tests. Uh, there's a difference between creatively correct and wrong, and we're gonna come
- 17:19
back to that when we talk about grading your agents later or rather grading your evals. Um, there's also a second way of categorizing your evals into two categories, uh, which is, uh, capability evals and regression evals. A capability eval asks, uh, "Can my agent do this thing at all?" Um, capability evals are expected to mostly fail. Your agent is expected to get a very low score on them, uh, because they give you a hill to climb. It gives you something for your agent to try and get better
- 17:49
at. Uh, once you've climbed that hill, your capability eval turns into a regression eval. A regression eval is an eval that you expect your agent to pass completely or nearly completely every single time. So as your agent matures, you'll be giving it capability eval after capability eval and slowly turning them into a suite of regression evals that you run all the time to make sure that, uh, previous behavior has not regressed.
- 18:17
This is what an eval result looks like. Uh, every eval produces a score and a label. In this case, it's a numeric score, just often just zero or one. Binary scores are very easy for, uh, LLMs to do. Uh, and it comes with a human-readable label like correct or incorrect, or valid or invalid. Um, LLM judges add a third thing, which is the explanation. Um, code evals don't produce explanations 'cause they are just code. Um, but LLM-- an LLM as a judge will say not just that
- 18:47
something is incorrect, but it will say why it decided that something is correct and this is incredibly valuable. Uh, because if you are living in a world of coding agents, you can take a whole bunch of explanations from a bunch of evals and pass them back to your coding agent and say, "Here is all the reasons that you messed up in the last test. What could you do that would make you better at this?" And your coding agent will just go, "Cool. Thank you for the feedback. Improve your application," uh, and it will get better at it next time. This is something I'm gonna show you, uh, in today's workshop.
- 19:17
It is how you turn a failing eval into a prompt improvement. This is a real judge explanation. Um, in this example, it, it's uh, a travel planning agent, um, and you're running a correctness eval. The judge doesn't just say incorrect, it tells you exactly what's missing. You asked for a budget flight, it didn't give a budget breakdown. It says, the agent gave destination, destination info and recommendations, uh, but the user asked about budget travel, and there were no cost estimates. That explanation makes the eval actionable,
- 19:48
uh, because you now have a concrete failure, you know what to fix in the prompt, uh, and the explanation is what makes evals into a useful debugging tool and not just a scoreboard. Uh, and if you're seeing the same explanation across fifty different traces, uh, you know you have a systematic problem and not a one-off edge case, which makes it the first thing you should fix. So you have to take your eval explanations, uh, and categorize them into types of failures, uh, and count up the categories. This is called coding. Again, we can make up words from the L-- from the
- 20:18
ML word-- uh, from the ML world. Um, but categorizing things is known as coding. Um, and of course, each explanation is natural language, uh, so it is hard to do this coding, uh, unless you use yet another LLM to do the categorization for you. We'll be seeing how that works, uh, towards the end of today. Um, and like I said, you can hand a whole pile of these explanations to a coding agent and let the app fix it for you. Uh, this is the full loop that we're
- 20:48
going to build today. We're going to, uh, instrument, trace, eval, and iterate. Each step feeds the next. You define what you want, uh, you build it, you measure how well it works, you ship it, you monitor it in production, and you iterate based on what you see. Every traditional software product goes through a loop like this, uh, but for AI, the measurement step is where most teams fall down, and evals are how you measure. Evals are the connective tissue, uh, across this whole life cycle. Uh, so
- 21:18
let's get started building it. Step one is setting up tracing with Arize X-- AX. Before we can run evals, we need something, uh, some data to run our evals on, uh, and you can't evaluate what you can't observe, so observation comes first.
- 21:33
So hopefully you've conquered the Wi-Fi, uh, and you have our notebook at this point.
- 21:41
And Dat is running around to make sure, uh, that you do. Uh,
- 21:48
let's go to our very first code cell, which is where I install our dependencies. Uh, Claude Agent SDK is self-explanatory. That is the Claude Agent SDK. Uh, OpenInference, OpenInference Instrumentation Claude Agent SDK is the auto-instrumentation package for the Claude Agent SDK. This is how it knows to capture traces from Claude agents, uh, without you changing your application code. Uh, as I mentioned earlier, you can use OpenAI instead if you want to, uh, or some framework. Um, we have packages for
- 22:19
all of those as well. You just need to swap them in. Um, Arize and Arize OTEL are the packages that send, uh, your traces to AX and let you read them back. Uh, and there's a Phoenix package in there that we use as a utility package. Don't worry about it. Um,
- 22:36
the Anthropic package lets you use Claude both to power our agent and to judge its output later. Uh, so, uh, go ahead and run this cell if you haven't already and while it's installing, I'm gonna talk about what it is that we're going to actually build. Um, the Claude Agent SDK, if you haven't already used it, is Anthropic's framework for building agents. So Anthropic's answer to LangChain and things like that. Uh, it can use tools, it can search the web, it can maintain conversation context across turns. Uh, OpenAI has their own, uh, agent
- 23:06
SDK, and of course, there are whole agent frameworks like CrewAI, LangChain, and Mastra. And, uh, AX is compatible with all of those. So no matter what framework you've built, uh, your agent in, uh, it is already instrumented and you can just turn on logging and all of your traces will light up. Um,
- 23:26
hopefully, your install is done by now, which is, uh, optimistic timing on my part. Um, next, set your keys. Uh, in this case, I have stolen my keys from, uh, Colab. If you are feeling naughty, you can just paste them in directly to your cell and put them there. Um, three things go in here. Like I said earlier, you need an Anthropic API key, uh, which is going to power our agent's LLM calls, and it's also going to power the judge later. You need an Arize API
- 23:56
key, which you get from your settings, which authenticates your traffic to the AX service, and you need an Arize space ID, which tells AX which, uh, workspace to put all of your data in. Uh, if your keys aren't working, the usual cu-culprit is a copy-paste error. Uh, and of course, Dat is still running around to help you if, uh, your keys aren't working. Dat is-- has been described as a walking security hole. Uh, so he will definitely give you a key if you can't figure out how to get your keys to work. So now we
- 24:26
will scroll to the register section. Uh, I'm gonna bump up my fonts here.
- 24:36
Is there any chance you can put the QR up one more time?
- 24:39
The QR up one more time. Sure. Probably.
- 24:51
So where's the user data stored?
- 24:56
You don't actually put the values in there, does the user data pull it or-
- 25:01
Uh, yeah, I didn't put-- Mine, mine lives in Colab, so Colab's secrets feature pulls it out.
- 25:09
All right. I'm gonna go back to where I was. Hopefully, I can do that.
- 25:24
All right. Uh, so we've installed dependencies, we've talked about the SDK, we've talked about the secrets, uh, and now we're at the register. So this is the magic. Uh, anybody who works for Arize will tell you that you only have to add two lines of code to your application in order to instrument your call, and this is-- these are the two lines of code. Uh, what we're doing here is we're passing in our Arize space ID, our API key, uh, and a name for it to trace
- 25:54
everything to. Um, the register function is setting up OpenTelemetry. OpenTelemetry is the industry standard for application observability, uh, and it has a layer on top called OpenInference, which adds LLM-specific attributes, things like prompt text, completion text, token counts, which model was called, which tools were invoked. Um, we point it at AX, uh, and we give it a project name, and that is how your traces get grouped in the UI. Uh, in this case, we're also passing batch equals false
- 26:24
because, uh, when you run in a notebook, uh, you want to send your telemetry as soon as it happens and not batch stuff up, which it does for efficiency in production. Um, that one line at the bottom, the Claude Agent SDK instrumenter, tells the Claude SDK to send a span to AX every time it makes an LLM call or invokes a tool, uh, and that is now instrumented. The reason this works with so little code is because OpenTelemetry is an open standard that nearly everybody uses. So the Claude,
- 26:54
the Claude SDK au- authors, the OpenAI SDK authors, the LangChain authors, all of them have already written the code inside of their SDKs that calls OpenIn- that calls OpenInference and sends data back. So all you have to do is say, "Hey, I'm an OpenInference collector. I live here. Send it to me."
- 27:13
One more bit of setup is that the register call sends data to AX, uh, and the Arize client reads data back out of AX, uh, so we'll use it to pull our spans back into the notebook. Uh, that is what this line is about, so you need to make sure that you've put in the same keys there or, uh, copied them across.
- 27:33
I have to use this whole Colab. How are you pasting in the keys? Like, that's manual. Sorry.
- 27:41
Oh, you need to make a-
- 27:42
Pasting, but-
- 27:42
You need to make a copy of the notebook from the File menu to be able to edit it.
- 27:45
I did that.
- 27:51
Oh,
- 27:55
thank you.
- 27:56
Cool. Um, all right. So if you've run that, run those cells, then we are ready to build. Uh, before we build, one thing worth knowing about, uh, everything we're going to do by hand today, you can also get your coding agent to do for you. The reason we are doing it by hand is so that you have a fundamental understanding of what it is that you're doing, and you're not just vibing your way to success. Um, but we publish a set of skills called, um, the Arize Skills, uh, which know how to do all of this for you at the skill level. So you can install your Arize Skills,
- 28:26
uh, and, um, run evaluations, create datasets, uh, instrument your app in the first place. Um, they all use, uh, AX's recommended patterns, and you install them once with a single NPX command. They work in Claude Code, they work in Cursor, they work in Codex, and dozens of other coding agents. Um, so we're gonna do stuff manually today, but in production, in a real-life en- environment, we are using our own skills every day to do this stuff.
- 28:54
Uh, so now let's build our agent. Uh, when I asked if everybody had written an agent before, everybody put their hand up, so I'm not gonna spend a lot of time explaining how this agent works. I'm going to assume that you already have an agent somewhere, and you're just trying to instrument it and make it work for you. Um, the fake agent that I'm using today is a financial analysis chatbot, uh, using the Claude Agent SDK. You give it a stock ticker and a focus area, and it searches the web for the latest information about that stock ticker, uh, researches real current financial data, and
- 29:24
writes a report for you. Uh, this is a u- a real use case that whole startups are built around, although, uh, my agent version is extremely simple. Um, I'm assuming, uh, that you built an agent already. Our agent works in two turns. Um, there's first a research turn that uses tools to search the web and gather data, then there is a writing turn that compiles the research into a readable report. Uh, so let me walk you through that. Uh, here is my research prompt and my write prompt. These are
- 29:54
deliberately extremely simple. They are going to cause errors later, and those are the things that we are going to use our evals to debug. Uh, and then I'm setting up my Claude agent opt- options. I'm using, uh, Claude Haiku, uh, as the agent underneath because Haiku is, you know, capable, but it will make mistakes, and we want some mistakes so that we can fix them. Uh, and it's also fast and cheap, so I don't burn a lot of money, uh, of your money when-- as we do this. Um,
- 30:22
the permission mode controls, uh, what the agent is allowed to do autonomously. There is a catch hiding in the, in the specific choices that I have made and the specific allowed tools, uh, that I am giving this agent right now, as we're going to find out later. Uh, and now let's look at our two turns. Um-
- 30:43
Turn number one is research. There's a bunch of boilerplate in here that I'm not going to explain, which is just about outputting the output so that we can see what it's doing as it's going. Uh, turn two is writing the report with, again, a pile of boilerplate. Um,
- 30:59
the prompts are critical. So this is where I want you to start doing your own changes in your own notebook. I have given it extremely basic prompts at the top. If you can-- see if you can think of a better prompt just off the top of your head, uh, that is going to do a better job than the prompts that I put in, uh, at solving this problem. What would be better than what I put in to, uh, do the research in the first place? What would be better than what I put in, uh, to do the report writing? Um, and if you one-shot it,
- 31:29
then you're gonna have, uh, less stuff to do later on. Um,
- 31:35
one of the things we do here is we wrap the whole thing in an OpenTelemetry span. Uh, this is because if I don't do that, then Claude Inter-- Cla-Claude senses as two-- the two turns as two separate spans, and we wanted them to be grouped together in the UI. Uh, so I've wrapped them in a span together so that they come out as a single agent turn. Um, this is just saving us some time later. I wanted to be clear about why I'm doing this, uh, bit of code. Um, and now we can
- 32:05
run our agents. Uh, I have run my agent in advance because I wanted to know exactly what it was going to do. Um, but now is the time to kick off your own agent. Um, I've asked it to analyze Tesla, specifically its financial performance and growth outlook. Um, you should go ahead and run this cell. It'll take a minute or two because the agent will actually search the web, uh, and the LLM is doing multiple rounds of reasoning. Um, the agent receives our research prompt. It decides it needs to search the web, which is a
- 32:35
tool call. It reads the search results. It decides if it has enough information and maybe searches again. That is the agentic part. You didn't say, "Do a web search, take the web-- information from the web search, and write a report about it." The agent is deciding, is this web search enough information? And if not, I will write-- I will run more web searches. It will run an in-indeterminate number of web searches to get enough information until it has decided what enough means, uh, and then it will write the report. This execution path is non-deterministic by
- 33:05
design, and that is the key here to understand. Uh, if you run this again with the same input, your agent might search for different things. It might find different results. It might write a different report. Uh, and both reports might be good. Um, or one might be good and one might be bad, and that is exactly why we need evals. We can't predict the output from the input alone, uh, and every single one of those decisions is being captured as a trace. Uh, so I'm just gonna pop open my last twelve
- 33:35
hours or my last twenty-four hours, I guess. Cool. This is a pile of traces. You're going to understand what all of this is later, but here you can see the Tesla one that I ran earlier. This is what it looks like when I go through. You can see every single turn, uh, it's calling a skill, it's doing a tool search, it's doing a web search. Uh, you can see every single thing that it does. You can see that it ran four web searches before finally deciding that was enough information. Then it did the second, uh, call in the step, and it wrote the
- 34:04
report. Uh,
- 34:09
one of the things that it does, uh, that we're gonna see later is, uh, I told you there was a catch in the tools that I gave it. Uh, here's this step that was unexpected when we were putting this demo together. It tried to write the report. I told it to write a report. It tried to write the report as a markdown file to disk. It is living in a Colab notebook, so there is no, uh, there's no file system for it to write to. Uh, so the write step would fail. Uh, and this is something that you need evals for. You need evals to detect when your agent is doing something
- 34:39
helpful but incorrect. Uh, because this fails silently. It gives you output anyway, but it's doing this unnecessary extra step of trying to write to disk, uh, without you knowing that it was there. Um, so back to the Colab. This is a pretty good report. Um, it's a financial performance table with real numbers, growth drivers, risk factors, analyst price targets, and a conclusion. Um, it's not bad for a first
- 35:09
pass, but it-- not bad as a vibe, and we are here to replace vibes with actual numbers. Um,
- 35:17
so like I said, uh, you can see all of that stuff in the, uh, traces. We didn't have to do anything to get all of that information. We didn't have to, like, instrument every single line to get the LLM call and get the tool call and all of that stuff. All of that is built into Claude Agent SDK, just as it is-- like it's built into OpenAI SDK. Uh, and each one of these rows in this table is a span. Um, traces reveal every decision the agent made. Without them, all you see
- 35:47
is I gave it a prompt, and I got a report. With traces, you see every step. Uh, if you click into any span, like I said, you'll see exactly what the model received as input and exactly what it returned. Um,
- 36:01
and that is what observability means in practice. Not just did it work, uh, but how did it work, and where exactly did it go wrong? Uh, when we run our evals in a few minutes, uh, we're gonna pull the input and attribute-- uh, output attributes out of these spans and feed them to our evaluators. The input becomes the evaluator's input. The output, uh, becomes what gets graded.
- 36:26
Uh, now we need to generate some test data. To run meaningful evals, uh, we need more than the one trace that you've generated so far. We need a body of data. Uh, in the notebook, there are twelve test queries covering different tickers and analysis types. Uh, you should kick those off now. They take about five to ten minutes to run, assuming that you've got Wi-Fi and everything's working for you. Um-
- 36:49
I'm gonna show you what they look like, uh, and
- 36:55
oops, this is them here.
- 37:01
So the key thing here is, uh, diversity. I've given it a bunch of obvious queries that should work. I've also given it some edge cases. For instance, I've given it one where I gave it two tickers, Apple and Microsoft, and told it to run a comparison, uh, which is something it could theoretically do. Uh, I've also given it the same, uh, tickers multiple times, uh, in places so that we can see, uh, the non-deterministic output, uh, when we say, uh, Amazon profitability trends and
- 37:31
outlook versus Amazon AWS performance and profitability. Um, different tickers and different questions, different levels of complexity are important, uh, to generate test data that's going to cover the range of things that users will actually ask. Although, spoiler alert, in reality, your users are going to ask really weird things that you're not going to be able to predict in advance, and that is another one of the reasons that evals are important. Um,
- 37:58
the, uh, Rivian query, R-I-V-N, uh, asks about a company with much less public data than Apple or Microsoft, and the Coca-Cola dividend yield query is a very different kind of analysis from a goth- growth stock query. Um, again, this is diversity that is important. Um, diversity in your test set is how you catch, uh, things that your agent is going to be not good at.
- 38:22
Uh, so while yours are running, mine have already loaded, so I can show you all of my data. Uh, these are all of my spans from all of those runs. Um, you can see, uh, at the bottom, there are thirteen, uh, the-- which is the initial Tesla run and twelve test queries. Each one is a complete execution of the Claude, uh, financial analyst. So now we have the data, but before we write any evals, uh, we need to actually, uh, look at the data.
- 38:53
Um, this is error analysis, uh, and the whole step is one instruction. Read your traces before you write evals. Before you write a single evaluator, start with your data. Uh, read or... a dozen or more of your traces end to end. This is the most important practice in this entire workshop. Uh, focus on the ones where something went wrong. What was the input? What was the output? What specifically is broken? This sounds extremely old-fashioned, just read the input, uh, but it is one of the
- 39:23
highest value activities in agent development. Uh, Anthropic, for instance, invested in tooling specifically for viewing eval transcripts, and their team spends a whole lot of time just looking at eval transcripts every day. Um, a trace tells you whether the agent made a genuine mistake or whether your graders rejected a valid solution. Uh, if you automate before you understand your failures, then you're going to create an eval that measures what's easy to measure instead of what actually matters.
- 39:54
You need requirements first. Before you can categorize failures, you need to know what success looks like. You can't say it doesn't work if you haven't defined what it works means. Uh, for our financial analyst, what does a good report look like? Uh, it should reference the correct ticker. It should include some recen-- real recent actionable financial data, uh, and include actionable recommendations. It should distinguish between forward-looking analysis and historical summary. Those are our success criteria, and they seem obvious, uh, when you write them down,
- 40:24
but lots of teams never do. Um, they ship an agent and then react to complaints instead of defining the bar up front. Uh, writing requirements down turns vague disappointment into specific testable criteria, and those criteria are what turn into your evals. Here's the thing about defining success that it's important to note, though, which is that it is not something that your engineers can do or not something your engineers can do alone. The definition of good usually lives in the domain knowledge, and the domain knowledge lives in the people
- 40:54
that your company will often laughably refer to as non-technical. So your product managers, your QA, your support team, those are the people who know what the definition of good really is. They know what the failures are going to look like. Those are the people who should be helping you write your evals. Uh, OpenAI put it really nicely. They said that people management skills are AI skills. So clear goals, uh, direct feedback, knowing what your value proposition is, uh, those skills matter more than ever when the system is probabilistic.
- 41:23
Uh, and this is one of the things that AX is built around. Anyone on your team can read traces or add annotations and contribute to your criteria. This is not just an engineering surface. You're expected to have other members of your team in here.
- 41:38
A quick note on where to get this test data. Uh, we have these traces because we already ran the agent ourselves, but what if you're building something new and you don't have real traffic yet? Uh, you can do what we just did, which is you can use synthetic data. You can have an LLM generate diverse queries across your expected categories. So, uh, research Tesla financial report performance is one phrasing. What's going on with Tesla stock is another query that's asking the same thing. Yo, is Tesla a buy right now is a third way of expressing the same query.
- 42:08
Uh, they're all the same intent, but they look very different. So vary the phrasing, the complexity, the level of specificity. Uh, a domain expert should review your set of queries, uh, because LLM-generated queries tend to cluster around obvious phrasings and miss the weird stuff that real users write. You also need to include edge cases in your test data, so things like non-existent tickers, multi-part questions, jailbreak attempts. Uh, those might be one percent of your traffic, but they are the one percent of your traffic that ends up, you know, being a PR
- 42:38
disaster in the press. Your test data should look like production data, not what you wish production looked like. Uh, so synthetic data can get you started, but as soon as you have production data, production data should be what you're running on.
- 42:52
So let's actually examine our traces. Uh, like I said, the first thirteen traces are the first thirteen traces that I wrote. Um, most of them produced long structured reports in line. So Tesla, Apple, Nvidia, stuff like that. Uh, three to seven thousand characters, uh, of reports right there in the trace output with executive summaries, valuation tables, explicit buy or hold recommendations. On the surface, most of these look fine, but like I said, three of the
- 43:22
thirteen were doing something weird, which was, uh, they were, uh, the repor-- output in them is much shorter because it just says, "Oh, I wrote this to disk for you," uh, and the output went to disk. Uh, so instead of having a full report, uh, it has a summary of the report and the, the actual body of the report went to the non-existent disk. This is what the eval is designed to find. Um, the agent decided to write a file with no
- 43:52
permission. The write silently failed, and it told us the report was saved. Uh, and that is the kind of pattern that you only a-see if you actually read the trace. Um, so this is, I believe, one of the three sources. Nope, that one had worked.
- 44:09
Oh, no, the Wi-Fi. Here we go.
- 44:25
So this one's showing another kind of failure where it just got stuck in a loop forever. It was trying to find information about... Who was it trying to find information about? Microsoft. Uh, and it just did an endless series of web searches and took forever. Um,
- 44:39
this is what it tried to write to disk, uh, and this is what it actually wrote to output. Uh, right, here we go. This is, "The report is saved as Microsoft Cloud Segment Financial Report dot MD." This report is useless 'cause it's referring to a report that doesn't exist. Um,
- 44:58
there's also, uh, an example of a confidently wrong. If we go to the Rivian report, uh, which is R-I-V-A-N. Where did it go? There we go. Uh, Rivian has the same problem that the Microsoft one does, but even inside of the full report that it tried to write to disk, there's a whole bunch of really confident information about a private company that you can't possibly verify. Uh, these might be hallucinations. They might not be hallucinations.
- 45:27
But we haven't done anything to check whether or not, uh, this report is really grounded in reality. Uh, so we're gonna do an eval about that later. There is a structured way to do this reading, uh, which, as I mentioned earlier, is called coding. There's open coding and axial coding, um, and it comes from qualitative research. Open coding means you read each trace, and you write down what you see. So stuff like vague recommendation, made up a number, talked about the wrong ticker. Uh, you're
- 45:57
not trying to be neat. As a programmer, it is very, uh, tempting to try and come up with categories in advance and say, "Oh, this is a, you know, retrieval failure," or whatever. Uh, don't do that when you're doing open coding. Just write down what the problem was in as natural language as possible. Uh, and, uh, then when you find the failure-- Sorry. And then, uh, you wanna do axial coding afterwards. Axial coding is when you look at, uh, all of the open codes that you've put together, and
- 46:27
you-- then you do the categorization. You say, "Okay, now there are five things that look similar. I'm going to give them all a single category and, uh, turn them into categories." Um,
- 46:39
the second pass, the, the, uh, axial coding pass is structural. Uh, and like I said, the tend-- the tendency, the temptation is to try and go straight to axial coding, and you should resist that. When you find a failure, you should ask why it failed. So the response was wrong is a symptom. Did it get bad search results? That is a retrieval failure. Did it get the right data, uh, but the wrong conclusion? Then that is a reasoning error. Did it make up a stock price? That's obviously hallucination. Did it answer outside of its domain? That is a
- 47:09
scope violation. Did it do something that you didn't tell it that it should be able to do? Each root cause, each axial code, uh, points to a different fix. If you don't know the right cau-- If you don't know the cause, you can't pick the right remedy. Once you've categorized, uh, which I did here,
- 47:30
uh, you can sum them up into, uh, a table. So,
- 47:40
uh, confession time, I didn't actually read all of those traces and code them. I just gave them random axial codes. Uh, but what you end up with is this table of root cause frequency. Things like looks good, possible hallucination, reasoning gap, unverifiable data, missing recommendation. It's tempting to look at the top one, the possible hallucination, and decide that that is the most important one to fix. But in reality, uh, you're going to want to, uh, balance between, uh,
- 48:10
frequency and severity. So if one time in a hundred, instead of doing a financial report, it gives the user instructions on how to make a bomb, you fix that one first, because that is the most severe possible failure. Uh, so you have to, uh, multiply the severity by the frequency to get to which ones you decide to fix first. Uh, and that is why you look at the data before you write your evals. Uh, one more concept before we get started on that is the Swiss cheese model,
- 48:41
um, which is a concept from safety engineering. E-- Imagine each layer of defense as a slice of Swiss cheese. Each slice has holes, uh, gaps where problems can slip through. Uh, but if you stack enough slices, the holes don't line up. So what gets through one layer gets caught by the next. Uh, your code eval catches format issues but misses semantic problems. Your LLM judge catches reasoning gaps but misses subtle hallucinations. And your human review catches the subtle stuff, but it can't scale to every trace. No single eval in that set, uh,
- 49:11
catches everything, and the combination is what gives you coverage. Uh- So stack the cheese. Um, now that we know what's wrong, let's automate the checking. Um, the next step is code evals. This is no model configuration, no API calls, uh, just Python. We're going to write the simplest useful eval that I can think of. Uh, our agent is supposed to analyze a specific stock ticker. What I'm going to do is write a code eval that checks whether that
- 49:41
stock ticker actually showed up in the report, uh, which is a test that my initial demos absolutely failed a number of times. Uh, an LLM judge would be overkill for this. You're looking for a specific string. It's somewhere in the report. You don't need an LLM to judge whether or not you mentioned the ticker.
- 49:58
So there's two ways to do this. There's two ways to do everything in AX. One is programmatically, uh, and one is via the UI. Uh, in the, uh, notebook, you can see the programmatic way. What I'm going to show you on screen is how you do it in the UI.
- 50:19
So this big button in the top right is what you want. You want to add an evaluator. In this case, I've already added my mentions evaluator.
- 50:29
But you will want to add an evaluator which will give you this screen. Uh, here you're going to see basically some Python. Um, in the notebook, I've given you the Python that you can use, uh, to either do your, uh, programmatic version or your version in the UI. Uh, this is a very simple Python function which returns, uh,
- 50:54
an evaluation result, which is... looks exactly like those evaluation results that I showed you earlier. Uh, it has a label. It has a score. Um,
- 51:06
so, uh, once you've written what your code evaluator does, you're going to have, uh, parameters to your code evaluator. In this case, I've called them query and report, and you're going to have to tell AX, uh, which variables in your trace are those variables. So there's this extremely fun UI, the single trace UI, uh, where you look at a single trace and you can literally just click, uh, the specific thing in your
- 51:36
output that you want to be that variable and you tell it, "Okay, that one should be the query and that one should be the report," and then it's mapped.
- 51:47
Um,
- 51:49
you should also, uh, be changing your evaluator to, uh, work on, uh, a trace as opposed to a span. That is what gives you that UI. Um, once you've done that, you need to save your code evaluator to our evaluator hub, uh, and then you hit save to save the evaluation.
- 52:17
So assuming you've done that... Yeah.
- 52:20
Can we talk a little bit on the evaluators, on the concept and what kind
- 52:28
of
- 52:28
Sure. Uh, so he asked for more detail on how code evaluators work. Uh, so code evaluators are-- they're completely deterministic functions. You're just taking input and output and saying, "Was this output good according to some definition of good?" So in this code, uh, I'm doing very, very simple Python. I'm just saying, uh, look for all of the words that are in capital letters, uh, strip out the ones that are obvious acronyms that I already know, and of the remaining words, are any of
- 52:58
them the, uh, the acronym which is the stock ticker that I'm looking for? That's all it's doing. Uh, lots of code evaluators are even shorter than that. They're something like, "Is this output five hundred characters long or less? Uh, is this output parsable JSON?" You're doing something very simple, very deterministic, uh, for the purposes of having that first line of defense in your Swiss cheese.
- 53:23
Um,
- 53:29
so I talked about how to add stuff already. Uh,
- 53:33
in the notebook, you can run the eval programmatically. Uh, this is how to do-- this is-- these are the instructions on how to do it via the UI if you followed that. The alternative t- is to do it in the code where you can just run it directly, which is, uh, easier to do if you're following along in a notebook. Um, a code evaluator returns a label and a score. Like I said, there's no LLM needed. There's no API call. It's instant. Uh, and, uh, if you're wondering why we're wrapping it in suppress tracing,
- 54:03
it's because, uh, the evaluation is itself, uh, a call to the, uh, agent SDK which would then get stored as traces and I didn't want traces of my tracing happening because that gets too meta. Uh, so I told it to suppress tracing when I'm running an evaluation. Uh, if you look at this programmatic run, uh, you will see that twelve passed and one failed. Uh, there was a missing ticker for Amazon. Uh, the specific trace that went wrong,
- 54:33
uh, is when I asked-- Uh, you'll no- you'll notice I told, told you earlier that I made two queries to Amazon, one where I asked about its profitability and one I asked where about, uh, AWS's profitability. Uh, if you click into the traces, uh, which I could attempt now, um, you'll see that the report that asked about AWS, uh, wrote a report that was entirely about AWS, only about the AWS branch of Amazon rather than,
- 55:04
uh, all of Amazon. Uh, and so it failed to mention the Amazon stock ticker entirely. It talked only about AWS. So it's a really subtle failure 'cause it did a bunch of research. It found out a whole bunch of stuff about what was happening at Amazon. It just didn't think about what was happening about Amazon as a company, uh, only about AWS because I gave it that extra prompt like, "I also want to know about AWS profitability." Um,
- 55:29
so, uh, looking at those results back in AX, um, they show up as annotations next to our traces. So you can click into any one of these traces, uh, and you'll see evaluations. I've run a lot more evaluations, uh, than you have so far, but you can see my mentions ticker evaluation has run with the span. It's run at the span level, it has hit a pass, uh, and it's given a score of one. Um, the notebook has a little
- 55:59
helper called Log Eval to AX, which is doing that work for you. It takes, uh, the spans that it ran offline and the tests that it ran offline and pushes them back to AX for you. Um, so now every score-- now every eval has, uh, a mentions ticker evaluation, uh, that you can look at and filter by, uh, and look for the failure-- failing ones. So why does this matter? Because it is your first line of defense. Uh, if your financial analyst agent writes a beautiful report about Mi-
- 56:29
Microsoft when the user asked about Tesla, uh, which is a thing that actually happened to me when I was building this, uh, then no amount of answer-- of eloquence from your agent, uh, matters because the answer is wrong. Uh, so you can catch this with five lines of Python. Um,
- 56:46
the ticker check catches, uh, really basic failures, but it catches them really cheaply. That is why it is important. Um, code evals aren't just toy examples. Often, you're going to want to know that your output is JSON, uh, or that it has a certain length, or you're going to want to avoid forbidden phrases like, "As an AI language model." Uh, and those are production-critical checks. A code eval doesn't have to be a simple string operation. Code eval means the grading logic is deterministic, uh, but it can do complicated things. It can query a
- 57:16
database for you, it could've hit an API and get the actual stock price to make sure that it wasn't hallucinating a stock price. Uh, anything where the grading always gives the same answer for the same input can be a code evaluator. One important principle is to grade what the agent produced and not the path that it took. Uh, I referred to this earlier. There's a common instinct to check that the agent followed a specific sequence of steps. Did it call this tool first and then that tool in that order? Uh, and, uh, practitioners have found
- 57:46
this to be, in practice, too rigid. Uh, agents regularly find valid approaches, uh, that the eval designer didn't anticipate. Uh, so, uh, if you grade the path, um, you punish creativity. Instead, grade what the agent produced. Did the output have the right information? Did the final state match what you wanted? Uh, our ticker check doesn't care how the agent found the data, it just checks, is the ticker in the output? Uh, that is the right level of abstraction for most code evals.
- 58:17
So now let's get to built-in evals, LLM evals. This is where it begins to get more interesting. Uh, code can check whether the ticker appears, but it can't check whether the analysis is good. Is the financial data accurate? Is the report complete? Are the recommendations well-reasoned? Uh, these are semantic questions, and for semantic questions, you need a judge. Every LLM as a judge has three parts. Uh, it has a judge model, which is the LLM that is doing the grading for you. It has a prompt template, which, as I mentioned, is
- 58:47
also called a rubric, uh, which is the criteria that the judge applies, the rules of what defines good. Uh, and it has the data, which is the examples being evaluated. AX keeps these three things separate, which means that you can mix and match. You can try the same criteria with different judge models, the same judge model with different criteria. Uh, it is modular by design. The good news is that AX ships with a built-- with a bunch of built-in evals, so you don't have to write your own. There are a bunch of types of evals that are always, uh,
- 59:17
that are often going to apply to, uh, most AI applications. So we've already written the prompts for those, and you don't have to write them yourself. Uh, correctness, for instance, checks whether a response is factually accurate. Faithfulness checks whether the response stays grounded in the source material. Uh, and there are evals for tool selection. Did the agent pick the right tool, uh, and tool invocation? Did it pass the right arguments to that tool? Stuff like document relevance, refusal detection, many of the things that you'd want to check in an agentic application,
- 59:47
uh, come out of the box. Uh, so the first eval that we're going to run, uh, is a correctness eval. Um,
- 1:00:00
and spoiler alert, it's not going to work. Um, setting it up is very easy. We give it an LLM. In this case, uh, I've chosen, uh, the Anthropic LLM, uh, Sonnet 4.6, because it is better than, uh, Haiku, which is the thing that is doing the work. You often want to pick a bigger, slower model to be your judge. Uh, and then I've just, uh, given correctness, ev-evaluated the LLM, uh, and, uh, run it right here. Um,
- 1:00:35
like I said, uh, once we've run the evaluation with evaluate data frame, uh, we've also, uh, turned on suppress tracing here so that our evaluation doesn't send a bunch of extra traces to, uh, AX. Um, it gives us this output, which is a bunch of evaluation rows. Um, these are not very readable. This is why you need AX to look at them. Um, the judge reads the input and output from each span and grades it against the rubric. Um,
- 1:01:07
and we can see that every score is zero. If I scroll across...
- 1:01:15
Status completed, name correctness, score zero. The reason it's doing this is because correctness is judging against the, uh, LLM's definition of, of what is correct and what is true, uh, which means that it is relying on the LLM's training data. What we're doing in this agent is a whole bunch of live web search to get current stock prices and up-to-date financial information, which means the agent is talking about things that happened in 2026-
- 1:01:46
But the LLM that's doing the judging only has information that stops in January. So if you read the explanations of our correctness eval inside, uh, inside of AX, you can scroll to evaluations, you can look at correctness, you can say that it is incorrect. Uh, I will save you a bunch of reading of a very densely packed column. What it's doing is complaining that there's a bunch of stuff that comes from the future that it couldn't possibly know. Uh, so our correctness eval here has completely failed. This is not what you want it to do. So instead,
- 1:02:16
we are going to use a faithfulness eval. Uh, faithfulness is a better eval for our use case because, uh, the judge actually has context. In faithfulness, you give the, uh, LLM as a judge the information that it is being based on, uh, and then say, based on the same information that the jud- that my agent collected, did it do a good job of writing the report? This is-- This way, they are on a level playing field. They both have the same amount of information, and your slow, complicated judge agent, uh, can
- 1:02:46
judge whether your, uh, quick, fast, uh, worker agent, uh, was, uh, doing a good job. Um, so let's look at the faithfulness cells. Uh, just like everything else, these can be done online as well as offline. So, uh, there is a faithfulness built-in evaluator, uh, which isn't in this view. There is a faithfulness built-in evaluator which you can pull up,
- 1:03:17
uh, and you can give your input, output, and your context to it. Uh, again, you can do this in code as well, and if you're using the skills, this is probably what's going to happen. Your skills will say, "Oh, I'm not going to use the UI. I'm going to use the code," uh, which is why I show it here. Um, so the first thing we do is we extract the research context from the traces. As you'll remember, we do, uh, a two-step process. So I take the output from the first step where it does the research, and I attach
- 1:03:47
that to all of my spans, uh, as, uh, the context for it- the judge to work on. Uh, and then I run the faithfulness evaluator. So, uh, here you can see spans with context. These are the same spans that I just pulled out of AX. I attached at the column level all of my context, uh, and then, uh, I called evaluate dataframe with those spans, and I t- gave it my faithfulness evaluator,
- 1:04:17
which I instantiated here just like I instantiated the correctness evaluator. This gives us, uh, much more interesting data. This gives us unfaithful seven and faithful six. Uh, so, uh, roughly half the time, my very simple report is hallucinating. It is coming up with things that were not in the source data according to Sonnet. Um, and that is--
- 1:04:46
this is a useful eval. This is an eval where we're failing half the time. This is a capability eval where we can climb the hill, modify our prompt, and get it better at writing a grounded report. This is exactly the kind of eval that you want. Um,
- 1:05:04
so that's two built-in evals and two very different signals, and there's a really important lesson here, which is correctness gave us zero out of thirteen, and faithfulness, uh, told us that reports aren't grounded in their sources roughly half the time. The difference isn't that one eval is better than the other. The difference is, uh, the use case. There are lots of use cases where that correctness eval that I showed you earlier works just fine for some use cases. Uh, it's just not in this specific use case where we are doing live, real-time research. Um, so it's important to know the
- 1:05:34
question that you're actually asking.
- 1:05:38
Built-in evals are your starting point. They give you an immediate signal without any prompt engineering, uh, but even with a good built-in like faithfulness, uh, there are a bunch of things that it can't check. Uh, our financial analyst, uh, should produce actionable recommendations. It shouldn't just summarize data. Um, so not just, you know, give you a financial report, but tell you what to do with the financial report. Should I buy, sell, or hold? Uh, no built-in template checks for that. This AX is-- does not cover
- 1:06:09
every single use case, so you have to build a specific eval that checks for actionability. You have to build an, uh, custom eval that is about your domain. If you're writing real evals, you're absolutely going to have to write a custom eval with a definition of good that we didn't think of. So that brings us to the next step, which is writing a custom eval rubric. This is where you make evals truly specific to your application. And writing a good eval is complicated, so I'm going to spend a little time talking about what makes a good eval rubric.
- 1:06:39
Uh, we recommend four parts of a custom rubric. The first is to define the judge's role. Um, you should spell out ex-explicit pass and fail criteria, you should label the data with XML tags, uh, and you should define the output choices outside the prompt itself. Uh, I'm going to add one more thing on top of those, which is actually kind of controversial amongst pro- practitioners, which is I tend to include examples of good and bad output. Uh, some practitioners believe, uh, that
- 1:07:09
including examples will make your, uh, eval overfit. Uh, but let's walk through all of those. So the first part is defining the role. Um, you'll notice this, uh, what is by now a tired trope, which is you are an expert financial analyst evaluator. That is not the important part. Tests show it makes some very small difference, uh, to whether or not, uh, your eval works. The important thing is to give it the context of what it is looking at. You are looking at a financial report, and
- 1:07:39
what I am expecting is financial analysis. Those are the important pieces of context that you are giving the LLM there. Um, you're also telling it that it should be-- that what you want is something actionable. So that brings us to part two. So, uh- This is where most people underinvest. Don't say a good response because that is, uh, an aspiration. That is not a criterion. A good response is helpful and accurate. Sounds like you're adding more detail, but you're still being just as vague because helpful and accurate don't have any
- 1:08:09
meaning as far as the LLM is concerned. Uh, instead, you should be listing exactly what makes a report actionable. So list exactly what makes it actionable and not actionable. Every criterion should be specific and observable, something that if you were a human looking at this report, you'd be able to know whether or not it was there, and you'd be able to say yes or no. Uh, so contain specific recommendations. That's something you can definitely say yes or no. Includes forward-looking analysis, not just historical data. That's a clear distinction that you can say yes
- 1:08:39
or no to. Uh, and the important thing here is that each of these criteria maps to something we saw in our analysis. If we... When we were reading the reports, uh, we saw-- we actually observed when we read our traces, or at least when I read our traces, uh, that, uh, these are ways that these reports are going wrong specifically. That is not a coincidence. Your error analysis is what tells you what criteria to write. You don't come up with them from whole cloth because it's not going to occur to you. Where you get that data from is from reading your
- 1:09:09
traces and going, "Oh, it didn't do that. My axial coding said it does this specific wrong thing." Uh, so specific criteria produce consistent judgments, uh, while vague criteria produce coin flips. The third thing you should do in your rubric is you should label the data with XML tags, especially if you are using Anthropic. Anthropic loves XML and finds it very easy to understand. Um, your rubric has variables, the input and the output, and you want to make sure that it can tell the difference between your instructions
- 1:09:39
and the input and the output, and XML tags are an excellent way to do that. Um, so put a user query tag around the input, a financial report tag around the output. Uh, clear boundaries reduce the chance that the judge confuses the question with the answer. It is especially important if you're including examples, uh, to label the examples with XML tags because otherwise it will go, "Oh, this is exactly what I'm supposed to write," and just spit out the example back to you. Um, which brings us to the examples. This is a part a lot of people skip.
- 1:10:09
Uh, in my opinion, it is the part that most improves quality. LLMs are okay at following instructions. They are increasingly good at following instructions, but what they are really good at is copying an example that you've already given them and replacing the things that are relevant to the current situation. So if you give it an example of exactly the kind of report that you're looking for, it's gonna follow that format exactly. That might or might not, depending on your domain, be a good idea. If your, if your example is too rigid, then the LLM will
- 1:10:39
always give the same report, and you have reduced its scope for creativity. Uh, but on the other hand, if it is being too non-deterministic and your report is being crazy, uh, then giving it an example will ground it in what exactly you are expecting.
- 1:10:55
Uh, here's what an actual example looks like in the template. Uh, it has everything. It has, uh, percentage revenue growth, a PE ratio. It identifies a concrete risk with a number to back it up, and it gives a specific recommendation. Uh, that is what actionable looks like. Here is a non-actionable result. Uh, it's not wrong. Uh, everything it says is correct, uh, but it doesn't say anything useful. It doesn't have any specific data. It doesn't have any specific risks. It just says investors should consider various factors. Uh, that's not a recommendation. Uh, so the
- 1:11:25
judge now has a concrete reference for what you actually mean by the labels that you are giving it. Uh, and this can dramatically improve consistency, especially on the edge cases where the report, uh, would have otherwise been somewhere in between.
- 1:11:40
Uh, and this is, uh, an important part. It is very tempting if you're writing this rubric to say, "And now spit out the word correct or incorrect," because that is what you want it to do. You do not need to do that anymore. Uh, the way that it actually works under the hood is we have given the LLM a tool that says, "You call this tool when you are done," and it will call that tool with correct or incorrect when you are done. So you don't need to tell it what output, and in fact, telling it to spit out a s-specific string as output is more likely to confuse it and
- 1:12:10
make your eval fail. So just in the UI, uh, set your, uh, choices, uh, or when you are setting the-- when you're creating it programmatically, set the choices, uh, and let the LLM do its thing. Um, when you are coming up with these choices, it is best to keep those choices simple. Uh, binary is best, two clear labels. Uh, numeric scales, like a scale of one to five, uh, they are tempting, but the LLMs are bad at them. What is the
- 1:12:40
difference between a three and a four? You would have to write a whole bunch of rules for the LLM to know that. Uh, it doesn't know that automatically. Uh, whereas you can give it a pass/fail, and it will be pretty good at deciding what is a pass and what is a fail. Uh, so if you really need more categories, you can do like a partial result. You can say a zero point five score is where it did a, a half-assed job. Uh, but by default, go for a zero or one. Um, one more practical tip is chain of
- 1:13:10
thought for judges. Ask the judge to reason before it scores and have it explain its thinking first. Uh, this measurably improves, uh, accuracy. Uh, when the judge has to articulate why something is actionable or not, or why it is correct or not, uh, it makes fewer mistakes than when it just picks a label, uh, without thinking first. Uh, this explanation is also very helpful for you, uh, but it also helps the judge think better. Just one of those weird things about LLMs. Uh,
- 1:13:40
so if you are coding along, now is the time to find your actionability template. Uh, this contains all of the example-- all of the instructions that I, uh, was just talking about. Uh, can you do better than me? Can you one-shot your way to a perfect eval, uh, by modifying this eval template? Um- As with, uh, all things in AX, you can do this via the UI as well. So you can create an online evaluator, uh, and you can create a no-new LLM as a
- 1:14:10
judge. In this case, I've already got it, my actionability evaluator, uh, which has literally exactly the same text, uh, that I just showed you in the notebook. Um, and as before, I had to pick a trace, uh, scope. Uh, and because it's an LLM as a judge, I had to set up a template, uh, sorry, I had to set up an LLM for it to use in the UI, uh, which means that you need to give it your Anthropic key again. Uh, and
- 1:14:40
I picked my inputs and outputs via the single trace method, uh, and then I saved it. One of the things that you can do with online evals, uh, is you can set them to run continuously against new incoming traces. This is something you'll almost certainly do in production. Uh, and if your evals are expensive or you have a whole lot of traffic, you won't want to do it on all of your traffic. You can, uh, set it to serve to, uh, rate only a percentage. That's the sampling rate. So you can say, you know, "Just look at one percent of my traffic," and that
- 1:15:10
will still give you, uh, a whole bunch of eval data without breaking your budget.
- 1:15:17
Uh,
- 1:15:20
so back to the code version. Wiring it up is just like wiring up the other action-- the other, uh, evaluators that we've given. Uh, we've created a classification ev-evaluator. We've given it a name, actionability. We've given it an LLM. Uh, we've given it that template that I just showed you, and we've given it those choices. This is us programmatically doing what we would have done in the UI, where we said, "Your choices are correct or incorrect," uh, actionable or not actionable, scores of one or zero. Uh, and then, uh, we ran
- 1:15:50
evaluate data frame. That gives us, uh,
- 1:15:56
the data that I was skipping over because it was already run earlier.
- 1:16:02
So you can see actionability, actionable. If we look at our evaluations tab, we can scroll down to our actionability, and we get this really great explanation about why the report is actionable or not actionable. Uh,
- 1:16:21
lots of our reports scored actionable, uh, but several are not, which is awesome because it means that we have, again, we have a capability eval, uh, and a hill that we can climb. Um, so in the notebook, uh, this is where you log your results back to AX, or if you've run it online in the UI, your evals-- your eval results are already in the UI as you've seen. Um, so you can see now in ev-- in AX that every span has multiple eval scores. It has mentions ticker, it has
- 1:16:51
correctness, it has faithfulness, it has actionability. Uh, and you can filter and select, uh, for specific, uh, eval results. So
- 1:17:10
eval.actionability.label equals not actionable.
- 1:17:21
That finds you the failing set. This is the set that you're like, "Okay, my eva- my eval says that my code messed up here. I can turn this into a dataset, and I can start running, uh, new prompts against this. I can run experiments against this," as we're going to see. Um,
- 1:17:40
before we move on, I wanna flag a cue-- a few common anti-patterns, uh, that turn evals from a useful tool into noise. Uh, the first one is to treat your eval prompts like code, and I mean that literally. They're very sensitive to wor-to wording. Small shifts in your eval prompt, uh, can lead to dramatic changes in your eval results. Uh, so you should version them. You should test them on examples where you know the right answer. If the judge, if the LLM as a judge disagrees with your human labels, uh, on forty percent of
- 1:18:10
examples, then your eval is probably wrong. It's not that the, your results are wrong, it is that you have written the eval incorrectly. Uh, AX has a prompt playground where you can iterate on, uh, rubrics without touching your code, and it's very useful for doing this kind of iteration of your eval. Uh, the second anti-pattern is the God evaluator. Um, it is tempting to build one big mega evaluator that checks everything at once. It'll look for accuracy and tone and completeness and policy compliance and formatting all in a
- 1:18:40
single prompt. Don't do this because it is a nightmare to calibrate. Uh, when your God evaluator says fail, which of those five dimensions failed? You don't know. So suddenly you're like, "Oh, well, I'll have it spit out labels." And then you're like, "Oh, it's spitting out fifteen different labels." Suddenly, it's way too complicated. Uh, you can't tell, you can't judge by it, you can't score by it. It is much, much simpler and easier to split this up into, uh, one evaluator per dimension. I've already done this today. I have a ticker check, I have
- 1:19:10
faithfulness, I have actionability. Three separate focused evals, and each one tells you something specific. Uh, the combination gives you a richer picture than any single eval ever could. If you find yourself writing a rubric that lists six different things that the response should do, then stop. That is six evaluators, not one. Uh, the third one is know your guardrails and your North Star metrics. Some evals are guardrails. They're ship blockers. If an agent hallucinates a stock price, that is a hard fail. Don't ship on
- 1:19:40
that. Uh, others are North Star metrics. They are aspirational. So always recommend complementary investments is a nice to have in this case. It is not a deal breaker. Uh, so know which of your evals is which. It changes how you act on the results, and later it changes how you set up your production monitors. A guardrail, uh, regression should page someone in the middle of the night. A North Star dip is something that you look at in a weekly report.
- 1:20:06
Which brings us to the next section, which is: Can you trust your judges? This is a question we've been dancing around. I've mentioned it a couple of times. Uh, it's time to evaluate your eva-evaluator. We built an LLM judge, and it told us that some reports are not actionable. Uh, but can we trust it? How do we know that the judge is correct? Uh, the key mental model here is that your judge is a classifier, so you can treat it as a classifier. It takes an input, in this case, a financial report, and it makes a prediction, actionable or not
- 1:20:36
actionable. That prediction can be right or wrong, so just like any classifier, you can measure its performance by comparing its predictions against ground truth, in this case, a human's judgment. Uh, you can check the judge's homework. Um, the thing about applying jud-human judgment is that it is a ton of work. You have to read actually every single output. Uh, you have to think hard and apply a label. Uh, AX does what it can to help by letting you define annotations and attaching them to spans, uh, right in the UI.
- 1:21:06
So let me show you what that looks like.
- 1:21:11
When you are looking at a span, you can click Annotate Span, and you can add and remove annotation configs. In this case, I've already created one called Human Actionable, and you can set this span as being actionable or not actionable. And then you can just use these little arrows to flip through all of your traces and set them as actionable or not actionable, which I have conveniently already done. Uh, this is the kind of work that you can give to a non-technical member of staff. That is why the UI looks like this. They can just click in, read the report, and
- 1:21:41
click a checkbox. They don't have to write any code. They don't need to know anything other than their domain expertise of whether or not this report should be considered actionable or not actionable.
- 1:21:52
What we're doing here when we add human annotations is building a golden dataset. Building golden datasets is incredibly helpful, uh, because it's how you measure, uh, whether your evaluator is actually doing its job. Uh, the way to do it is the same thing, uh, the LLM did. Give yourself real concrete criteria and stick to them. So don't say, "This was good, this was bad." Uh, be specific the same way that you've told the LLM to be specific. You failed at this specific thing. You as-- I-- This thing was supposed to be present, and it wasn't. Uh, and
- 1:22:22
eliminate the chance to get lazy, uh, when you get tired of labeling. Um, as I said, uh, the labeling is extremely, uh, time-consuming, so I did it already. Um, once I've labeled some spans in the UI, uh, I can re-export the spans into the notebook. Um, and this is a nice thing about AX, is that the human annotations that I added by hand come back as columns, uh, in the export. Um, so
- 1:22:52
we filter down to just the rows that I labeled...
- 1:23:02
here.
- 1:23:05
Uh, and now we're going to compare, uh, what I said was actionable and not actionable to what the judge says is actionable or not actionable. Um, before we talk-- before we do that, though, uh, a little bit more about golden datasets. Uh, keep your tasks unambiguous. If your agent scores zero percent consistently, that's almost always a broken task. That's almost always, uh, you've written your rubric incorrectly. Uh,
- 1:23:35
if you've made your task something that only a human can do or worse, that somebody-- something nobody could really do, uh, then you're gonna get a zero score. Um, so for each task, create a reference solution, a known working output that passes all of your graders. This proves that the task is solvable. And you should also test in both directions. You should have cases where a behavior should occur and cases where it shouldn't. So if you test, for instance, does it search the web, uh, then you will end up with
- 1:24:05
a, uh, an agent that always searches the web even when it doesn't need to, whereas you should also have a test that says, "Did it not search the web for this very obvious answer where it didn't need to do that?" Uh, and if you're doing this for real, you should split your labeled data that you just created, uh, into a development set and a test set. Uh, so use maybe seventy-five percent of your labels to itel-to iterate on your eval, uh, tweaking criteria, adjusting examples until the judge agrees with you. Then hold out the remaining twenty-five
- 1:24:34
percent to run the judge and see that it actually gets the scores that you are expecting. Uh, that's how you know that the judge is actually generalizing to new examples instead of just fitting to the specific dataset that you gave it. Uh, this is the same principle as train test splits in machine learning. Don't overfit your evaluator to your golden dataset.
- 1:24:56
Now let's look at our actionability judge.
- 1:25:01
I have run the judge on some examples, and it has come up with, uh, a list of where they agree and disagree. Uh, spoiler alert, I just hit actionable and not actionable at random, uh, so that I would get lots of disagreement because actually Sonnet is pretty good at this stuff. Um, that gives us, uh, an agreement rate of forty-six percent, six out of thirteen times. Um...
- 1:25:29
In real life, uh, this is exactly where you dig in. This is where, uh, you would ask, "Is my rubric wrong, or were my human labels wrong?" When the judge disagrees with you, you should read its explanation and decide whether or not the judge was right or you were right. Uh, that is the most valuable output of meta-evaluation. It usually reveals an ambiguity in the rules that you set down. To fix a rubric, you should read the explanations on the disagreements, find the ambiguity, uh, and tighten the criteria. So instead of
- 1:25:58
just includes forward-looking analysis, I'd write, includes forward-looking analysis with specific recommendations or guidance. That's more precise. Uh, it rules out the case where the report talks about the future but never tells the judge what to do about it. Uh, then I'd rerun the judge on the same examples and see if the disagreement goes, goes away. Uh, this is rubric iteration. Your eval prompt is an application. Uh, it is another LLM application. Uh, so it needs testing and iteration just like your agent does. Uh, and that is what I meant when I said that you should
- 1:26:28
treat your evals like code.
- 1:26:32
So now I'm gonna talk very briefly about an ML concept called precision and recall. Uh, precision asks when the judge says not actionable, how often is it really not actionable? Uh, and recall asks, of all the reports that are actually not actionable, how many did the judge catch? Uh, with a tiny sample like, uh, twelve examples, these numbers are gonna jud- jump around a lot. Uh, but even with the twelve that I have, uh, you can see whether the judge is in
- 1:27:02
the right ballpark or completely off. In most scenarios, you want to, uh, you want to prioritize recall. These two measurements are at odds with each other. If you are-- If you prioritize recall, then you're going to get false positives where the judge has said something is wrong when nothing is wrong. That is much better than accidentally letting through things that are wrong. Uh, however, some-- in some cases when you are tuning you will say, "Actually, I want precision.
- 1:27:32
I want it to be absolutely right when it's right, and I don't care if it misses some things." There are domains where that is the output that you want. But most of the time, uh, you want to prioritize recall. Um,
- 1:27:47
a few known pitfalls with using an LLM as a judge. Um, one is position bias. If you present it with two options to judge between, uh, the judge tend to fav-- tends to favor whichever comes first. Uh, this depends on the model. Sometimes the model depend-- uh, prefers whichever comes last, but it always has some kind of preference to position. Um, there's also length bias. Uh, LLMs like longer responses. They tend, they tend to score those higher and say that, say that they are better even if the length-- extra length is just
- 1:28:16
filler. Uh, and the last one I love, which is confidence bias. The judge gets fooled, uh, by a response that sounds confident just like humans do. If your LLM says things that are wrong but it says them in a really confident tone, uh, your LLM judge is more likely to believe them. Um, there's also self-preference bias. If you use the same model to do the generation as you use to do the judging, they tend to like their own output. So if possible, use both say, OpenAI and Anthropic. Use one to
- 1:28:46
do your agent and one to do the judging, and they are less likely to agree with each other.
- 1:28:53
Uh, self-preference bias is why I used Sonnet as a judge, uh, for output generated by Haiku. Uh, but it works even better if you use a different, uh, labs model entirely.
- 1:29:06
And remember that the benchmark here is human performance. Humans are not going to be perfect at judging whether things are or are not actionable. They are gonna get it wrong quite a lot of the time. What you're trying to do is come up with an eval that is reasonably close to human judgment. Uh, human judgment also varies from human to human. Human interru- inter-rater reliability, so two humans judging the same thing, uh, is often as low as point two or point three. So two experts with the same output and the same rules
- 1:29:37
disagree a surprising amount of the time. So if your judge-- if your LLM judge achieves that, uh, higher consistency than humans, then it is doing better than humans. The judge, uh, sometimes disagrees with me is not by itself a reason to distrust it. You-- What you're looking for is a judge that disagrees with you all of the time. That is the danger signal. A judge that disagrees with you some of the time is a judge, uh, correctly handling a genuinely ambiguous task.
- 1:30:07
Uh, and also failures should seem fair. That is a principle from Anthropic's eval teams that I want to leave you with. When a task fails, uh, it should be clear what the agent got wrong and why. If you look at a failing te- trace and you think that answer looks fine to me, the problem is probably the eval and not the agent.
- 1:30:26
So, uh, now we've built evals, we've tested them, uh, let's use them to actually improve our agent. This is how you close the loop. Uh,
- 1:30:39
here's the problem with one-off fixes. Um, if you've found some failures, you've read some explanations, you know what to improve, you change the prompt, and then what? How do you know, uh, whether the fix actually worked? How do you know if it didn't break something that was working before? If you just run the agent again on a couple of examples and eyeball it, then you're back to vibes. Uh, you need a structured way to compare before and after. And as I mentioned earlier, this is what experiments are for. As I showed you earlier, uh, you
- 1:31:09
can filter your traces down to just the ones that are failing, and you can turn them into a dataset like this. So you select all of your traces, you click Add to Dataset, and you create a new dataset. In this case, uh, I've called it AIEWF Financial Demo Fails, uh, and I've already created my failure dataset. But this gives you a smaller set of places where you know that your eval is failing so that you can then run, uh, your changed prompt against
- 1:31:39
just them. This is faster and cheaper. Um, this is a regression test set. Uh, sorry. You should also save a regression test set rather. You should get the ones where it is not failing and make sure that those are also a dataset that you can check less frequently to make sure that when you've changed your model-- Sorry, when you've changed your rubric, uh, that you haven't accidentally made it fail in places that it was succeeding before. Um...
- 1:32:08
And your data sets are not static. Um, pre-production, when you don't have users yet, your data sets are probably going to be synthetic queries you generated yourself, uh, or queries that you got an LLM to generate for you. That is fine. It gets you off the ground. As soon as real traffic starts coming in, you should turn your, uh, that into your eval sets. Uh, real failures should replace imagined ones. The union of your failure cases and your passing cases is what people call your golden data set, uh, the curated labeled set you
- 1:32:38
trust as ground truth for measuring quality. Uh, now we're going to improve our agent, uh, but we are not going to hand-write the fix. Uh, as I mentioned, you can do all of this with skills. You can do all of this with a coding agent. I'm assuming it being twenty twenty-six that absolutely everyone is using a coding agent to do all of their coding. Uh, so why would you modify your prompts by yourself when you can get Claude code to pull down your failing traces, uh, modify your
- 1:33:08
code for you, uh, and improve your application for you? That is what we're going to do here. Um, Claude will pull down all of the failing traces. It will read all the explanations. It will turn the explanations into actional- action that it takes on your code base. Um, and here's the important thing. Nothing here is guessed. Claude isn't inventing improvements out of thin air. You haven't just told it, "Get better at this." You haven't just told it, "Make no mistakes." Uh, every change it proposes is-- traces
- 1:33:38
back to a specific failure explanation. Uh, so if the judge said that it lacks specific recommendations, uh, Claude will read that and rewrite the prompt so that it does. Uh, if the judge said that it presents risks without supporting evidence, then it will rewrite the prompt so that it fixes that too. This is data-driven prompt engineering, uh, and now we can automate it. Uh,
- 1:34:02
so notice what we've passed in. Uh, we've passed in our requirements, what we were trying to get done, uh, and we have passed in, uh, an improvement prompt telling it, uh, what you should be working on and how to get better at it. This is yet another LLM application and yet another thing that is non-deterministic and yet another way that it can go wrong. You could, if we had more time, meta-evaluate your meta-evaluator of your meta-evaluator and make sure that your meta-evaluator is doing the right thing. But we're not doing that here
- 1:34:32
because it gets too complicated, and also we only have so much time. Um, but, uh, good requirements keep Claude honest, so you should be spending time making sure that this, uh, evaluator, uh, is as accurate as you can make it. Uh, so in this case, we are taking these prompts, and we are wiring them into, uh,
- 1:34:58
our two paths. Uh, we're taking the recommendations that we got from Claude based on pulling the, uh, failing traces, uh, and we are passing them back to Anthropic and saying, "Based on these explanations, write a better prompt. Write a better research prompt, write a better write prompt," and then we're gonna feed them back into our agent. Um,
- 1:35:22
and now we're going to run an experiment. So, uh, over in the UI, we have data sets and experiments. We have my demo financial fails. Uh,
- 1:35:33
we can,
- 1:35:37
assuming I can scroll, wire up an improved agent, uh, using those new improved prompts, so you can see the improved research prompt and the improved write prompt. This is otherwise exactly the same agent that it was before. Uh, and you can create an evaluator, uh, with an experiment. Uh,
- 1:35:57
so what we're doing here is we're giving it a task. A task can be anything. So a task, it could be some subcomponent of what your agent did. So if we were running an evaluator that was about, uh, tool calls, then your task could be just call this tool and see what the output of the tool is, and then you would evaluate that. In this case, my task is run the entire agent again from start to finish, but it doesn't need to be. Uh, and then we are running the same evaluation results, uh,
- 1:36:27
against our, uh, against this new running of the task. So we've changed the agent, we've rerun the agent, and we are rerunning the eval, and we're turning that into an experiment. This is what the experiment looks like. Experiments.run. We've given it a name. We've given it the data set that it should work on. Uh, we've given it the task. We've given it the evaluator, and we've told it to run. Uh, this gives us the results which are
- 1:36:57
here.
- 1:37:00
So for every single run, uh, we've got the output that it got and the new eval score. So, uh, as you can see, this was a data set, if you recall, where previously all of them were not actionable. Uh, the eval has run, and you can see that some of them are now actionable. So we have improved. And if we look at the overall score here, actually, I've nailed it because I had, you know, a lot of chances to get this right. I have one-shotted my prompt, uh,
- 1:37:30
from getting-- being wrong fifty percent of the time to being ri-right one hundred percent of the time. Uh, real evals won't work like that. Uh, real iteration won't, won't run so smoothly. That is why this is a graph. What you're expecting to do is run experiment after experiment and slowly watch the graph climb as your evals get more and more capable. That is the hill that I'm talking about that you're trying to climb. Uh,
- 1:37:58
so the key abstraction behind experiments is that a task is just a Python function. It takes an example from the data set, it runs whatever you run-- want, and it returns an output. So AX doesn't care- Uh, what the task does internally. So it can run the full agent or just one component, like I said, or it could call an API. As long as it takes an input and returns an output, you can run experiments on it. The power of experiments is controlled comparison. So you take the same inputs, the same evaluators, and the only thing that changed is the
- 1:38:28
agent's prompts. Uh, that means any differences in scores is attributable to your change. You're not wondering, did it score higher because of my prompt change, uh, or did it score higher because the web search happens to return better results this time? Well, actually, you're still wondering that a little bit, but 'cause the agent is non-deterministic, so it could have done ten web searches instead of five web searches. But you've eliminated a major source of variation, which is the test cases.
- 1:38:55
So like I said, uh, my agent has one shot at this, uh, which is great, but in reality, uh, it's not going to be that simple. Um, the eval iterate cycle that we're talking about here, where you go through, uh, run an experiment, change your prompt, run the e-experiment again over and over and over, that is where the real value lies. Uh, not in the score itself, but in the cycle. You run the evals, you look at the failures, you read the explanations, uh, you identify a pattern,
- 1:39:25
and you adjust your prompt. Or, as in this case, you get Claude to adjust your prompt for you. Um,
- 1:39:33
did it get better is a question that can only be answered by evals. You have taken the definition of good, and you have codified it in a way that is incredibly useful. Uh, so this allows you to move from "I think it's working" to "I can prove it's working." A practical question that comes up fast, uh, is how many samples do you need? Um, for workshop scale experiments like today, twelve to twenty examples gives you a directional signal. For actual shipping decisions, you're gonna want, uh, more. You're gonna want something like two hundred to four
- 1:40:03
hundred is a good target. But note that two hundred to four hundred is still like a human-sized number that you could possibly accumulate ev-- over a couple of weeks. Uh, it's not tens of thousands or millions like you would use if you were training an ML model. Um, there are diminishing returns. Uh, to halve your mar-your margin of error, you have to quadruple your sample size. Uh, so at some lev-- at some point, you have enough signal to be able to get along, uh, to get along. Um, when you're iterating, where do you
- 1:40:33
invest? Uh, there's a hierarchy. Um, data quality fixes have the highest impact. If your agent is searching the wrong sources or your knowledge base has stale information, then no amount of prompt iteration is going to help. You should fix the data first. Uh, then prompting improvements are next. Few shot examples, explicit instructions, the kind of stuff that I mentioned today, uh, constraints on what the agent should and shouldn't do. These are often the highest ROI change, uh, to your agent. Um, model selection
- 1:41:03
comes third. Sometimes a more capable model solves problems that prompting can't, uh, but it will also cost more and be slower, so that is a trade-off that you have to make. Uh, the thing to not spend a lot of time on is hyperparameter tuning. Things like temperature and top P, it's very tempting to just go for them because they're easy to tweak. Uh, they very seldom make a high leverage change like prompting would or improving the quality of your data would. So you can try them, but they should be the last thing you try. A practice worth knowing is
- 1:41:32
eval-driven development. Write the eval before you build the feature. Uh, if you want your agent to always verify customer identity before it does something, before processing a refund, for instance, then write the eval that checks for that first. That gives you a clear, measurable definition of what good looks like, of what done looks like. Then you can build a feature, uh, until the eval passes. This is the same philosophy as test-driven development.
- 1:41:59
Um, and something I love about eval dri-driven development is that the people closest to product requirements are best positioned to define success. So product managers, customer service, uh, people, even salespeople can contribute to eval tasks. They can tell you what looks good and what doesn't look good in practice. Uh, they don't need to write code. They just need to describe in natural language what good looks like to them, and you can boil that down into your eval rubric. Uh,
- 1:42:30
and so far, uh, I've shown you how to do things both offline and online. Uh, we've been doing everything in a notebook, uh, but you're not really done when your dev experiment trace passes. Uh, you're done when your agent, uh, keeps performing in production on traffic that you've never seen. Um, and it's worth spending a couple of minutes talking about that, even though, uh, we're gonna stick in the notebook, uh, for this workbook-- for this work session. Um, the first thing is
- 1:42:59
online evals. Uh, like I said, uh, and like I showed you earlier, online evals are exactly what they sound like. They take the evaluators that you wrote today, uh, the ticker check, the faithfulness, the actionability, uh, and run them automatically on incoming production traces, the same evals running on new data. Uh, I showed you how to run this on a trace. AX lets you s-- do this at three scopes. You can run your evals at the span level. So you can say this particular LLM call, this particular tool call, run a trace on it every single
- 1:43:29
time. Sorry, run a eval on it every single time. Or you can run it at the trace level, where it has the full tree to look at, uh, and it can evaluate everything from start to finish. Um, you pick the scope that matches the question that you're asking. Um, and you don't need to run evals on every single prediction trace. I, uh, production trace rather. Um, LLM judges cost money, and at production, uh, scale, those costs can add up. So you sample ten percent of your traffic or one percent of your traffic, uh, and that gives you a
- 1:43:59
directional signal without you having to run an expensive LLM as a judge on every single iteration, every single, uh, instance of your application. Um
- 1:44:11
If you don't want to write your evals by hand, uh, we have good news for you, which is that, uh, we have built an agent into AX itself. It is called Alyx. Um, you can describe your eval in plain English, or you can get your non-technical, uh, colleagues to describe a pro-- describe the definition of good in plain English, and Alyx will actually manipulate the UI for you, uh, create an eval, write the rubric, settle the scores correctly, get all the configuration correct, and then just tell you that it's
- 1:44:41
done. Uh, this is especially useful for non-engineers. Um, and the full loop in production looks like this. You go from application, uh, which produces traces, to online evals, which grade those traces and produce labels and explanations. The eval labels feed into monitors, which we didn't cover today. Um, the monitors can alert your team when something goes wrong. Um, the team can investigate, find the failing traces, save them as a regression dataset, um, improve the
- 1:45:11
agent, run an experiment to verify, and ship the fix. And then the loop can run again. Um, as new production traffic accumulates, you will accumulate new failure modes, you will accumulate new edge cases. Uh, you're always going to be i- uh, iterating on your evals and iterating on your production monitoring.
- 1:45:33
But it is a payoff that compounds. Every failure that your online evals catch becomes a new test case. You save it to a dataset, and now it's a regression test for your next round of changes. Over time, this creates a dataset that is unique to your application. Uh, so your specific failure modes, your specific users, your specific quality bar, nobody else has that data, uh, and that is a competitive advantage that grows every single time your agent runs.
- 1:45:59
There's one last pattern that I want to leave you with, uh, which is the most powerful one I know. We did the simplest version today. We fed the judges' explanations back to Claude and had it rewrite two prompts. This works beautifully for two prompts, uh, but a real system isn't two prompts. It's prompts scattered across many files, retrieval settings, tool definitions, all tangled together. Rewriting that isn't a single API call. It's a job for a coding agent, uh, that can see your entire repository. Um, so this is the pattern. You want to export
- 1:46:29
your failing traces from AX along with their explanations, uh, and hand the whole batch to Claude Code or Cursor or another coding agent as context. Uh, this is where the Arize Skills plugin that I mentioned at the beginning, uh, works really well because you don't have, because you don't have to know how to call our API. You don't have to know how to do any of that stuff. You give it the skills and then tell Claude Code, "Hey, pull down some traces, see what's wrong, and ma-make my application better," and it works. Uh,
- 1:47:00
there should be two guardrails on that approach first. However, first, you should feed it your requirements too and not just the failure explanations. So the goal isn't make the evals pass because otherwise it will cheat. It will include all of your test data into its eval, uh, and automatically pass all of your evals. So you have to make, uh, um, your-- You have to make sure that your requirements to your Claude Code are clear that you don't want it just to pass the evals. You want it to do this specific thing, which then makes the evals pass. Um, and
- 1:47:30
second, you should tell it to find themes and not chase individual failures. It is very tempting for Claude Code to look at one particular failed trace and write a whole bunch of code to fix that particular failed trace when there are actually ten failing traces that are all the same cau-cause that it should be focused on instead. Uh, I'm not gonna do all of that time, all of that because I only have twelve minutes left. Uh, but, um, once you have evals running, this is one of the highest leverage things that you can do with them.
- 1:48:01
So let's step back and look at the whole shape of it. Uh, production produces traces. Online evals grade them and produce explanations. A coding agent reads the explanations and proposes fixes, and experiments verify that the fixes work, and then you ship. The production traffic that comes back gets graded by the same online evals that started the loop. Uh, and this is the software development life cycle closing in on itself. Uh, your AI software is helping you improve your AI software, and AX is the substrate, uh, that ties it all together.
- 1:48:31
Your traces, the evals, the explanations, the datasets, the experiments, the online evals, and the monitors.
- 1:48:39
So just to recap, we instrumented an agent with two lines of code. We traced it, and we read our data. We wrote a code eval and built in LLM evals. We wrote a custom rubric from scratch. We validated our judge against human labels. Uh, we saved failures as a dataset. We had Claude improve the prompts. We ran an experiment to prove that it worked. Uh, and we saw how AX takes all of that into production, uh, with online evals. Uh, this is the full loop, and now you know all of it.
- 1:49:09
However, you don't have to do all of it at once. This has been two hours of extremely c-- dense information, and I'm aware of that. You can start very, very small. Uh, you can start by reading your traces. Just add those two lines of code and start looking at your traces, and that is already, that is immediately going to give you information about what your agent is up to that you didn't know. Uh, fifteen minutes of reading real outputs will teach you more about your application than an hour of building dashboards. Then write one code eval. Code evals are easier than LLM evals, and they're faster,
- 1:49:39
and they're cheaper. Uh, and make it for the thing that matters most. So correctness, faithfulness, whatever fits your use case. Run them, look at the results, see what patterns emerge, and build from there. Evals are infrastructure. They are not an afterthought. Some teams create evals at the very start of development. Ad-- Others add them later once they are at scale. Uh, but either way, uh, the important thing is that you treat evals as a core part of your system. They should be as routine as unit tests and the value compounds, but
- 1:50:08
only if you keep investing. Each time a regression shows up before it show-- reaches your users instead of after, you'll understand why all of this work is worth doing. Don't hope for great. You can get to great systematically by specifying it, measuring it, and improving towards it. So now is the time to try this for real on an app of your own. Uh, you've already got an AX account, I hope, by the time the Wi-Fi kicked in. Uh, the docs at arize.com,
- 1:50:39
uh, have companion notebooks with runnable code for everything that we covered today. Uh, and if you want your coding agent to do the heavy lifting, you can install the Arize plugins, uh, the Arize Skills plugin with one NPX command. Uh, and that is in-- that link is in today's notebook. Uh, if you've made it all the way to the end of this, you can also get a free year of Arize Pro, uh, using that code that's at the, on the screen right now, and that is it. Thank you all for your time and attention.