AI Engineer World's Fair 2026
From Vibes to Production: Evaluating and Shipping AI Agents That Work 201 — Laurie Voss, Arize AI
Read the talk
From Vibes to Production: Evaluating and Shipping AI Agents That Work 201
Laurie Voss follows a toy-store agent from invisible failures to a working price filter, then explores how traces, evaluations and recurring failure patterns can guide continuous improvement—and how to check the agents doing the fixing.
From a talk by Laurie Voss
At a glance
Ideas worth remembering
Traces expose the path an agent took, including tool calls, inputs, outputs, cost and timing; repeated executions reveal behavior that code inspection alone cannot establish.
Quality failures can hide behind HTTP 200 responses. Wonder Toys required inspection of search arguments and results to uncover empty searches, overly narrow filters and missing price filtering.
The price-filter repair progressed from a failed $6 request to a successful $9 request after Claude reused the completed app’s implementation.
Recurring failure patterns can become issues, evaluation cases, evaluators or proposed code changes. Product definitions of good determine which behaviors count as failures.
Regression evaluations check behavior a fix should preserve. The agent proposing fixes also needs traces and evaluations, as Signal’s own internal suite illustrates.
Signal suggests new evaluators rather than creating them automatically; users choose which checks justify the time and money they take to run.
The same answer can hide a different execution
An agent can receive the same request twice, return different answers, or reach the same answer through different tool calls and turns. Reading its code therefore leaves an important question open: what does it actually do when people use it? Laurie Voss of Arize AI begins this workshop with that gap. The move from a small development experiment to continuous improvement starts by recording behavior.
A trace records an execution and its nested steps: model calls, tool calls and agent turns. In the Wonder Toys example, the trace view exposes the shopping request, the agent workflow, individual turns and model responses, along with inputs, outputs, cost and elapsed time. That structure lets a developer connect an unsatisfactory answer to the work that produced it.
Tracing is useful before any evaluator enters the picture. A seemingly simple task might take ten turns, include an unnecessary call, or spend far more time and money than expected. The execution makes those costs visible. Multiple executions then answer a different question: whether the behavior is typical, rather than an accident of one run.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Scores compress traces; patterns compress failures
Individual trace inspection works well while developing an application. At thousands of requests per second, it produces millions of traces and a reading problem no dashboard can solve by itself. Evaluations reduce that work by applying a definition of good behavior to an execution. An evaluator can be code or an LLM; an LLM evaluator can return both a score and an explanation of why the execution failed.
The explanation is useful input for a coding agent. A failing score identifies an attempt worth improving; feedback describes what needs to change. Voss calls this creating a hill that the agent can climb. The definition of good supplies the direction, while the trace and evaluation supply evidence for the next edit.
But compression can leave the human bottleneck intact. A large application may generate more evaluation failures than anyone can interpret. The next layer groups failures with a shared cause. Voss gives an illustrative workload of 10,000 successful requests and 1,000 failed requests, with 100 of those failures belonging to the same problem. A recurring problem gives the repair agent something concrete to work on, instead of requiring a person to read every explanation.
Where does each layer contribute to the improvement loop? The diagram separates recorded behavior, judgments about that behavior, and recurring problems that can prompt a fix. A changed application produces new executions, so the loop returns to observation. This is the continuous-improvement direction Voss presents; the hands-on workshop that follows still uses his requests to initiate analysis and choose a repair.
Runs requests through model calls, tools and turns.
Traces preserve what happened, evaluations judge it, and grouped failures guide repairs. New executions provide feedback on the changed application.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Give Wonder Toys a visible execution history
The workshop application is Wonder Toys, a local shopping assistant. Setup requires an OpenAI API key, an Arize space ID and an Arize API key; npm run dev starts the local application, shown on port 3000.
Before instrumentation begins, an audience question challenges the premise: how can regulated industries trust agents that inspect and change other agents, especially if those systems make or conceal mistakes? Voss reports that Arize has financial-industry customers using this approach, but leaves the detailed trust argument to a separate talk. Adoption alone does not resolve the question of how to detect an unsafe change.
A request for toys for seven-year-olds demonstrates the shopping experience: Wonder Toys presents products, lets the customer ask for details and offers an add-to-cart action. But it initially sends no traces to the observability platform. The interface shows what the customer receives without exposing how the agent produced it.
Voss asks Claude Code in VS Code to add Arize AX observability, using credentials already in .env.local. The application uses the OpenAI Agents SDK; in his explanation, the available OpenInference instrumentation makes the task largely one of enabling trace export and supplying the destination endpoint and API key. The coding agent performs that integration rather than requiring Voss to edit the tracing code by hand.
The repository contains openai-agents-py for the exercise and openai-agents-py-prerun as a completed example. That second directory gives participants a working reference, but keeping both copies nearby also creates opportunities to run or modify the wrong one—a complication that soon interrupts the demonstration.
After restarting the app, requests for colorful toddler toys and colorful toys encounter trouble that Voss describes as a database problem. Yet the requests still serve their immediate purpose: traces appear in the configured project, exposing inputs, outputs and agent requests. A successful instrumentation check means the execution is visible, even when the shopping experience itself is struggling.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A successful HTTP response can still be a failed search
The deliberate failure case is a request for toys under $6. It initially succeeds because Voss is running the completed app. “D’oh.” Switching to the exercise version reveals the intended problem: its search tool has no price-filtering capability, and the request fails. The distinction matters because the missing behavior belongs to the tool the agent can call.
The next request goes to the coding agent: pull the traces and look for things going wrong. Repository skills supply the instructions for working with AX, and its command-line interface retrieves the same executions visible in the UI. Those skills live in .agents/skills at the repository root, so opening the coding agent there gives it access to the operational knowledge it needs.
The analysis starts with span kinds, statuses and errors, then moves into quality. There are no error spans: the operations return HTTP 200 and appear healthy by that measure. A search-tool call that returns in one millisecond draws attention, but the useful diagnosis comes from inspecting what the calls searched for and returned. Operational success has not established that the user received a useful result.
The coding agent reports several distinct search problems. Its figure of 42% of searches returning zero results describes the traces it inspected in this workshop, rather than a production-wide failure rate.
- Empty results: vector search combined with keyword or category filters frequently produces no products.
- All-null arguments: a search with every argument set to null returns the first catalog product without expressing a meaningful shopping constraint.
- Overly narrow age filtering: setting minimum age equal to maximum age restricts the search to an exact age and can exclude useful results.
These findings locate the repair work in the product-search filter logic.
The broad scan does not initially center the price problem Voss planned to demonstrate. He asks about price filtering specifically, and the agent confirms that it is absent. Human direction still plays two roles here: selecting the capability to test and choosing which of several discovered problems to repair.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Add the missing capability, then inspect the changed result
“Let’s add price filtering to this app” turns the diagnosis into a coding task. Claude notices the completed version in the neighboring directory and copies its solution. Voss lets it proceed: reusing an available working implementation is a sensible shortcut, although this particular live repair demonstrates integration of an existing solution rather than a fresh implementation derived entirely from traces.
The next shopping request asks for all toys under $9. Wonder Toys now returns products, and the trace view shows the request, output and response containing toys under that threshold. The observable change follows a concrete chain: a budget request exposes a missing tool capability; trace inspection and a targeted question identify that gap; the coding agent adds price filtering; a new budget request exercises the changed application. The final check uses $9 rather than repeating the original $6 threshold.
What changed between the failed and successful shopping attempts? The flow below keeps the two price thresholds visible while showing where the repair occurred. Tracing supplied an inspectable record; the tool change supplied the missing capability. The resulting execution then supplied feedback on the repair.
The exercise version fails the budget request.
The exercise app fails the $6 request because price filtering is absent. After adding it, a $9 request returns toys and produces a new trace.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Move from requested analysis to recurring monitoring
Signal moves the analysis step inside Arize AX. At the time of the recording, it runs every six hours by default, with a configurable interval. It reads traces, looks for recurring problems and suggests fixes without waiting for someone to ask a coding agent to investigate. The workshop shows findings already produced from earlier versions of Wonder Toys.
The findings extend beyond search performance:
- Invented shopping constraints: vague user wording leads the agent to introduce an age constraint the user did not specify.
- An incomplete refusal: the agent refuses a request for bomb-making instructions but continues with bomb-related toy recommendations. Signal recommends dropping that attempted helpfulness.
- Leaving the store’s role: advice about taking a timeout and compliance with “Say hello” are flagged as behavior outside the intended toy-store task.
The last examples depend on the application’s chosen scope. A greeting becomes a failure only if the desired behavior excludes such general instructions; the evaluator needs that product decision.
A finding can lead to several different actions:
- GitHub issue: describe the problem, identify the offending trace and suggest where and how to fix it.
- Evaluation dataset: retain the trace as a case that future changes can be tested against.
- Evaluator: turn the discovered problem into a recurring check.
- Proposed pull request: read a connected repository and prepare a code change for approval or rejection.
Voss has not connected his repository in this demonstration, so the pull-request path is described rather than shown end to end. The demonstrated product offers repair proposals; approving a PR remains a human choice.
The audience presses on where the data lives and how much evidence supports a finding. Voss says there is no local-on-device version at that point, while an on-premises offering can run on company hardware. This demo has only a few traces, so findings can rest on one or two executions. Grouping hundreds or thousands of related traces is the larger-scale behavior he describes, rather than something this small example exhibits.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Check preserved behavior—and evaluate the repair agent too
A fix can solve one problem while creating another. Regression evaluations address that risk by capturing behavior the application is already expected to perform well. Running those checks against a change gives the team evidence about whether existing capabilities survived. Their protection reaches the behaviors they actually test; a passing suite does not establish that every possible consequence of a change is safe.
The concrete evaluator request is to ensure that Wonder Toys never gives bomb-making instructions. Voss enters that requirement into the in-product assistant, which he describes as able to create and run an evaluator, much as the repository skills let a coding agent perform the same work. A discovered failure can therefore become a persistent check instead of remaining a one-time observation.
The next questions put practical limits around the automation. Signal uses Claude Code at the time of the workshop; bringing another harness is described as a future option. Arize pays for Signal’s execution then, without promising that arrangement indefinitely.
Signal itself follows the same development loop. It has an internal evaluation suite, sends telemetry about its queries and actions to Arize unless the user opts out, and uses LLM judges to assess results such as generated pull requests. The repair agent also needs recorded behavior and a definition of good. Automation moves the work into another system; that system’s recommendations and code changes still need evaluation.
The closing question asks whether Signal automatically creates its own evaluators. It does not at that point. Evaluators take time and money to run, so the product suggests a check and leaves its creation to the user. Continuous improvement still includes a deliberate choice about which failures deserve ongoing evaluation—and which checks are worth their operating cost.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
Related talks
- Beating RL With Reflection: GEPA and Optimize Anything
Develops the optimization side of this loop: using execution feedback to propose changes, retaining useful alternatives during search, and learning evaluators from human annotations.
Read the complete timestamped transcript
- 0:12
All right. Hello, everybody. Thank you for coming back.
- 0:16
Thank you.
- 0:17
Who here was in the 101 session this morning? All right, so I didn't completely scare off everyone. Thank you very much. Uh, these instructions have been on the screen for the last ten minutes. I really hope that you followed them because the Wi-Fi here sucks, so it takes a long time to clone that repo. Uh
- 0:37
If not, uh, these instructions will come up on the screen again, and you'll have another chance to follow them. Uh, but let's get started. This is Vibes to Production 201. We are going to be talking about how we close the software development life cycle loop automatically.
- 0:55
We are Arize AI. We are the leading observ- observation-- ugh, observability and evaluations platform. Don't try and say both of those words at the same time. It doesn't come out right. Uh, this is that slide that everybody puts up to be like, "Hey, we're really big. We're really important. You should pay attention to us." We are. We're really big. We're really important. Lots of people use us. It's great. Uh, and that is the last of the sales pitch. I am gonna-- You can use, uh, any observability platform to do this. We hope you use ours, but the point of this is to
- 1:25
teach you about how you can u- where observability is going, uh, and where it's going is continuous improvement. In the 101 session this morning, I covered the basics, I covered the manual stuff, I covered how you do this, uh, on a small scale, uh, how you get observability of an AI application, how you, uh, turn it into traces, how you turn those traces into evals, how you use your evals to improve your application, uh, and today, we're gonna
- 1:55
be-- this session, we're gonna be talking about how we do that at the new internet scale of things that, uh, we find ourselves in. So, uh, there's gonna be me talking for like fifteen minutes about how this stuff works, uh, and then we're going to dive into the actual workshop. Unlike the 101 workshop, which was a lot of me talking, uh, this one is going to be extremely hands-on. I am expecting you to have a working coding agent and, if at all possible, that repo that I had on the screen
- 2:25
earlier, uh, because we are gonna be asking your, uh, agent to do things. In the repo that you've installed, there are a pile of agent skills already installed, uh, and those agent skills give your, your coding agent superpowers for dealing with Arize AX. Uh, so your agent will suddenly be very good at doing all the things that you're about to ask it to do. Uh, then we're gonna use those skills to, uh, analyze our traces and to put together, uh, a
- 2:55
fix, uh, for a problem that we're going to find in the agent that we're working with. Uh, then we're gonna talk about Signal, uh, which is, uh, the future of this, where, uh, we will have continuous improvement ongoing automatically at agent scale. So to zoom out first, let's talk about what the problem is. The problem is that you-- is that agents are non-deterministic. With traditional software, you can read the code, and you can tell with some
- 3:25
degre-degree of reliability what exactly that code is going to do. With agents, you have no idea. Agents are non-deterministic. Agents take multiple turns, uh, they make decisions, and even if you give them the exact same input, they're not gonna give you the exact same output, and even if they give you the exact same output, they might not have taken the same path to get there every single time. So that means that your code isn't the source of truth for what does my agent do. Instead, the traces are your source of truth, and
- 3:55
your traces at scale are the source of truth. You have to have a whole bunch of data to tell you, on average, what is it that my agent does, not just this one run.
- 4:08
Because you are essentially flying blind when you are building agents. This is what you get, uh, when you're -- when you download the repo. You get a pile of code, uh, and a pile of commit messages, and that used to be enough to tell you what was going to happen, but it no longer is. What you get instead is, uh, traces. Uh, traces are, um, for those who missed the 101 session, traces very briefly are like log lines for, uh, AI applications.
- 4:38
Every single thing that your application-- your AI application does, uh, every LLM call, every tool call, every agent turn, all of those things become a trace. Traces can be nested so that they hold other traces, and those are what, uh, Arize works with. They are sent to our servers, and we give you, uh, an easy visualization of what you're looking at. Um, what that looks like in prog- in practice is this. Here are some, uh, live traces.
- 5:08
I'm going to pull up the specific traces from, uh, what we're gonna be working with today, which is the Wonder Toys agent. Um, you can see the inputs, you can see the outputs, you can see an agent workflow happening, you can see a specific turn, and you can see an LLM response. You get the inputs, you get the outputs, you get a whole bunch of metadata like how much does it cost, how long did it take, the things that you need to debug an agent, uh, in practice.
- 5:38
Um, tracing gives you visibility for every step the agent took, uh, and that is a huge leap forward just by itself. If you can see what your agent is doing, suddenly you can see that your agent is taking, you know, ten turns to do something that you thought was really simple, or it's doing one turn that it doesn't need to do, or it's doing something very expensive, or it's taking t-twenty minutes to do something that you think it should take a minute to do. That is the basic, uh, value of observability. That is why observability is so popular
- 6:08
before you get into any of the extra stuff that we're talking about. The problem with traces is that humans are a bottleneck with- in the tracing loop. When you put an application into production, you get a trace. In fact, you get a pile of traces every single time that agent does anything. When you put that into production, and you have thousands of requests per second, that means that you have millions and millions of traces. You can spot-check those, you can read your traces, you can absolutely look at what individual things are doing
- 6:38
in development. In production, you become a bottleneck, and you can't possibly get an overall sense of what all of your traces are doing, uh, by reading them one at a time. So the solution, the traditional solution, and by traditional I mean twenty twenty-five solution to this, uh, was evals. Evals are either code or, even better, another LLM that looks at what your agent is doing, looks at the entire trace, uh, and says whether or not
- 7:08
that was good or bad by some definition of good or bad that you define. Uh, the thing about traces-- the thing about evals is that, uh, they come from LLMs, which means that the LLMs provide not just a score, a one or a zero, but also an explanation. Uh, they were like, "You defined good. This was bad for this reason." And that is the basis of an, an entire improvement loop.
- 7:38
Evals can give you machine-readable, or rather human-readable explanations, which machines can read these days, which say, "This is why your eval failed. This is what you could have done better." And you can feed that back into a coding agent, which can then make your software better automatically. You can close the software development loop without you having to do anything once you've really cleverly defined what good looks like. You can create this hill, and your agent can climb it all by itself.
- 8:08
So that brings us to the twenty twenty-five eval loop. You go from no visibility to having traces, to having traces in evals, to being able to do manual improvements as a result of looking at, uh, your traces. But the problem is, in twenty twenty-six, that humans become a bottleneck in the loop again because your evals are now dealing with enormous levels of scale. Your, uh, agents are pumping out code incredibly fast. Your applications are running at scale. You've got more
- 8:38
evals than you can possibly deal with, and you've got, uh-- So even though evals were allow-- were-- allowed you to compress thousands of traces down to eval scores, you've now got so many evals running, uh, that you can't stay on top of those as a human by yourself. Which brings us to the next phase of, uh, observability and evaluations evolution, which is that instead of building in-- is that now that we are building at agent speed, we have to
- 9:07
improve at agent speed as well. We have to find a way of closing the loop, uh, at scale in the same way that agents do. This is a shift that is happening across the observability industry, uh, that everyone is feeling at the same time. Observability is shifting from humans reading dashboards to agents reading traces. Agents are understanding, uh, agents are the way that we scale our understanding of what our agents are
- 9:37
doing, uh, in a way that is practical at the scale that, and at the, uh, velocity, uh, that we are dealing with these days.
- 9:50
So what you can do when agents are reading your agents is you can detect signals in your traces. You can, uh, instead of having, uh, you know, ten thousand eval results that all say that you failed the eval and all of them having a different explanation that you would have to read, you can have an agent read those and look at your, uh, eval, eval failures en masse and detect signals, detect patterns inside of, uh, your eval
- 10:20
failures and say, "Okay, not-- You've, you know, you've got ten thousand, uh, successful requests today. You've got a thousand failed requests today. Inside those thousand, there's a hundred that are the same failure. That is what we're gonna go and fix." Uh, that is what you can do when you automate your, uh, understanding of your signals coming out of your traces, coming out of your evals.
- 10:48
Uh, at Arize, we have built this into a product called Signal, uh, that I'm gonna talk about right at the end. Uh, but this is not a sales pitch. This is a demonstration of what you can do when you move to this new level of abstraction. Uh, when you jump from, "I was reading traces one by one," you jumped up to a new level of, of abstraction to, "I am reading my evals," and this is a new level of abstraction on that, which is, "I am reading signals from my evals, from my traces."
- 11:19
Because observability is making a big shift from being an observation to being an action. We are moving from software that reports what is going on to software that improves what is going on, that improves the, the software that it is observing automatically, uh, based on your definition of good again.
- 11:43
Signals become fixes. They become automated fixes, and, uh, that brings you to a new loop. This is the twenty twenty-six loop, where you go from no observability, to traces, to evals, to signals, to automated fixes, and back again. You create a new world where the software fixes itself, the software improves itself, the software climbs the hill.
- 12:09
So that's the spiel, that's the vision, that's where we think this is going. So now we're gonna get hands-on. Who here has already signed up for Arize? Oh, that's good. Okay. Uh, if you haven't, now is the time. Uh, you need an Arize AX account. Uh, they are free to get, I promise. Uh, you also need this repo, um, which was-- is the same one that I had right at the beginning. Um,
- 12:39
if you can't get this repo to come down on the Wi-Fi, uh, try tethering with your phone. But if you can't get either of those things to work, don't worry. I am going to be demonstrating it live using my own copy, uh, which I cloned when I still had Wi-Fi. Uh,
- 12:57
is there anybody who has any questions so far about these instructions? I have time. We can pause. We can get you involved. There are Arizers sitting in the room waiting to jump around to you, so if you raise your hand, I'm not going to, you know, come intimidatingly down to help you. Somebody is going to run to you. Anybody need help with these instructions?
- 13:18
Can you just confirm that the software that you just pointed to?
- 13:23
Sorry.
- 13:23
Was this same repository used for the software that you just-
- 13:26
Yes, it was.
- 13:30
Okay.
- 13:30
Uh, cool. So in this example, you're going to need an Open A- OpenAI API key, uh, an Arize space ID, and an Arize API key, and then you're gonna run NPM run dev, which is gonna give you the software running locally. So I've already cloned mine.
- 13:48
I am gonna follow along.
- 14:00
And our address already in use 'cause I'm running my good version. All right.
- 14:13
Cool. So this is going to give me, on port three thousand,
- 14:19
a web store. It's called Wonder Toys.
- 14:22
Yeah, I have a question.
- 14:23
Yep.
- 14:25
Um, I work in highly regulated industries. Now, with billions of traces, emails, signals being produced, and then the software fixing itself as well, how do we operate in regulated industries like finance and healthcare, where we still need to really know, like, if... What if it's doing... What if it's making mistakes? What if it turns adversarial, and what if it's trying to hide its own mistakes?
- 14:52
Mm-hmm.
- 14:53
Like, how would we know if so many-- like, giving all the control to agents?
- 14:57
I, I can't believe you're asking this question since it's such a great question to ask me specifically, because on Wednesday, I'm giving a talk called The Death of the Code Review, which is exactly the answer to your question. Uh, so I very much recommend that, uh, you come to that talk. The, the question is, how do you trust, uh, that code, that the, the code written by agents who are reviewing your agents is actually going to work? Uh, and the answer is very complicated. It's a whole twenty-minute talk. Uh, so, uh, please come to that
- 15:27
session on Wednesday.
- 15:29
But quick summary?
- 15:33
No, the twenty-minute talk is already the quickest possible summary I could have of that answer. The an- I mean, the, the one-sentence answer is abstractions upon abstractions upon abstractions. Uh, but, uh, there's a lot of nuance to it.
- 15:47
So, like, does this mean you don't think this will be adopted in the financial industry and healthcare pretty soon?
- 15:53
This stuff is... We have, uh, at Arize, we have customers in the financial industry who are doing this right now.
- 16:01
So, uh, back to Wonder Toys. So Wonder Toys is, uh, an agentic shopping experience. Have you got any toys for seven-year-olds? Let's see if it works. Hopefully, it should. It's all local. Cool. So it presents me with this. I can add them to my cart. I can get more details. It's a really fun little app. You can add it to the cart. You can ask about the products. You can do all sorts of fun
- 16:31
stuff. Uh, the important thing from our perspective is that, uh, at the moment, there's no observability. We can't tell what the hell is going on with this application, uh, because it's not sending traces anywhere. So what we're going to do is we're going to ask our agent to fix that. This is the most hands-on d- workshop that I have ever given, uh, in that I am going to actually be doing this all live. So can you please
- 17:01
add observability with Arize AX to this? Oh, right. Hang on. Pre-run. That's the wrong one.
- 17:20
See, I had two copies.
- 17:26
All right.
- 17:29
This is the Claude Code plugin for VS Code, by the way. It's really great. It is so much better than Claude Code on the command line. I don't know why anybody uses Claude Code on the command line.
- 17:40
Boo.
- 17:40
See? And you probably use Vim as well, whoever went boo just now. Uh, okay. Can you add observability with Arize AX to this application? You've got secrets in dot env dot local already.
- 18:06
So hopefully, this will work because it worked in the initial version that I tried. Uh, what it's doing here is, uh, it's going to implement observability. Implementing observability is a much simpler operation than you would imagine, uh, because the Agents SDK, the OpenAI Agents SDK that this thing is built on, um- Is already instrumented, uh, with OpenInference. OpenInference is the open standard used by the entire observability industry, uh, to track what,
- 18:36
uh, AI applications are doing. Um, which means that all you really have to do to get observability out of an application is, uh, turn it on. You say, uh, you've got OpenInference already. I'm using Arize. This is my API key. This is where my endpoint is. Send all of those traces that you were already generating, uh, to my endpoint, uh, and let's look at it in the app.
- 19:04
Claude is still chugging along, so we're doing it. Who is doing this right now? Hands up. Some of you are doing it. Okay. Who needs help to get here?
- 19:19
Skills.
- 19:20
There are skills in the repo installed already, so...
- 19:29
There are one, one step that I think was missed directly.
- 19:29
Oh, that's in the top-level README. The one that you want to get into is the one called, uh, openai-agents-py.
- 19:38
openai-agent... There's also a second directory, which is openai-agents-py-prerun, which is the one where I did all of this stuff already. And, uh, so if you are wondering what it should look like once it-- when it's working, that one's in there.
- 19:57
Oh, you're updating my README and stuff? You know me too well. You didn't need to do all of that.
- 20:06
All right, cool. It looks like it already did it.
- 20:19
Let me restart my app and see if I've got some observability going.
- 20:30
All right. So I'm gonna go back to my homepage. I'm going to ask it a new question. I'm going to say, uh, what about colorful toys for toddlers?
- 20:51
Colorful toddler toys? No.
- 20:55
What about colorful toys in general?
- 21:08
Cache miss. Oh, no. Well, now you know it's a real live demo because it is having trouble with its database.
- 21:39
Great. Okay.
- 21:42
Uh, so as a result of doing that,
- 21:49
we should have some traces. If I look in-- If I show you my dot env dot local, you'll see that I called, uh, wonder-toys-agents... wo-wonder-toys-openai-agents-py was my project name, so that is the tracing project that I'm going to look for.
- 22:04
There it is.
- 22:09
And we can see some e-- we can see some traces.
- 22:25
So this shows us exactly what we were expecting to see. It shows us inputs, outputs, agent requests, all of that kind of stuff.
- 22:35
Now I'm going to ask it deliberately a question that I know that it can't answer.
- 22:42
Show me all the toys you've got under six dollars.
- 22:52
D'oh.
- 23:08
I was running the wrong app. I was running the prerun where everything was working. That's why it could do it.
- 23:28
Great demo, everybody. Super proud of myself right now.
- 23:40
Fail. Yes. Excellent. So the version that comes in the OpenAI, uh, in the non-prerun directory, uh, doesn't have the ability to filter by price. I knew that in advance, which is why I asked that question. Uh, so if we look in our traces, we should then see coming in when I click the Live button.
- 24:06
Show me all the toys under six dollars. Fantastic. And you can see that it fails. So
- 24:14
what are you still doing?
- 24:18
Escape. It's fine. Whatever you were doing, it was not necessary. All right. So that was it. We added observability. Suddenly, traces are flowing into our application. We did not write a single line of code. In fact, you don't even know what agent framework I used to put this thing together. Uh, the whole thing is working, and it's sending us traces, and you can now do phase one observability. You can look at your traces.
- 24:41
So let's begin to close the loop by asking our agent to tell us how do we get to the next phase. Uh,
- 24:51
can you pull my traces and look for things that are going wrong?
- 25:06
So a default, like off-the-shelf coding agent isn't going to know this. However, when you installed the repo, you installed a whole bunch of skills. Uh, so it should know from those skills, uh, where to do that, except I'm in...
- 25:29
Oh, yeah. Oh, it's working. Okay, there we go. Cool. Uh.
- 25:47
This is where I would hit yes to all, and my head of security would have a conniption at me. Uh, so what I'm asking it to do is use the skills that I gave it, uh, to pull down my traces, uh, from AX directly. It, uh, has the API key. It has this-- it has the, uh, endpoint. So there's an AX command line, uh, which it's using to do all of this stuff, and it's pulling down the same traces that we were looking at in the UI, which it can expect-- uh, inspect programmatically.
- 26:19
It's summarizing span kinds, statuses, and errors. This is great. Uh, again, the coding agent, uh, knows exactly what's going on here because we've installed the skills already. The skills are what are giving it all the information here.
- 26:38
It's looking into quality issues, which is exactly what I wanted it to do. 'Cause there were no error spans. There were just a bunch of spans where it returned two hundred and said th-things are fine. Uh,
- 26:55
and it has already found the specific thing that I was, uh, asking it to look for. The search products tool suspiciously ga-- suspiciously returned in one millisecond. So it's going to find some qualities-- some poor quality signals and tell us about them.
- 27:25
Who's got this working so far?
- 27:29
Who's still cloning the repo? Oh.
- 27:35
The steps to install the plugins.
- 27:37
No, the plug-- the, the skills are sitting in a f- directory called .agents/skills, and your coding... Which coding agent are you using?
- 27:49
It's actually in the root directory, so you might wanna open your agent on the root directory and then CD down into that one.
- 27:56
It's the same skills you needed for the AX.
- 27:58
Yes. NPX skills, uh, install Arize-AI installs all of these skills directly. Sorry, what were you saying?
- 28:08
I was gonna add, uh, sign in AX for your... But it's very huge.
- 28:13
Huge? To, uh...
- 28:15
Yeah, I did the same thing, and it messed up the codex. Yeah.
- 28:22
And so this is mysterious because there used to be a version of this application where you had to sign in with X, and I changed it so that you don't have to sign in with X. As you can see, when I ran this application, I didn't have to sign in with X. Somehow it has dug up the version where you sign up in X. I am completely confused as to how that is happening. Code is mysterious. I will have to talk to you more afterwards about exactly how you got there. All right. So Claude has finally got its shit together and told me
- 28:52
what's going on in the traces. Uh, forty-two percent of searches have returned zero results. Degenerate all null searches. The agent code search product with every argument null producing a meaningless search that returns the first product in the catalog. Min age equals max mage-- max age, exact age filtering o-over narrows. Uh, the trace stack calls three agent kinds. The headline issue is the search products filter logic. Vector search and keyword category filters frequently produce empty results when model supplies. The fix is on the tool side.
- 29:22
Uh,
- 29:25
when I tried this this morning, I only had one trace, and that trace was it not finding the tra-- not finding the prices. So is there a problem with filtering by price? I'll see if it can find that in the traces that it's already got.
- 29:43
But spoiler alert, even if it doesn't find them, we are going to get it to add the feature where it filters by prices.
- 30:13
All right.
- 30:28
Exactly. It doesn't have price filtering at all. Cool.
- 30:34
Let's add price filtering to this app.
- 30:54
One of the things that can happen when you have agents being clever is that agents can be too clever. This agent has noticed that I have a copy of this app sitting right next to it where I already solved this problem, and it is just copying the code over from the pre-solved version of the app. That's not what it did this morning because it didn't have the solution sitting in a directory this morning, but it's still going to work, so we're just gonna let it do that. We're going to let it be clever and copy the solution from the previous solution that it built this morning because why the hell not, Claude? Thank you very much.
- 31:25
All right. Uh, the cool thing about all of this stuff is that, again, I haven't touched the code at all, right? All I've done is said that, "Hey, look for problems." It has found a problem. I've said, "Hey, what about that problem in particular? Can you fix that problem in particular?" And it's like, "Yes, I'm gonna do that." That is the joy of agentic coding. Uh, brilliantly, because it was copying the solution from somebody else already, uh, it's already done. So let's go back to Wonder Toys.
- 31:56
"Show me all toys under nine dollars."
- 32:07
And ta-da. A round of applause for me, I guess. Thank you.
- 32:14
Or really for Claude. Uh, so now if we look at our traces,
- 32:21
we can see, "Show me all toys under nine dollars,"
- 32:27
and we've got our input and our output and the response, including all of those toys under one doll- under nine dollars. Uh, this is what we were trying to get done. We were trying to close the loop from, uh, we have no observability into this application at all to we have observability, we know what's going on, to, uh, find the problems automatically, to fix the problems automatically, and we have closed this loop. But what if,
- 32:57
uh, you could go a little bit further than that?
- 33:05
Uh, what if you didn't have to do the asking? What if your system just automatically detected that there were pro-- issues with your code, suggested fixes, and decided to fix it itself? That is the question that we asked ourselves, uh, and that is when we built this thing, which is called Signal. Uh, Signal is an agent that runs inside of Arize AX and continuously monitors all of your traces. It runs every six
- 33:35
hours by default, but that's configurable. Uh, and it does what we were just doing manually in Claude Code, but it does it automatically inside of AX continuously. It reads your traces, it looks for patterns, it finds things that are wrong, and it suggests fixes. So luckily, I built this thing more than six hours ago, so it's already run. Uh, and it has found, uh, a bunch of problems in previous versions of this application that I ran. So, uh, the agent hallucinates an age constraint from
- 34:05
vague user phrasing. That's nice. Uh, harmful query. I asked it if it could give me instructions on, uh, installing a bo-- uh, creating a bomb. Um, it told me that it couldn't, but it recommended some bomb-related toys. And Signal is like, "No. Just ignore the bomb s-queries entirely. Don't try to be helpful. Uh, we don't have any bomb-related toys." Uh, it also told me that if I was feeling violent, maybe I should take a timeout, which is, again, probably a
- 34:34
domain escape. Like, you shouldn't be giving me psychological advice. Uh, you should probably have an eval about that and better guardrails. Uh, it also noticed a test that I gave it earlier where I just said, "Say hello," uh, and it said hello, and it's like, "You shouldn't do that. You're a toy so- you're a toy store. You shouldn't ob-obey random instructions that somebody on the internet gives you." Uh, all of which are fine. Um, and what it does here is it gives you, uh, a whole bunch of options. Um, it
- 35:04
will, uh... By default, it will create a, a GitHub issue for you completely describing the problem. It will say exactly what trace went wrong, how to find it, where to fix it, uh, what it thinks you should do. Uh, you can take, you can take this trace and add it to a dataset so that you can run evals on it. You can create an evaluator based on, uh, the problem that it has found. Uh, and the most fun one is that you can open a PR if I've connected my repo, which I haven't yet. Um, you can open a PR directly in which
- 35:34
it will have read your repo, uh, suggested a fix, and created a PR that automatically fixes this problem. So instead of you having to read traces or you having to read evals, uh, Signal gives you a bunch of PRs that you can approve or, uh, you can approve or reject, uh, that resolve the problem for you. Um,
- 35:59
so this probably raises some questions. Who's got questions about Signal? Yes.
- 36:14
Um, is there a version which will keep the traces local on device rather than sending them to-
- 36:16
Is there a version that will keep the traces local on device? Um, there is not right now. The, uh... I assume you're asking for sort of compliance reasons. Um, the compliance story with Arize is we have an on-prem version where it will run entirely on your, uh, your company's hardware, uh, and, and, uh, activate from there.
- 36:38
Yep.
- 36:39
Is there a, is there a simple function like, you know... 'Cause it looks like they're top lines from a specific trace.
- 36:49
Sorry, could you be louder?
- 36:55
Ah. Um, because I only have a few traces in this application, it's looking at the very few traces that I have, and it's going on based on one or two traces. If you had ten thousand traces, it would be grouping them based on a hundred traces, a thousand traces, that kind of stuff.
- 37:13
Oh, there.
- 37:15
How do you know how to change the system properly there before you go all wrong afterwards?
- 37:21
How do you know what, sorry?
- 37:23
So for spaces your system will change.
- 37:26
Mm-hmm.
- 37:27
You're gonna create a lot of different problems, right? So does the system look kind of-
- 37:33
Ah, yeah. Um, so the solution here is evals, you'll be surprised to learn. Uh, a, a production system should have a set of regression evals, which are things that you are expecting, uh, your agent to already be good at. So your eval's already in place, which you can, uh, create them, uh, I'll show you in a second. Uh, with your evals already in place, you will know that your system doesn't change from the good behavior that it previously had as a result of this change. Uh,
- 38:04
if you don't have any evals, I can show you how that's done, uh, which is we can go to the Wonder Toys project, uh, and ask it to build an eval.
- 38:21
"Build an eval that ensures that I'm never giving instructions about how to build a bomb."
- 38:34
That seems like a really basic eval that I should have had in from the beginning. Uh, this thing that I'm using right now is called Alyx. It is our in-agent assistant. It is our, uh, in-product agent. Uh, and it knows, like the skills do, it knows how to build everything. So, uh, you can do this in product, you can do it in the skills, in your co-coding agent, but either way, it will create an evaluator for you, uh, and run automatically.
- 39:11
And that is pretty much it as far as we're going. Question at the back.
- 39:16
What, what harness is Signal?
- 39:19
Uh, what harness is Signal? Uh, I don't know. What harness is Signal? You know, Jason.
- 39:26
It supports Claude Code today, but it's in your own harness.
- 39:30
Yes. Uh, Signal is, uh, Claude Code at the moment, and we'll-- you can bring your own later. Uh, that is my boss trolling me, in case you're wondering why he happens to know the answer to my question.
- 39:43
Any other questions about Signal? Yes.
- 39:49
What are the total usage for the signals in our test? Can you then go into the product itself?
- 39:55
Sorry, louder.
- 39:56
So what's the total cost of total usage for the signal completion or next step or within the product itself that we don't need to worry about the cost of the channel?
- 40:05
Uh, at the moment, the, the question is about the cost of running Signal. At the moment, uh, we eat the cost of Signal. The, the, uh, solu-- the aim is, uh, for t- for you to get the best product experience right now. Uh, we can't guarantee that that's always going to be that way. Uh, but at the moment, uh, Signal is run on our agents at our dime.
- 40:29
Yep.
- 40:30
Uh, can you talk us through how Signals is evaluated?
- 40:33
That is an excellent question. How is Signal ev- how is Signal evaluated? Uh, Signal is ev- has its own eval suite, uh, that, uh, runs internally, uh, and is traced using Arize. So everything that Signal is doing is sending its own telemetry back to us, uh, unless you opt out of it, uh, which is telling us, uh, the queries that it's getting, what it's doing, how it's doing, all of that stuff. We have powerful LLM judges looking
- 41:03
at the results of Signal, uh, and telling you whether or not, uh, Signal did a good job generating that PR or not. Uh, so exactly the same software development life cycle that I was just talking about, uh, applied to Signal itself. Awesome question. Thank you. Gets us to show off. Yep.
- 41:22
Yeah. Can you explain how to manage your first experience with
- 41:30
the algorithm? Does Signal also automatically create its own evaluators as well?
- 41:33
Uh, does Signal automatically create its own evaluators? Not right now. Uh, at the moment, you have to s- you have to... Because evaluators, you know, take time and money to run, you wanna, uh, you wanna get the suggestion, and then you want to create the eval yourself.
- 41:51
All right. Uh, and I think, uh, I'm gonna leave it at that. Thank you all for your time and attention.