AI Engineer World's Fair 2026
Why Building an Eval Platform Is Harder Than It Looks — Braintrust
Read the talk
Why Building an Eval Platform Is Harder Than It Looks — Braintrust
Hossein Niazmandi follows evals from a spreadsheet to a production feedback loop, showing how collaboration, large traces and competing query workloads turn a simple testing interface into a systems problem.
From a talk by Hossein Niazmandi
At a glance
Ideas worth remembering
Offline evals establish expectations; production observability tests those expectations against real interactions.
A custom UI improves access, but experimentation also needs editable candidates, comparable results and data that supports analysis over time.
The improvement loop turns observed failures into offline tests, checks changes for regressions and returns the chosen version to production.
An eval platform must support both immediate investigation of arriving traces and long-running analysis of large, nested data.
Coding agents can run evals and propose changes, while humans review the resulting alternatives and choose what ships.
Testing before launch only establishes a hypothesis
An agent can respond differently to similar inputs, and that variability makes confidence harder to earn. Hossein Niazmandi, who leads solutions engineering for the West at Braintrust, frames agent quality around two connected activities: testing behavior before deployment and observing behavior once real users arrive.
- Evals before production. Run examples, experiment with the agent and test whether its behavior meets the intended standard. These results establish a hypothesis about how the deployed application will behave.
- Observability in production. Monitor actual interactions to see whether that hypothesis holds. Real users supply requests and circumstances that offline development may never have covered.
The same flexibility that lets an LLM handle different domains and user needs creates operational risk. Inconsistent behavior can damage a brand; an inappropriate statement or action can create compliance problems; difficult debugging increases maintenance costs. Evals reduce uncertainty before launch. Continuous monitoring supplies the evidence needed to keep reducing it afterward.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The spreadsheet captures the visible part of the work
A spreadsheet is a reasonable first step. It acknowledges that agent behavior needs inspection and gives the team somewhere to record results. The smallest eval system needs three things: examples that invoke the agent, a way to execute those examples, and a way to inspect outputs or scores. An example can contain a prompt, a request or whatever context starts the run.
That simple arrangement is the tip of the iceberg. A useful platform eventually needs better datasets, scoring systems, review processes and debugging tools. It also needs a way to connect pre-production tests with production behavior. Each addition changes what the system must retain and what people must be able to do with it.
The people involved change too. Engineers implement and debug the agent. Product managers who previously wrote product requirements now participate in evals. Subject-matter experts bring the domain knowledge needed to judge the product’s behavior. An interface that works for the engineer running a script may leave the people who understand the desired outcome outside the process.
This expands evaluation into an operating workflow. Testing once and shipping does not finish the job: the team keeps improving agent quality as the application runs. The platform has to support that continuing work, rather than merely preserve the output of one test run.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A nicer report still leaves experimentation unfinished
The target is an improvement loop: observe production failures, turn them into test cases, improve the agent offline and check that the changes do not introduce regressions. This loop has to continue because both user behavior and the agent itself change over time. The stages leading toward it reveal why a small internal eval tool tends to grow.
The first stage is a loop over input examples. Execute the agent, record its output, tweak something and run the examples again. Its advantage is speed: a team can begin with almost no barrier to entry. Its weakness is the work surrounding those runs. Comparing results over time, organizing human scoring and collaborating with domain experts remain manual and awkward. The spreadsheet records attempts, but offers little support for systematic experimentation.
The second stage adds a custom UI, often built by a product engineer. Product managers can now enter the eval process more easily, and iteration can become quicker. But presentation alone does not solve review, collaboration, data persistence or analysis across a long history of runs. Niazmandi’s verdict on this stage is concise: “Still just a reporting tool.”
The third stage lets nontechnical users change the system being evaluated. They can adjust the system prompt, switch the underlying model or vary parameter values, then compare the resulting behavior side by side. That is a meaningful capability change: users can investigate alternatives rather than only read logged results.
Yet those comparisons still use a curated offline dataset. A good result on that set tells the team what to expect on those examples; it does not reveal the steps the deployed agent takes or where live interactions fail. Without production visibility, the team can ship a carefully tested version and still operate blind.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Production failures become the next test cases
The fourth stage connects experiments to production traces. Log the input, output and each execution step so a poor interaction can be examined rather than represented only by a disappointing final answer. The team defines dimensions of success and failure, uses the traces to understand what happened, and selects failed scenarios for offline evaluation.
Those scenarios then become reusable tests. The team iterates on the agent, checks the proposed improvement alongside regression cases and ships the chosen version. Production supplies fresh behavior for the next pass. The practical gain is a connection between a live failure and the evidence used to decide whether a proposed change helps.
Where does a production failure go, and how does a fix return to users? The cycle below follows that movement. Traces supply diagnostic information; selected scenarios become tests; offline comparisons guide the next release. The return arrow matters because a release creates new evidence rather than ending evaluation.
At this point the team owns a product containing much of its development cycle. A proof of concept can make that look easy. Maintaining the same workflow at production scale exposes the storage and query problems hidden beneath the interface.
Log inputs, outputs and execution steps; identify poor interactions.
Production traces feed offline tests; tested changes return to production, where new interactions restart the loop.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Large, nested traces create two query workloads
“Agent traces are nasty.” They contain semi-structured JSON, deeply nested fields and full text. Niazmandi describes some interactions producing tens or hundreds of megabytes—“a movie’s worth of data”—compared with kilobyte-sized heartbeat observability records. The storage problem includes payload size, data shape and the ways people need to read it.
His claim that typical cloud data warehouses can collapse under this load comes from the workloads he describes, rather than a comparative benchmark in the recording. The concrete requirement is to query large, irregular traces while also keeping newly arriving production data useful immediately.
- Immediate investigation. Query traces as they arrive to understand why users are having poor experiences and respond quickly.
- Long-running analysis. Examine larger bodies of data to improve the agent experience or support humans annotating the outcomes of evaluation judges. These reads also include large aggregations and point-in-time snapshots.
Braintrust built BTQL as an abstraction layer through which people can query that data. The query interface is itself another component to design and maintain. The recording explains why the layer exists, without specifying its underlying database design: the platform must make the trace data accessible while supporting these different workloads. A polished results page is only one part of that system.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Agents run the loop; humans choose what ships
The platform’s users now include coding agents. A headless interface lets a person request traces from the last twenty-four hours where users had poor experiences, then ask the agent to run evals. With access to the codebase and underlying infrastructure, the coding agent can execute the tests and log their results. This shifts some of the labor away from engineers and product managers who previously performed each step themselves.
Trace analysis can also help discover failures the team did not think to label in advance. Braintrust’s proposed mechanism is to run inference over stored tracing data and surface patterns such as silent failures or users repeatedly prompting for the same request. Repetition is a useful signal of possible frustration; it still needs interpretation to establish what went wrong.
Follow that repeated-request scenario through the loop. Initially, the observable behavior is a user asking for the same thing again and again. Analysis of the stored interaction can flag that pattern for investigation. Once the team understands the failure, the request and relevant context can become an offline test case. Candidate changes can then be evaluated against that case and the existing regression set. The visible change is in the workflow: an interaction that might have remained an unnoticed production trace becomes a case the team can rerun and compare. This is an illustrative path through the mechanism, not a reported successful fix.
That automation still depends on the rest of the platform: a database, role-based access control to manage permissions, and data masking. These are additional responsibilities for the team building the system. Supporting more participants and more production data requires controls as well as analysis.
The ending assigns the final decision to a human. Coding agents can iteratively change the application, suggest improvements and run different eval iterations. A person reviews the outcomes and decides which version best serves what they want to ship. Faster iteration makes comparison more important: the agent produces alternatives, and the human chooses the production behavior.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
Related talks
- Beating RL With Reflection: GEPA and Optimize Anything
Explains a concrete optimization loop in which execution traces guide candidate changes, evaluators score the alternatives, and human feedback helps define the judge.
Read the complete timestamped transcript
- 0:12
Thank you for joining this session on why building agent quality platforms is hard. My name is Hossein. I lead the, uh, solution engineering organization on the west for Braintrust. Um, I spent fifteen years in solutions between Salesforce and Databricks. I've been a lifelong technologist. I've built computers since I was five years old, so I have this, like, insatiable appetite for technology, which finds me kind of on the bleeding edge of AI here at Braintrust.
- 0:42
For those of you, I know some of you were in the room earlier when Jess did her session, I saw some hands about evals. How many of you know what Braintrust is and what we do as a company? Just show of hands. Perfect. I see one or two. So at a glance, Braintrust as a platform is focused around agent quality, and the core idea is helping teams build and maintain confidence with the AI features or the AI agents that they're shipping.
- 1:12
And there are two key pillars of agent quality. The first is evals, which we talked about earlier. Evals is what you do before your agent reaches production. This is when teams experiment, they test the behavior, and they build the confidence that they want to have before they ship to production. The second pillar is observability. Observability is once your agent is in production and is interfacing with
- 1:42
real-life users, it's creating interactions, and the goal is to take that hypothesis that you built in your offline development process and then reinforce that in your production scenario through continuous monitoring. Evals and observability are closely related to the same problem. One happens in production, the other one happens in development, and both is the understanding to improve agent
- 2:12
quality. Now, I won't spend much more time here. Let me close that. Let's talk about why we're here, which is,
- 2:25
why are evals important? You know, evals matter because LLMs by nature are non-deterministic. They are highly variable. That variability is exactly what makes them so powerful. They can reason across many different domains, they can solve different kinds of problems, and they can handle a wide range of user needs. But that same flexibility also creates risk. Agents use
- 2:54
LLM as their brain, and agentic experiences are starting to become the primary way in which users will interact with companies. So teams need confidence in their agents that they will behave reliably and produce the outcomes that they expect. And without evals, companies face real risks. Brand, brands become a risk if there's inconsistent behaviors,
- 3:25
compliance if the agent says or does the wrong thing, and there's a cost and maintenance risk if the systems are hard or unreliable to debug or control. The goal of evals is to reduce that uncertainty before launch so that customers have a good experience and the agents behave the way that you expect them to.
- 3:51
So you might be thinking, okay, well, evals is just a UI on a spreadsheet, right? And I guess I should ask the audience, just show of hands, how many of you have run evals and logged the results into a spreadsheet? Has that happened? Okay, I see one or two. So there's no shame in that, by the way. It's actually the, a good first step. The important thing is that teams acknowledge that this is a real problem, and they need a way to understand how agents behave
- 4:21
based on different inputs. At the simplest level, an eval system tests against... Or it has three different kinda criteria. It has a way to execute agents against test inputs. Second, a way to view the outputs or the scores, even if those are being logged in a spreadsheet. And third, a set of test inputs or examples that can invoke the agent. An input example is
- 4:51
whatever information is needed to start the agent run. It could be a prompt, a request, or context that causes the agent to act. Now, spreadsheet-based evals are not wrong. However, they're often the first step for a useful version of eval workflow. But of course, as you can imagine, there's more to the iceberg than just what we just talked about. And
- 5:21
if evals were only, you know, run the agent, view the output, and then score it in a spreadsheet, this talk would be over right now. But there's a lot more that happens behind the scenes or underneath the iceberg. There are many supporting teams that eventually need to build better datasets, scoring systems, review processes, debugging tools, and the way to connect pre-production testing
- 5:51
with production behavior. We'll touch on some of these today, and anything I don't get to, you feel free to come up afterwards and we can chat about. And this is where things start to get complicated. The underlying technology is complex. LLMs are not simply deterministic. Agent quality is a multi-persona problem. It's not just software engineers and AI engineers. There are PMs who used to build PRDs that are now running evals. There are SMEs
- 6:21
who have the domain expertise of the product that you're building.
- 6:27
And evals themselves are just one part of the development and operating workflow.
- 6:35
And gone are the days where you build once with unit tests and regression tests, you ship to production, and you don't think about it. Evals are a way for you to continue hill climbing against one target, which is agent quality.
- 6:49
So let's talk about the different stages of the, the eval platforms that we see.
- 6:56
And before we do that, this is the North Star for many teams. They want to build the improvement loop, which is I have a feature or application that's in production, I can observe the failure modes that happen based on dimensions that I define. I wanna be able to grab those failure modes and create test cases so that I can iterate offline so I can improve my agent quality without introducing new regressions. And then you continue iterating
- 7:26
over time because just like in classical ML, drift becomes a real thing. The way users interact with your agents will change over time. The way you build your agents will change over time. So what does phase one look like? There were some hands that were raised around spreadsheets, and it is not-- there's nothing to be shameful about. This is where many people start. And a s-- the basic setup is pretty simple. It's a for loop with a set of
- 7:56
input examples and a way to execute the agent. And with that, you can run the same examples that you tweak and see how the agent responds with the outputs over time. The bis-biggest advantage is accessibility and time. You can start this with basically zero barrier to entry, but the re- the returns diminish pretty quickly. At this stage, you're pretty much only documenting. You can't really do true experimentation,
- 8:27
and let alone do any ki- type of analysis over time. Common limitations are analytics are hard, they're manual, human scoring is valuable but difficult to do, collaboration is very weak, and domain experts are usually boxed out of these type of results.
- 8:48
So again, spreadsheets are a great starting place, but difficult to evolve. Then someone, usually a product engineer, you know, they see this problem and they think, "Okay, I can build a nice UI, a bespoke UI to help serve-- to solve this problem." And while it might be nice at a glance, um, it helps you bring in the different personas that you would otherwise not have access to. This allows PMs to now
- 9:18
kinda get into the loop of doing evals.
- 9:24
But the limitation that you might have thought about, it continues, that the documentation, the review process, though it might l- visually look nicer, still collaboration becomes difficult. And though iteration cycles become quicker, it is still not easy for users to persist that data and do long-horizon analytics on. Still just a reporting tool. Then the next evolution becomes,
- 9:54
"Hey, I want my non-technical users to have an interface where not that they just log the results, but they should be able to experiment. I wanna be able to tweak a system prompt. I wanna change the underlying model. I wanna ch-- iterate on some parameter values." So being able to do this type of side-by-side comparison becomes the next evolution of the eval cycle that we see. But there's still one primary gap, and that is
- 10:24
that the test cases in which you use to do these kind of comparisons are still a curated set that arise in your offline evals. But what happens in production? I've done my iterations here. I have a good idea of what I can expect to happen in prod. But once I ship it to production, I'm operating in the blind, so I have no idea what steps my agent is taking and when those failures are actually happening in my live environment.
- 10:55
And when we think-- kinda going back to the flywheel, when we think about what are the best systems that enable this, well, you should be able to, starting at the top, observe the failure modes, log every input, output, every step of your trace execution that your agent takes, analyze, understand what went wrong. You create the dimensions of failure and success because you know the outcomes. Then grabbing those failure modes, those different scenarios where your users had poor interactions,
- 11:26
and creating evals in them in your offline process, iterating on them, so you can create the improvements without developing new regressions, and then shipping to production and hill climbing against this.
- 11:40
And when you think about the final outcome of this, well, you get teams that are building the flywheel. You get that entire development cycle inside of your platform, but then the problem becomes you have to h- you have to maintain it. You own this, this product that you've built. And especially, this is easy to do when it comes to smaller scale or POC and demo environments. But agent traces are nasty.
- 12:10
They are semi-structured JSON. In some cases, we see teams that are logging a movie's worth of data in each interaction, hundreds of megabytes. So being able to query that type of data can be difficult, and your typical cloud data warehouses kinda collapse under this type of volume and load.
- 12:33
And this is the same problem that we ran into. I won't spend much time talking about Braintrust, but when you have production traces coming in, you want to be able to query that data in real time as it lands in there. Because if your users are having poor experiences, you wanna be able to know in real time what is happening, why are they having these poor experiences, and how can I remedy that as soon as possible? You also wanna be able to do long-running
- 13:02
queries. If you think about being able to fine-tune your, your agent experience or even have humans who are doing a- alignment by annotating your, your judges' outcomes. Two different workloads. So we built an abstraction layer called BTQL, which is a s- another form of complexity because you want an interface for people to be able to query that data.
- 13:31
So it's not necessarily that it's just a UI or UX problem, but building an eval platform truly becomes a systems problem.
- 13:44
And there are a novel set of issues when the ChatGPT boom happened a few years ago, where I need to do real-time ingest. I have huge payloads that are h- tens of megabytes, hundreds of megabytes in size sometimes, compared to traditional heartbeat observability, which are just kilobytes in size. The structure and the shape of the data is different. Having deeply nested semi-structured full text is
- 14:14
difficult to query, and the read patterns are different as well. You wanna be able to aggregate in large volumes, but also be able to do snapshots of s- of, of in-time data.
- 14:28
And building the right system should allow you to not only empower the AI engineer, but also the PMs as well as the SMEs. And more recently, we've seen that agents have become a first-class citizen of these eval platforms, where you wanna be able to use natural language in a headless experience, where you can tell your coding agent of choice, "Find me all traces in the last twenty-four
- 14:58
hours where the user had a poor experience. Run the evals for me." And because these coding agents have access to all of your underlying infrastructure and your code base, they become the mechanism to run the evals and log those results. And then it becomes a question of like, so, so what? Why is this a-- why is this important? Well, the goal is everything that I've shown you so far, it requires the
- 15:28
humans, the, the engineer, the PM, to be deeply involved in this process, and it can be laborious at times.
- 15:39
At Braintrust, the way we think about this is we want to help you operate at scale. Rather than you having to think about these dimensions and-- of success and failure for your AI agent, Braintrust can surface these insights automatically. Because we log your tracing data into our platform, we can run inference on them, and we can tell you the unknown unknowns. "Hey, what are, what are scenarios that when my agent is s- failing
- 16:09
silently?" Or, "What are scenarios that my users are experiencing frustration by prompting my agent for the same request over and over again?"
- 16:20
And it's a lot of things that we didn't talk about, things like the underlying database that powers it or having to build RBAC into the system to manage controls and permissions or data masking. You know, each one of these features proves that this is not just a UI and a spreadsheet, but rather it is a, a systems problem that powers it.
- 16:43
And as we think about the evolution of, of the improvement loop, I said this earlier, where humans were doing the improvement loop, and now we see coding agents where they could iteratively make changes and suggest what kind of improvements you should be making to your application. And ultimately, the human's responsibility is to review the outcome. If I have different iterations of these evals that are run by my coding agent, I can
- 17:13
look at the outcome and make the decision of which version of it is the best for what I wanna ship to production. So with that, that is the end of my talk. I appreciate you all coming out and learning about why it's hard to build eval systems. Thank you.