AI Engineer World's Fair 2026
Advanced Workshop: Mastering AI Observability — Doug Guthrie, Braintrust
Read the talk
Advanced Workshop: Mastering AI Observability
Doug Guthrie connects tracing, online scoring and topic discovery to a practical improvement loop: find support-agent failures, turn them into evaluation cases, and give a coding agent enough context to propose reviewable changes.
From a talk by Doug Guthrie
At a glance
Ideas worth remembering
Traces supply the intermediate evidence; scorers and topic labels make relevant failures easier to find within a large stream.
Custom facets can reveal domain-specific workflow problems that default issue classification misses, as the support-agent record-lookup example shows.
Production failures belong in evaluation datasets so fixes can be checked against both new edge cases and previously preserved behavior.
A coding assistant needs repository access plus selected trace evidence to propose concrete changes. The CLI and improvement skill connect those two kinds of context.
Automation can package code changes, reasons, regression cases and evaluation evidence into a pull request; a reviewer still decides whether the proposal should ship.
Remote evaluations let collaborators change exposed prompts and models in a playground while agent execution stays on the evaluation server.
Capture the steps before trying to explain the failure
An agent can generate a great deal of telemetry without making its problems easy to find. A final answer tells you what the user received; it does not explain which tool ran, what input it received, or where the workflow went wrong. Doug Guthrie, a solutions engineer at Braintrust, builds this workshop around that gap between collecting traces and using them to improve an agent.
Tracing records the intermediate steps that lead to an output, including tool inputs and outputs. An untraced part of the system remains a black box for both an engineer and a coding assistant investigating a failure. More detailed tracing creates more data to examine, but it also supplies the evidence needed to distinguish a bad final response from a broken step upstream.
The workshop uses a deliberately simple Python support agent built with the OpenAI Agents SDK. That choice makes the integration concrete without making the framework the lesson. Guthrie describes other integrations, including OpenTelemetry and custom orchestration: the essential requirement is to capture the agent’s work, whatever produces it.
The intended development loop starts with a change, evaluates whether it improves or regresses behavior, and deploys it when the evidence supports doing so. Production then supplies the next set of observations. Guthrie’s compact instruction is that “production should inform development.” At scale, this requires another step between observation and editing: turn many traces into a smaller set of useful findings.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Scores measure the failure modes you already know
A scorer turns a question about behavior into a result that can be compared or filtered. The question might be mechanical—did the agent call the required tool?—or subjective—did this conversation resolve the user’s problem? Different questions need different scoring mechanisms.
- Code checks: Validate a schema or check whether tool A ran before tool B. These checks avoid an LLM call and are generally fast and cheap.
- LLM judges: Give a model explicit criteria for assessing an output or conversation. This allows judgments that are harder to express as deterministic rules.
- Human calibration: Review the judge’s results and reasoning to decide whether its judgments match the intended quality standard.
- Code with model calls: Put branching or other complex logic in a code-based scorer, invoking an LLM where the logic needs a subjective assessment.
A judge needs maintenance. Its explanation helps a reviewer understand why it assigned a score and how the criteria might need to change. Guthrie recommends starting with harsh binary judgments and difficult examples; a perfect result can mean the dataset or scorer is too forgiving. Examples in judge prompts and a stronger scoring model can help, though he presents these as practices that have worked with customers rather than rules for every workload.
Offline evaluation runs a candidate against test cases before release, potentially in CI, to compare improvement and regression. Online scoring applies judgments to incoming production traffic. A routing score of zero, for example, gives an investigator a way to isolate failed cases from the larger stream. The scoring criteria provide the signal; filtering makes that signal usable.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Topics looks beyond the existing test suite
An evaluation suite covers the behaviors its authors thought to test. Users can still encounter failures outside those cases. Braintrust’s Topics feature looks for patterns in trace data, groups them, and exposes their frequency. The resulting examples can suggest new evaluation cases or new scorers. Patterns in user intent can also reveal a capability users want that the agent does not yet provide.
A facet defines what to look for. The built-in facets cover task, issues and sentiment: what users are trying to accomplish, what goes wrong, and how users feel during the interaction. The model creates a taxonomy through the pipeline rather than requiring the user to supply every category in advance. A custom facet supplies a prompt for a more specific question.
How does a large execution trace become something an investigator can filter? The diagram follows the transformation. A preprocessor reduces the trace to the user, assistant and tool-call exchange. A facet summarizes that exchange for its particular concern. The backend embeds the summary, builds a topic map, and uses the result to label and classify traces. Those labels connect aggregate patterns back to individual examples.
Cost matters because a discovery process that sees only a tiny sample may miss the tail of behavior it is meant to uncover. Guthrie describes running Topics on 100% of traces as a design goal and quotes six cents per million input tokens and 40 cents per million output tokens for the hosted processing discussed here. Those are workshop-era figures, not a complete estimate of the cost of processing a particular application.
User messages, assistant messages and tool calls.
The pipeline summarizes traces through a chosen facet before embedding them. The resulting labels let queries select relevant traces without loading the entire stream into an agent’s context.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
What the pipeline can distinguish—and what can be customized
An audience question challenges the premise: if an agent repeatedly uses the same tools and workflow, won’t its traces be too similar for embeddings to distinguish useful patterns? Guthrie points to customer experience and the forthcoming demonstration, while acknowledging that the workshop data is scripted demo data. The practical question remains application-specific: whether the facets separate meaningful differences in the traces your users actually produce.
In response to a competitor question, Guthrie emphasizes integration flexibility and Brainstore, Braintrust’s purpose-built backend. He describes moving away from ClickHouse after encountering scaling problems with customer workloads. That is his account of the product’s architectural motivation; the workshop does not provide a comparative database benchmark. The relevant capability for the later demonstration is direct SQL access to traces and spans.
Customization starts before clustering. A custom preprocessor can include metadata or spans the default message-oriented preprocessor does not capture, and a custom facet can ask a domain-specific question. Long conversations also require settings that reflect how turns are distributed across traces and how long an interaction lasts. These choices determine what evidence reaches the pipeline and when it is processed.
The same approach can examine internal coding-agent sessions. Guthrie names integrations for Claude Code, Codex, Cursor and OpenCode, with Topics helping teams understand what engineers do in those sessions. Understanding usage is a starting point; he treats measuring the tools’ broader effect on engineering efficiency as an area that still needs additional machinery.
Deployment can separate the control plane from the data plane. The web app and control-plane metadata stay with Braintrust, while traces, datasets and experiments can reside in the customer’s infrastructure. This distinction matters for teams evaluating where their trace data will be stored; it does not describe a wholly self-hosted web application.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Wire the support agent and establish a baseline
Venue internet problems turn the hands-on portion into a walkthrough. The setup still makes the dependencies clear: create an organization, create the AIE-Workshop project expected by the repository, configure AI providers, and generate a Braintrust API key. The project name in the application and repository must agree so logs and evaluation assets land in the intended place.
Provider configuration supplies the models used by LLM judges. Guthrie adds OpenAI and Anthropic credentials, and describes custom providers for organizations that route model calls through their own gateway. A firewalled gateway may need network access configured as well. Ordinary judge calls use the organization’s configured models; Topics uses Braintrust’s built-in models.
Locally, uv sync installs the Python dependencies and creates a virtual environment. The example .env supplies a Braintrust API key and a default model available through the configured providers, with an option to choose a different judge model. The Braintrust CLI adds the operations needed later: query SQL, inspect logs, create datasets and run evaluations. Skills teach the coding assistant how to use those operations, while BT init creates a config.json identifying the organization and project.
After seeding the database and checking configuration with make ready, a chat interaction produces a trace in Braintrust. The integration sets the OpenAI Agents SDK trace processor to a Braintrust tracing processor, directing emitted traces and spans to Braintrust. Guthrie uses explicit setup because the example also accesses configured providers through Braintrust’s gateway; he notes a simpler auto-instrumentation option.
The initial evaluation loads a small dataset, invokes the support agent for each row, and stores the telemetry and results as an experiment. Its scores mix a code-based required-tools check with LLM judgments such as support resolution and tool-use quality. This establishes something to compare against after a change: the next experiment can show whether a fix improves the new case while damaging previously tested behavior.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Choose the scoring scope and the evidence each facet sees
Publishing the scorers makes the repository’s evaluation logic available for online use. The workshop’s push-scores command uploads them to Braintrust, where automations can invoke them on incoming logs. Keeping scorer definitions in code also allows a playground user to reuse a scorer whose source remains version-controlled.
The judge’s input should match the question being asked. The thread variable gives it the back-and-forth conversation, useful for multi-turn resolution. User messages, assistant messages, or a particular span’s input and output support narrower checks. Judging a policy lookup’s intermediate result requires different evidence from judging whether an entire support conversation succeeded.
- Span scope: Assess an intermediate step, using filters to select the relevant span. Selecting span scope defaults to the root.
- Trace scope: Access the spans within one trace. The demonstration uses this scope for communication quality, support resolution and tool-use quality.
- Group scope: Combine related traces when conversation turns are recorded separately. A shared conversation ID provides the relationship.
Guthrie sets judge sampling to 100% for the demonstration so the incoming examples receive broad scoring coverage. That setting incurs model compute for the selected judges. Sampling is therefore a cost-and-coverage decision separate from choosing the scorer itself.
Topics receives the built-in facets plus a custom support workflow issue facet from a Markdown prompt in the repository. The custom prompt outputs none when its no-problem criteria are met, and a regex excludes those outputs. This keeps the issue map focused on failures instead of filling it with successful interactions.
Processing settings control which traces qualify, what proportion is sampled, and which time window supplies the topic map. Timing and group scope also address conversations spread across many traces or a long period: the pipeline can use the shared conversation ID and an invocation interval suited to the interaction. These settings can be configured through the CLI as well as the UI.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A refund inquiry becomes a visible record-lookup failure
The import supplies enough data to examine patterns: BT sync loads 968 sample traces containing 5,260 spans from an S3 bucket. Each conversation turn is represented by its own trace. In one example, the current turn carries eleven previous messages; the OpenAI Agents SDK has attached that history, so it is already available without separately retrieving every earlier trace.
Consider the user checking the status of a returned-item refund. The incoming trace initially records the interaction. The configured automations then attach additional interpretations: negative sentiment reflects frustration with the pending transaction, while the custom workflow facet identifies an order-record retrieval failure. Both the specific order-ID lookup and the broader customer search failed to locate the order. A vague unsuccessful support exchange has become a case with a named workflow problem.
The trace views answer different investigative questions. Thread view shows the conversational exchange; timeline view helps locate bottlenecks and inspect information such as cache hits across model calls. Loop, the embedded assistant, can also create a review UI over trace data so another participant can examine agent outputs or calibrate judges.
In a project where processing has already finished, the topic map shows grouped patterns and their frequency. The built-in issues map has not appeared: Guthrie explains that map generation requires at least 100 matches or generated facets, and the default issues prompt does not fit this agent well enough to produce them. The custom support-workflow facet does reveal record-lookup failures. An absent default issue map therefore cannot be read as an absence of problems.
The useful change here is in what can be investigated, rather than a demonstrated repair to the refund workflow. Task, sentiment, workflow labels and online scores now give an engineer—or an assistant—ways to select examples and inspect the tool failures behind them. The custom facet makes a domain-specific failure visible that the initial evaluation and default issue classification had missed.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Query the patterns, inspect the spans, preserve the case
Manual investigation combines labels and scores to narrow the logs, then opens traces and spans to discover what happened. That workflow is useful but reactive: someone must keep choosing filters and reading examples. Guthrie asks Loop, “Help me understand what I should improve in my support agent based on recent logs.” Loop responds by issuing SQL queries against Brainstore to retrieve relevant context.
The labels do important work before the assistant reads the detailed evidence. Pulling every trace into context would repeat the original scaling problem. Queries can instead find relevant groups and representative failures, after which the assistant examines span inputs and outputs. Classification selects where to look; the execution trace explains what went wrong.
A selected failure can then become a dataset row for offline evaluation. The refund lookup problem illustrates the sequence: discover the workflow failure, inspect its trace, preserve a useful case, change the agent, and rerun the evaluation. A new lookup-failure scorer may also be needed so the same kind of problem can be detected before the next release.
An audience clarification makes the purpose precise: this is an evaluation dataset, not a training set. Adding a trace does not itself update model weights. It expands the cases used to test the agent. The evaluation combines a dataset, a task that invokes the agent, and scorers that judge the results; preserving older cases helps detect regressions while new cases extend coverage.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Give the coding agent both production evidence and repository context
The agent-improvement skill packages the investigation loop for a coding assistant: write SQL, view traces and spans, create scores, modify the agent, and run evaluations. After installing it, Guthrie starts Codex in the repository. The CLI connects the assistant’s local code access to the scored and classified traces in the Braintrust project.
Codex queries the 968 support traces and identifies scored failures, including a visible cluster of errors in the escalate-to-human tool. It then examines classifier outputs and representative failed traces, continuing to query and drill into individual examples. The demonstrated investigation uses aggregate signals to find candidates, then detailed execution evidence to understand them.
This division of work matters. Topics and scorers provide searchable observations; SQL retrieves a manageable subset; the coding assistant supplies access to the repository where a fix can be made. The workflow joins observability and evaluation through tools the assistant can invoke, rather than requiring it to infer a code change from a dashboard alone.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Sampling, model choice and access shape the operating loop
While Codex runs, the questions return to deployment. Hybrid deployment is an enterprise option in the workshop: traces, experiments and datasets remain in the customer’s infrastructure, while the control plane stays with Braintrust. In BYOC, Braintrust manages the customer-hosted data plane; in the other hybrid arrangement, the customer’s team manages it using deployment tooling such as Terraform or Helm.
Model flexibility has a specific limit. Evaluation can help test a swap between generation models without introducing regressions, and judges can use configured provider models. Topics summarization and vector embedding, however, are restricted to Braintrust-hosted models. Guthrie says the team previously allowed substitutions but got poor topic results; the hosted models were trained for this pipeline. Custom preprocessors and facet prompts do not imply that every model inside the pipeline is replaceable.
Classifications also support trends over time because they are stored as trace metadata. A saved view can filter for classifications.sentiment being non-null, group by its label, and chart sentiment or tasks over time. That moves the question from which patterns exist in a snapshot to how their frequency changes.
- Sampling as the agent matures: Guthrie observes higher judge sampling for newly deployed agents, with lower rates as known failures become less frequent and the marginal signal declines.
- Classification cost: Topics may replace some taxonomy-oriented judge work. He describes Dropbox experimenting with this because broader coverage was more affordable than its existing classification scores.
- Conditional scoring: A code-based scorer can call one scorer, branch on the result, call another, and return an array of score objects. The UI does not directly wire that branching graph; a filter can also gate invocation on the presence of a scorer span.
Projects group the logs, datasets and scorers relevant to an agent or use case. Monitoring is project-scoped in this demonstration; an organization-wide view requires queries that pull data across projects. Guthrie describes a first-class agent object within projects as forthcoming, rather than an existing capability shown here.
Evaluation assets should change when the agent’s behavior or responsibilities change, without requiring constant churn for a simple, stable agent. A prompt edit may need only a few additional examples; a new tool or runtime change may require a larger adjustment. Production observations give the team a reason to expand coverage rather than treating the original baseline as a complete model of user behavior.
Scoring can run in Braintrust or in the application’s own infrastructure. Published code-based scorers can bundle their dependencies, allowing Braintrust to handle asynchronous triggering. Teams that retain that machinery themselves can update existing spans in place by ID, attaching scores, metadata or metrics after the original trace arrives.
Trace access and improvement access need not be identical. Guthrie proposes restricting the raw-data project with role-based access control and allowing a service account to run the investigation. Engineers could then receive aggregate analysis, SQL queries and a proposed pull request without direct access to the underlying traces. The proposal separates who can inspect raw conversations from who can review a resulting code change.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The automated loop produces a pull request for a human decision
Returning to the coding session, Guthrie shows scorer changes, additions to the agent’s tools, and new examples in a dataset. The session also queries an evaluation experiment. The recording establishes that the workflow reaches code and evaluation assets, but does not present a numerical before-and-after result for the support-agent repairs.
How can this run without someone starting the terminal session each time? Guthrie shows a separate supervisor-agent example whose GitHub Action runs Claude Code with the Braintrust CLI and improvement skill. A schedule or event can trigger the workflow. The platform itself does not yet have a single button that performs the whole loop; the demonstrated automation assembles the available pieces.
The pull request packages what changed, why the assistant recommended it, evaluation impact, links to experiments, and regression examples drawn from the investigation. The reviewer then decides whether the proposal is meaningful. Guthrie reports that one customer generates five to 10 additional PRs per day with this approach; that is an output-volume observation, not a measure of merged changes or quality improvement.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Expose evaluation controls while keeping execution on the server
The final demonstration opens the improvement loop to collaborators who work on prompts and models rather than repository code. Remote evaluation exposes an evaluation task and its parameters to a playground. In the supervisor-agent example, those parameters include the system prompt and model, plus prompt and model settings for the math and research agents. A product manager or subject-matter expert can adjust the exposed controls without needing a custom application for every experiment.
Where does the agent run when a playground user clicks Run? The diagram separates configuration from execution. An evaluation dev server can run locally, with the demonstrated setup looking for localhost:8300, or be exposed in the team’s infrastructure. Braintrust sends a POST request to the evaluation endpoint; the agent’s code executes on that server, and evaluation results stream back to the playground.
Guthrie’s hosted example runs on Modal and uses a create-app function to expose the evaluations in an evals directory. The relationship is useful: the playground presents the parameters engineers choose to expose, while the server retains the agent’s tools and execution logic. Collaborators can explore prompt and model changes against the actual evaluation code.
The ending turns the same methods back on coding assistants themselves. Braintrust is working on ways to evaluate the skills those assistants use and understand their Claude Code and Codex sessions. The next object to inspect is the development workflow: which skills help, how sessions unfold, and whether the coding assistant is making engineers more effective.
Edit exposed prompts and models, then click Run.
The user edits exposed parameters in the playground. A POST invokes the evaluation server, where code runs; results stream back to the same playground.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
Related talks
- Beating RL With Reflection: GEPA and Optimize Anything
Explains a complementary improvement mechanism: use execution traces and diagnostic feedback to search over prompts, agent programs, skills and evaluators.
Read the complete timestamped transcript
- 0:13
Hey, everybody. My name is Doug Guthrie. I'm a, uh, solutions engineer at Braintrust. Um, I have a few colleagues, a couple colleagues in the room that, uh, can help if you get stuck anywhere, uh, within the workshop today. Also, um, I get interrupted all day long during calls, so if you wanna, uh, stop me at any point and ask a question, I'm up for that as well. Uh, hopefully, you've had a chance to grab the QR code by now. It just takes you out to this repo. This is what we're gonna be use, using today, uh, as
- 0:43
part of the workshop and going through understanding how you can do observability within Braintrust and sort of using all of the different features within there to build better agents. That's really why we're all here. Like, how do we actually create good quality agents? Um, I'll leave this up for, uh, ten more seconds or so.
- 1:21
I'm gonna have this again later in the, the presentation when we actually get into the workshop, so if you, if you don't get it, um, you can also Google, um, DP Guthrie in GitHub and search that re- uh, repository on, on GitHub, and you'll be able to find it.
- 1:36
Cool. So, um, we're all here to master AI observability, uh, and I'm gonna give you what that looks like in the context of using Braintrust. Uh, quick agenda. Uh, I do have some slides that I wanna sorta set the foundation for doing observability, and I think, um, what we've learned of how to, how to do it really well. So I'll sorta set the foundation for what that looks like, uh, give you a sense for how do we generate signal in this massive amount of noise. Um, you all have probably seen a trace or you can understand that the agents that
- 2:06
you're building are generating a massive amount of data, and being able to parse through all of that and actually derive some type of signal or insight is challenging. That's where the challenge resides. It's-- There's one side of actually being able to trace and then have a sorta platform that allows for observability. Then there's the other side of how do we actually do something with it and derive some type of insight. And so I'll show you what that looks like within Braintrust, and then we'll sorta walk through all of what, what I, what I described or what I've shown on, uh, on these slides in the workshop. Um, as
- 2:36
I said, uh, if you have a question, feel free to jump in. I do not mind. Uh, I have sort of, or at least I'll try, to save some time at the end for some Q&A as well. So feel free to save your question to the end.
- 2:50
So just a quick Braintrust intro if you're not, uh, familiar. Uh, we are an AI and obs-- We are an evals and observability, uh, platform. Our sole goal is to help our customers build better agents. Uh, we, we sort of, uh, help enable our customers understand what quality looks like and then how to improve it in a real measurable way. Um, one of, like, the, the more fun parts about my job is that I actually get to go speak to some of these customers that you see up there on the top right. I get to understand how they are
- 3:20
building their agents, and a lot of those learnings actually make, make their way into our product, make their way into my other engagements with customers. So I get to go actually talk to the engineers of Cloudflare, and I get to talk to the engineers of Dropbox, and I get to understand what they're doing, the, the types of scores they're creating, how they think about our new feature topics, what that sort of creates, the, the types of automations they're building on top of it. Uh, I hope to give you some sense of that today. Uh, I mentioned some of this, right? We are an evals and
- 3:49
observability platform. Um, you think about, like, evals and observability, I think they're, like, two diff-- two sides of the same coin. Uh, taken together, it's really what's going to allow you to create better quality agents. Uh, if you, if you hang around, and I mean, like, if you, if you aren't a Braintrust customer, if you're not using Braintrust today, would highly recommend coming back and looking at Braintrust in the next month, two months. Uh, where I think we see the, the industry at large today is, is a more, like, very
- 4:19
passive way of observing data. Uh, sort of what I alluded to of, like, being able to derive insight. Uh, Braintrust is moving, I think, a little bit more towards what we call active observability. It's, it's actually giving you the primitives, uh, the foundations to give you the right insights where you don't actually have to pull from your platform and understand what that means in the context of your agent and, and what you need to go do to go and improve it. Uh, so, uh, active observability is the, is the way in which we can start to bring those insights to you so that you can build better quality
- 4:49
agents. So quick, quick Braintrust stuff there. Um, I'll, I'll plug this again at the end, but we are-- we do have a booth at the expo. Uh, I think it opens around five. Come and find-- I'll be there tonight. Come and find me, ask some more questions. Uh, we'll be here all week as well. So the, the foundation to all of this, uh, the foundation to observability in general, like, I-I'm guessing most of you already know here, is tracing. You have to be able to understand all of the different steps that your agent takes, right? What are the, the, the things that
- 5:19
it's doing to get to its final output? And not o- not only that, but what are the inputs to those tools? What are the outputs of those? And if you s- if you start to think about, like, all of the different steps that your agent can take, this can balloon, um, pretty quickly. And being able to sort of parse through this, this is the challenge that we are trying to solve. But without that tracing, right, we're not actually able to understand what's going on, where quality is actually deteriorating. Uh, and so being able to, uh, one, trace all of this data and then derive
- 5:49
some type of insight, again, is where, where Braintrust is gonna provide some value. But i-i-it just in, in general, if you're thinking about observability, uh, if you don't have some part of your stack I-i in, in a trace that looks like this, then it's, it's a black box. It's gonna be very challenging for you to actually surface those quality issues or, you know, point, uh, Codex or Claude Code or whatever you're using out there to, like, understand what's going on. If that's, the, that data is not there, you can't actually surface those, uh, those places where the agent is falling down, where it
- 6:19
could be improved. So tracing is the foundation. Um, I think if you look at the eval side, and of course talking today the observability, it all sort of, uh, centers around us being able to, uh, capture all of the different things that our agent is doing.
- 6:35
Uh, so just another quick call-out on the Braintrust side. Uh, we, we have a lot of different ways-- So, uh, I'll skip ahead just a little bit. The workshop that we're doing is built on-- It's a, you know, dummy sort of support agent. The, the sort of agent isn't the thing that's important, uh, but it's built with the OpenAI Agents SDK. Uh, we have lots of din- different integrations with all of the different agent frameworks. We have integrations with all of the AI providers. We have different programming languages, so depending on how you are building your agent, whether it's the, the
- 7:05
example today is in Python, Python, but it could be TypeScript, it could be Java, Ruby, Go. Uh, so we are pretty flexible in that we allow you to sort of, um, bring whatever it is you're doing today from a programming language perspective or how you are building your agent, whether it's with a framework, whether you're using OpenTelemetry, whether you're sort of hand-rolling your own orchestration, whatever it is, uh, Braintrust has a way to hook into that pretty seamlessly. And I think if you're looking out across the market, right, if you're trying to understand how do I, uh, how do I trace this, this
- 7:35
agent, um, you're gonna have to have the ability to go across a lot of different things generally. And, uh, you can see over here, there's a lot of different ways in which you can integrate your stack into Braintrust. Uh, so again, I just have the dummy agent today is using the OpenAI, OpenAI Agents SDK, but that's certainly not what we're confined to.
- 7:57
Uh, this is something that I sort of-- I, I say the word flywheel like, I don't know, 100 times a day it, it, it seems like. Um, the, the whole idea and where I think we've gotten a lot of feedback from our customers is the, the way in which that, the way in which they are building better agents today is they are creating something like this, right? And what, what I mean by that is that production should inform development. Uh, you should have a really simple and easy-to-create flywheel that allows you to,
- 8:27
um, develop some type of change, fix something within your agent, test whether or not that change regressed the agent in some way or improved it. Uh, you should be able to have the confidence to go and deploy that agent to production. When you deploy, right, we need to be able to observe it. Uh, that observation though, we need to, like, tack something on top of that, which maybe you don't see necessarily called out specifically within this flywheel, but you need to be able to derive some type of insight at scale. Um, most of the customers, especially on the slide that I showed earlier,
- 8:57
they have a massive amount of scale, and being able to derive insight or having a human comb through that data would just be incredibly challenging, and so we need to layer intelligence on top of that, and that's one of the things that, that Braintrust will offer. But I think just if you're thinking about the, the process of building a better agent in general, you need a flywheel like this. You need to create something where production in some way informs how you go and change your agent. And so this is what I'll show you a little bit today when we, uh, go through the workshop, is how we can
- 9:27
actually surface some of those insights in an automated way from our, uh, support agent, bring those insights back to our development process, make some type of change, run an eval, right? That, that sort of loop, uh, becomes sort of, uh, it, it just incredibly important and, and hopefully it becomes a very trivial thing that most companies and most developers can just very easily plug into.
- 9:54
So when I say signal, uh, I just, in this context, I mean two different things that we will sort of wire up also during the workshop. Uh, I mean, uh, scores. So this could be, uh, you know, evaluators. This is just a way to measure the output of your agent in some way. Uh, you have some particular failure mode. You have something that you're trying to understand about the agent when it, um, when it received some input. Uh, we, we can create a lot of different types of scores.
- 10:24
Uh, one of them you see over on the far left is a code-based score. These are generally, uh, very fast, very cheap, uh, and very easy to run. Uh, you can imagine, like I'm trying to validate a schema. I am trying to understand was tool A called before tool B? Was it called in the right order? Uh, things that can be produced very, very quickly, uh, and without sort of having to go i-into an LLM. Uh, the more subjective though is that LLM judge. We wanna bring some sort of like human, uh, understanding to that, that, that trace,
- 10:54
right? To that interaction that just happened with that eval case or even, uh, online as well. And so you could actually, you know, give the LLM some criteria of how you want to assess the output or assess that conversation, and it should produce some score. Um, and then there's, there's probably actually a fourth one here, but the third one that you see over on the right is an LLM judg- L- LLM judge aligned with human review. What's also important, you could probably imagine, is if I have an LLM judge and it's producing results and we are sort of
- 11:24
making, uh, decisions based on th-those results, we should probably understand whether or not those results are good or not. And so there should be some element of human review. I think especially if you are using these judges in an online sense, right? You are scoring actual interactions of users with your agent, you should be able to provide some type of interface to users to provide that review to calibrate those, those judges, right? This isn't a set it and forget it type thing. Uh, your agent will evolve, so too will your scores. The, the fourth one that, uh,
- 11:54
I actually have an example of in the, uh, the worksh- the workshop is a sort of code-based score, but it invokes an LLM under, under the hood. Uh, you can imagine like there is branching logic or there's just, uh, more advanced or complex things that you need to do, uh, based on some type of logic that you want to encode in your code-based score. Uh, so you could do that as well. Uh, these are some of the things that I think I've heard in, in different conversations with, with our customers. Uh, I think, again, it's like take it with a grain of salt to some degree. Like, it's
- 12:24
not always gonna fit 100% of the scenarios every single time. It's, it's where I've seen a lot of success across our customer base. Um, just being able to, like, understand the reasoning that the LLM gave for producing a score, incredibly helpful, especially on that human review side, where it's like, hey, I can understand why it produced that. That will then inform how I can go change my scores in some way. Um, binary scoring, I think, uh, especially at the outset when we have sort of an agent that we are putting out into production, binary scoring I think is, is better
- 12:54
than, you know, A, B, C, D or, or so on. And you should be incredibly harsh. Uh, you should produce really hard examples. You should, you should-- If you get 100% on every single one of your evals, and, and we probably will today because it's, uh, sort of dummy use case, but if you get 100% of your, uh, 100% on your evals, uh, one, you probably don't have, uh, strong enough test cases in your dataset. Your scores probably aren't harsh enough, uh, so you should sorta look at that and, and s- and sorta calibrate
- 13:24
accordingly.
- 13:27
Um, another one that I've seen, you know, again, grain of salt, uh, in- including some examples in, in your prompts tends to, tends to help, and then perhaps using a stronger model for scoring than for generation, right? Let's bring a little bit, uh, closer to human subjectivity as possible to this. Uh, obviously, this can change based on the, the, the type of traffic, right? If we start to talk about online scoring and things like that. But, um, to that point, uh, there are obviously two different types of scores, at least the way that we see that, uh, here at
- 13:57
Braintrust. There is the offline, right? This is our, our evals. We are sort of, uh, trying to derive some signal of whether or not we should push this thing to production. This is going to give us indication of did it improve or did it regress? Um, this is gonna be, uh, you can imagine h- wiring this up to Excuse me. It's my daughter.
- 14:25
Um, yeah, so you could imagine, like, wiring this up, right, to CI. You could imagine wiring it up to some sort of automated pipeline, uh, again, where you're running it offline. And then on the flip side, the, the whole idea, like, the thing that we're talking about here is how do we, how do we derive signal from things that's actually happening in production? And one of those things that you can do is use those online scores and, and score that traffic as it's incoming into Braintrust. And again, the idea is to attach some type of signal where then you as
- 14:55
a user can actually go in and be like, "Hey, my routing accuracy score is zero in these examples. Is there some sort of pattern that I should be aware of based on this feedback?" Right? So I need to go from, like, this massive amount of data into something that is much smaller and easier for me to consume. The second signal is a, is a feature that we have in Braintrust called Topics. Um, the idea here is that you wanna, like, understand patterns within your data. You wanna go a little bit further beyond the, uh, the known
- 15:25
unknowns, right? So, like, your scores that you've already created, uh, in your offline evals, they are failure modes that perhaps you've already recognized. They are things that you already are testing for, but that's certainly not the, the scope in which your agent can operate or the failure modes that it can encounter. And so we need to understand beyond that, and Topics is one of those things that will help enable that. We should be able to, um, understand silent failures. So you can see there, like, these very sort of traditional scores like factuality and, and moderation. They're, they're gonna
- 15:55
catch a very small subset of errors and, and probably you have a little bit more, uh, bespoke scores that, that you're using. And then to that point, like, the evals that you're running, again, are using those, those subset of scores to understand improvement or regression. But there is a sort of wide swath of use cases that your test scenarios are not capturing, that your users are actually, uh, probing your agent with and trying to, like, get some type of output, and perhaps maybe it's not the right one. Um, so we wanna
- 16:25
capture these edge cases. We wanna capture the tail end, and we wanna structure, uh, then the flywheel, right? We wanna bring these examples back to our datasets. We wanna run evals. We wanna perhaps create new scores based on that feedback. So, so what is, what is Topics? Uh, Topics, uh, in general or at a high level is a way for us to understand patterns within your data. Um, we can create clusters of those patterns and give you a sense of frequency with which those things are occurring. Uh, also important to call out here is that there are both
- 16:55
built-in and custom facets. And, and what, what I mean by that is when you go into your Braintrust account, you can enable a feature called Topics, and out of the box, you can see task, issues, and sentiment. So I can understand, hey, what are, what is the common user intent? How are users feeling when they interact with my agent? And what are the common agent issues? So I can get that, that sort of taxonomy where I'm not necessarily providing that taxonomy up, uh, up front. The LLM, uh, creates that, uh, for me as part of this pipeline. But I can also bring my
- 17:25
own custom facets. I can, I can sort of build my own prompt, and I can build my own pipeline to create these patterns within my data, right? With respect to the things that I care about. Um, go, go back to this. Um, this is, you know, just an example of, of, of an illustration that you might see with the topics that were generated. Uh, but why, why would you use it? I think I sort of alluded to this a little bit, but you wanna be able to not have blind spots. Uh, you want to be able to understand all of the different ways in which your users
- 17:55
are interacting with your agents and then allow that to inform what you go do as a developer, right? This, this doesn't even necessarily have to be an issue or a failure mode. This could be, like, potentially, um, feature requests, new things that your agent could do based on the way users are actually interacting with your agent. Um, so this is, like, I think one of the ways in which I, I've seen a lot of our users start to use this, uh, is actually, uh, create this or, or sort of understand what's going on and inform what a roadmap should look like for the agent, right? It, it starts to uncover
- 18:25
things that we're not necessarily aware of with respect to these agents running in production.
- 18:31
Behind the scenes is, is a really complex pipeline. Um, you see over there on the left a preprocessor. So if you think back to my, my first slide with the, that trace, and that's, like, a very, very small trace, uh, relative to what I've seen our customers generate. What we need to do is pre-process that data into something that is meaningful to give to the LLM. Um, so we give you a, what we call a default preprocessor. We essentially take what w- what you'll see in our thread view, and you'll see that in just a little bit,
- 19:01
but it's the back and forth, right? It's the user, it's the assistant, it's the tool calls. We pre-process that massive trace into something that is meaningful, is, is much smaller. Uh, we then pass that to a facet, right? The facet is the thing that I, that I just described, issues, task, sentiment, your own built-in facets. What we're gonna do here, that, that, uh, pro- preprocessed data is going to be passed to that facet. We are going to then summarize it. That summary, we then create a vector embedding of that summary within our back end. Uh, on top of
- 19:31
that, once we have that vector embedding, we can now cr- go create what we call a topic map. This is sort of what you saw on the previous slide of the, the scatterplot. Uh, this is what now allows us to label our traces and classify them. And if I have those labels and those classifiers, it becomes incredibly trivial for me as a user, and also incredibly trivial for me as a user, um, using a coding agent on top of Braintrust to now start to filter through that noise using all of that really rich data. Uh,
- 20:02
important call-out, the pipeline that you're seeing, the facets and that embedding step is actually using a, a hosted model that we have. So we have just done a lot of work behind the scenes to create, um, not, not like a feature that you would just, like, show in a demo. Obviously, like, we'll, we'll see some stuff here today. But this should work at scale. You should be able to run this on 100% of your traces. That's, that's the sort of goal of our topics pipeline. It is, uh, just to get, like, m- maybe more detailed
- 20:32
than I should, six cents per, um, million tokens on the input side and 40 cents on the output side. Uh, compare that to, like, other models that are, that you might be using, uh, on the LLM Judge to create that same type of signal. You can create a pretty compelling use case, um, especially if you start to see value in the, the sort of topics and the labels and classifiers that we generate as a result of it.
- 20:57
Um, I also alluded to this a little bit, like, why'd we build this? Uh, we've met-- We have customers that, that generate, uh, massive amounts of data, and one of the things that I think we solved for very early on was being able to ingest that data and then being able to-- a- allowing our customers to query over that data in, in sort of sub-second speeds. What we need, what we didn't necessarily have at the outset was the ability to go do something with it, right? And this is where that active observability comes into play. Like, how do we now go derive
- 21:27
insight from all of this data, especially this data at scale? This is where the, the challenge comes into play. And so if you have this really complex pipeline that has been architected to both run at scale, um, and I mean that from a technical perspective as well as a cost one, uh, like I alluded to earlier, then we can actually provide a, a, I think a really interesting primitive and foundation that our customers can build on to inform that flywheel. Um, at the end of the day, the, the, the thing that we're trying to do is
- 21:57
enable our customers to build that flywheel and to build it in a way that actually, um, improves their agent quality, and they can do it very, very quickly.
- 22:10
All right. Hopefully, I, I didn't lose anybody there in the, in the presentation, and you all are ready for workshop. Um, quick show of hands. Anybody still need the QR code or the URL for the, the GitHub repo? Yeah.
- 22:26
Yeah, I'll keep it up here for 30 seconds or so.
- 22:54
Uh, my team uses LangSmith. I think LangSmith kind of sucks. Can you give me a list of things when my team says, "Well, why should we use Braintrust instead?" that I can go to them and say, "These are the things Braintrust can do that LangSmith cannot."
- 23:08
Oh, man. You're going, like, straight competitor here, huh?
- 23:11
Yeah.
- 23:11
Wow. Was not prepared for, for that. Uh, I am, I'm certainly not the guy, like, I will not be up here bashing anybody. Um, I, I do think that we, we have a, a stronger offering, and I think we have a stronger offering in a, in a lot of different regards. Um, certainly not in any particular order. One of them is the, the flexibility that I showed to you on the other slide. There is the ability to go across m- multiple languages, multiple programming, excuse me, multiple programming languages,
- 23:41
frameworks. Um, it's a very agnostic way to build your agent. Uh, LangSmith, I think, has some of that as well, but they are definitely geared towards LangGraph and, uh, LangChain and deep agents and th- and those types of things, and I think they have features that are only sort of for that. Um, the other thing that I think if you go back-- Actually, I used to have a slide on this, but Braintrust built their own database in late 2024, and, and the reason that we built this database is, is Ankur, our CEO,
- 24:11
recognized very early on the trajectory that we were going in from a agent perspective. And so he-- W- we, we were using ClickHouse behind the scenes at the time and recognized very quickly when we onboarded the likes of, you know, the, the, the customers that you saw in there, Netflix and Microsoft and Dropbox, that they-- ClickHouse just very quickly fell down at scale. And we had to go do something different. And so this is where the engineering team at Ankor built Brainstore. This is our own purpose-built backend. I think at the time, people thought we were
- 24:41
crazy. They thought like, "Hey, there are all of these, like open source alternatives. There's ClickHouse. Like, there are these things that actually, um, you know, allow us to do this and do this well." What you now see, and especially like you probably see, see on the LangSmith side, they've just built their own database. Um, Arise, they just built their own database like six months ago. I think people are starting to recognize, uh, that this is actually a, a problem that's not solved for with out-of-the-box solutions, and you need something custom. And so, in that regard, I think we're like a year, uh,
- 25:11
ahead of some of these, these competitors. And where that shows up is, is in topics, is in being able to generate real insights on data at scale. Um, there, there's probably some other o- other ones if you wanna, you know, come find me afterwards, we can, we can get into it. But, um, I'd say that, that last one, like the having our own backend and the things that we were able to build on top of that has really what-- uh, has really enabled us, I think, to separate ourselves from other players in the space.
- 25:41
Yeah.
- 25:44
Yeah, yeah.
- 25:45
Um, so, um, breakdown a little bit more into topics-
- 25:49
Yeah
- 25:49
... and you mentioned it, um, can you elaborate on, do you have examples to show how well it works or does not work? Uh, because, um, you know, I'm assuming you said, uh, look- looking at the slide, you said you're using embeddings to distinguish, uh, towards some of these benchmarks. The challenge here is that, uh, all of those, uh, traces
- 26:20
are going to look pretty much the same because they're all based on, you know, you're working on a particular problem, and so you have to get some benchmark or someone has a test set, and they all relate to one particular topic, which means that, um, there is not much to semantically distinguish, uh, between all these traces. So I'm curious about whether you have examples that show
- 26:50
how well it works.
- 26:52
Yeah, I mean, uh, like during the workshop... Uh, sorry if, if anybody couldn't hear, the, the question was just like examples of topics and, uh, a little bit more insight into how-- into perhaps the, the process and are we actually generating anything that is meaningfully different because the traces are relatively the same. They are going to be semantically similar. Um, I think you'll see this in the workshop, hopefully. Uh, I think the, the customers that are using this in production would, would probably disagree with you. That's why they are using it, is because they are, they
- 27:22
are seeing those, those differences, and they are seeing those show up as different classifiers, as things that they, they weren't yet aware of before actually running that pipeline. Um, I, I think part of it is just like, is getting your hands dirty a little bit and actually putting in your actual traces. Like, I have dummy demo data where I've, you know, used a script to hopefully create some differences that we create this. I think you naturally see even more of that when you have actual users interact with the agent. Um, the traces
- 27:52
look probably somewhat similar, right? We're invoking the same tools. It's going through the same, uh, workflow. But I think, you know, depending on what your agent is doing and f- and what your users are doing with your agent... A- and I think maybe to back up a little bit, I think a lot of us are building agents that are increasing in complexity or they're building systems that are increasing in complexity. And it, it creates, uh, I think some, some differences that, that will show up on, on screen. But you, you'll see when you create an account, you actually
- 28:22
have, uh, y- you have a free account. You can, uh, send data up to a certain point. You can even toggle on topics to see if it is actually generating some of these things. So would just encourage you, even after this workshop, to go play with it, um, and, and see if it does generate some interesting, uh, topics for you. Yeah, go for it.
- 28:42
Is it possible to use Braintrust for an agent that you are not serving to a customer, but you're using yourself, like Claude Code-
- 28:51
Yeah
- 28:51
... or, and analyzing whether they perform well or not?
- 28:55
Yeah. The, the question is like, uh, can we, can we use topics perhaps on like, uh, our own sessions with Claude Code and Codex and, and so on? Absolutely. One of the things that you'll see if you go to our integrations page is an integration with Claude Code, integration with Codex, uh, Cursor, OpenCode. And so you can-- uh, I'm actually working with, um, multiple customers right now where they are trying to do exactly that and layer topics on top of that. Uh, I, I think a lot of-- I don't know exactly where the question's coming from, but I think
- 29:25
what I've heard is that most engineers today, most teams of engineers are using those tools, and there, there needs to now be, I think, insight into the efficacy or how well these things are actually improving the workflow and, and, and those types of things. And I, and I think there's, there's probably some machinery that still needs to be put in place, but if you're just looking to, hey, understand what people are doing in those, those different tools, I think topics is a really good way to, to start to understand that. Yeah.
- 29:56
Um, can Braintrust be self-hosted or is it fully in-cloud solution?
- 30:01
So we have, uh, we have a few different hosting options. One is just full SaaS. That's what you're all gonna be using today. Uh, we have both a BYOC and a hybrid model. Um, and, and so maybe to like back up even further than this or than the question is like, Braintrust was actually architected from day one to support a hybrid, uh, deployment. And so by that, I mean, uh, the control plane, the web app, right, the, the metadata that we store is always gonna be hosted by us. The data plane, so your traces, your datasets, your experiments can actually
- 30:31
be hosted within your VPC, within your infrastructure. Uh, we deploy across all of the major cloud providers. Um, but yeah, you actually-- We have, you know, those customers that you saw up there, I would wager that 90% of them are, uh, self-hosted ones.
- 30:47
Yeah. Yeah, sorry. Green shirt.
- 30:49
How, how do you handle, uh, very long traces? For example, very long conversations that may, uh, pass through different topics? Do you split those long traces or
- 31:04
do you-
- 31:04
Yeah, so the, the question is, like, how do you handle long-running conversations with topics? Uh, some of it is, uh, I can probably show you this during the workshop, but there are different settings that you can configure that operate on, uh, that can operate on different traces, that can operate or can only be invoked after a certain time period. Uh, so depending on, like-- Generally speaking, I think there's an understanding of how long conversations tend to last that tends to inform the settings that you create. But there's, uh,
- 31:34
uh, getting into the workshop probably answers that a little bit better, uh, but I think we can answer that one. Cool. How about one more, and then we'll jump in.
- 31:42
Is the topic pipeline customizable?
- 31:45
Yeah. This, this, the prepro- preprocessor, we give you a default one. You can bring your own custom preprocessor. Uh, you have data that exists outside of, uh, what we can parse, uh, from input and output or what, what's going to and from the LLM. You have stuff that's in metadata. You have stuff that's within s-some other span that maybe we're not sort of pre-processing as part of that default preprocessor. You can absolutely create your own. Uh, the other one, again, like I think I mentioned, was the facets. You can bring your own facet. You can
- 32:14
build your own, uh, thing that you're trying to understand or to derive patterns of within Braintrust. But yeah, this-- I think one of the o- the other really cool things about this pipeline is it's completely hackable, uh, in the best sense of the word.
- 32:28
Cool. I see a 30-minute timer. I think we, we have till, we have till 11:00 though. Cool. All right, so I'm gonna assume everybody has the, the workshop, uh, or the, the QR code here, or at least the, the access to the repo.
- 32:55
Is that-- Should we zoom in a little bit?
- 33:04
Okay. Um, th-this could be a little bit different based on you having a Braintrust org or not. Um, like I said, uh, Noah is back here and can help if you, if you run into anything. Uh, also, you know, if, if you wanna, like, just yell at me for a little bit, that's, that's fine too. So the, obviously the first thing that we'll need here is a, is a Braintrust organization. And so I'll, I'll sort of walk, uh, walk through this, this guide with you all, and, and we'll start to understand, uh, exactly what I just described there. We wanna create that flywheel, and I think we'll, we'll focus a
- 33:34
little bit more on the online side, but just know that there is the, the flip side, uh, which is the evals, and I do have an eval as part of this that we can also chat about.
- 33:47
All right, so a little bit of a, a teaser there. It was not meant to do that, but we will create a new organization.
- 34:04
Um, if it, if it's your first organization that you've created, uh, you're going to navigate to that, that link. You may have some different steps at the outset. It may actually, uh, ask you to provide some AI providers and a, uh, generate an API key. Please go through that. I also have steps for those as well. Uh, you just see a slightly different, uh, flow because I already have organizations created here on the Braintrust side. If you s- if you see a screen like I do now, quick raise, raise your hands.
- 34:35
Do most people-- Are, are most people just g- creating a new Braintrust org for the very first time? Okay, cool. Awesome. I love to see that. Um, we've created the org. Is the next step to create a project?
- 34:50
Sorry?
- 34:56
Oh.
- 34:56
Internet issues.
- 34:56
Internet issues.
- 34:57
Yeah.
- 34:57
Love that. Okay. So nobody can participate.
- 35:01
I just got a text from someone in another session who's having the same thing.
- 35:04
Okay. Bummer. I don't know if anybody was here last year. We had the exact same, exact same problem. Um, well, we'll just go through this though, and I will try not to go too fast knowing that nobody else is, is doing this here. But, so you'll see the-- We, we created an organization. Uh, you'll also see we, we automatically created a My Project. Uh, you can, you can delete this if, if you'd like. Uh, we're gonna create our own project for, uh, this workshop.
- 35:35
Uh, I would label it this, AIE-Workshop. All of the sort of code in the repo, uh, assumes that you have done this. Now, you could certainly, uh, not do this and then change some other things in the repo, but just wanted to call that out. If you are following along, uh, just know that you'll need to, uh, make sure that the project that you've defined here is also matching what you have within the repo.
- 36:00
So the, the other two things that we'll need to sort of complete this is, uh, one, we will need to generate or create an AI provider. Uh, so if you go over to Braintrust and into your settings, you'll see a, uh, a place where you can actually configure your AI providers. Uh, I'll click this one up here. There, there should be a lot of different, uh, places for you to add some type of API key or, uh, what we just released very recently is Workload Identity Federation for some of
- 36:30
our major AI providers. Uh, but you can also-- What, what, what, uh, our organizations tend to also do or where I see them also build is, like, they'll have their own gateway, and they want to connect their, their gateway to Braintrust. They could do that using a custom provider. You sort of give the underlying spec that that gateway sort of conforms to, uh, how we are authenticating to that, and some other sort of URL authentication type. So generally speaking, uh, you can connect to a lot of different AI providers in a lot of ways. I haven't seen
- 37:00
anybody not be able to connect to anything, uh, from, from Braintrust. The, the other thing maybe just to... If, if you are sort of following along, uh, and you are using a gateway but that gateway has perhaps some firewalls, um, you probably see some, some errors on this side. Uh, just know that, you know, there are ways in which you can sort of allow certain IP addresses, uh, to those different places that perhaps only allow certain IP addresses, like a gateway. Um, I am going to stop
- 37:29
sharing just for a second so you don't see my API key.
- 38:07
All right. One more-- I'm gonna generate an API key also.
- 38:16
Okay.
- 38:33
All right. So just, uh, to recap what I just did here, I went over to AI providers. I added both OpenAI and Anthropic as, uh, API keys. Uh, the, the reason that you'll need this is, uh, if you do configure online scoring and you're using judges as part of that, uh, you will need to choose some type of model. So another callout here is that when you actually invoke those, those scorers, you are using your own models to do so. The only caveat to do that is that topics
- 39:03
pipeline which uses our Braintrust built-in models, um, for that complex pipeline.
- 39:13
All right, so we essentially just did, uh, sort of step zero here. We, uh, created a new org. We have a project called AIE-workshop. We have an AI provider, and then we have an API key. Uh, if you haven't done this, right, so we can now clone this repo.
- 39:55
Looks like I did not update something in the code.
- 40:07
All right, one more time.
- 40:27
All right, anybody thumbs up, thumbs down? Is this okay from a... You need to zoom in?
- 40:32
Cool. Okay, so let's, let's start going through this. Um, obviously we're gonna be using some, some different tooling here to, uh, to power this. I'm using uv as the, the sort of package manager to install all of my, uh, Python dependencies. So if you don't have uv already installed, you would use this command. Uh, I already do, so I'm just going to run, uh, uv sync and then add some dev dependencies as well. Uh, so that should create a virtual environment within your directory and then
- 41:02
should install all of the different, uh, things that are required to run this workshop. The next thing that we will do is, uh, create our .env file, and we already have one as a example. So you can copy that one over, and all you're doing in this one is you are adding an API key that you generated on the Braintrust side, and then you are adding a default model. Um, this is really only needed if you wanted to spin up the, the sort of chat UI, uh, oh, and then run, I think, evals as well. But these
- 41:32
are the, the two things that you would sort of create or modify within that .env that you just copied over, uh, to, to power and run this workshop. There, there are some other ones that, you know, there are some instructions in that .env file of what you could also, uh, change, like if you wanted to follow maybe those best practices that we talked about earlier of maybe creating or having a better or higher, uh, more powerful judge model, you could do that as well. Just if you're, if you're not aware what this should be,
- 42:02
there is the,
- 42:06
the, um, this is the, the .env that we just copied over. Again, I'm just modifying or adding my Braintrust API key, and then I added a, a default model, and this default model obviously is going to be something that is accessible via the providers that I, that I configured on the Braintrust side. So if I did OpenAI, I'm going to do a GPT model, Anthropic, uh, Sonnet, and so on. So just make sure that these, th-those two things, uh, mirror one another. The, the other thing that's gonna be really important, uh, for this workshop is to install the
- 42:37
Braintrust CLI. Uh, this is going to allow you to run arbitrary SQL to run evals, to view your logs, to create datasets- You think about that flywheel that I just described, a lot of those things are, are part of that flywheel, and it's something that we can plug into our agent to actually start to do itself, which you'll see at the end of the, uh, at the end of the workshop. So you'll see over here, I already have the, uh, the Braintrust CLI installed. Uh, I'm also gonna do a BT setup skills,
- 43:07
and I'm gonna use Codex. Uh, you might be using Claude Code or OpenCode or Cursor or Qwen or Gemini. That's supported also, but the idea here is that we are going to give our coding agent access to some skills to understand how to create a dataset, how to run evals, um, how to write SQL, and so on.
- 43:32
Hi. Yeah. It might be unpopular, but could we switch to light theme so it's not dark? Switch to what? The light theme so it's not dark. Oh, yeah.
- 43:43
That's a really good idea. Where-- I don't, I don't think I've ever actually done that in here.
- 44:00
Any idea? Um, maybe search in the help. I think you can do Control, Shift, E. What is it? Control- Control, Shift, E. B?
- 44:24
Uh, E. Control, Shift, E, and then search
- 44:43
colon E.
- 44:50
Let's do this.
- 45:21
I think it's
- 45:40
Command, Shift, E. Sorry? Sorry. Command, Shift, E brings the-
- 45:51
Command, Shift, B? Yeah. And then just type in colon E. Type in what? Uh, colon E.
- 46:10
All right.
- 46:14
I'm guessing you all, all were hoping to see me navigate around a terminal and not have any idea what I'm doing.
- 46:23
So we're, we're all good? This is better for everybody? Awesome. Sorry about that.
- 46:31
Okay, cool. The, the next thing that we will do is we will initialize, uh, the, the CLI. Um, I'm gonna do something slightly different because I have, uh, already initialized, uh, and I need to OAuth into Braintrust to access a different organization. So I'm gonna do... Oh.
- 46:56
So if, again, if you already have the Braintrust CLI and you created a new org, uh, you'll go through that, and then you'll just, uh, select that org that you just created. So for me, it was ABC DPG.
- 47:14
Cool. And then I will run that BT init, and I'll look for ABC. And so this is just going to create a, uh, config.json file in the directory that when BT runs, it knows what org and project to, to access. So you could imagine, uh, checking this into your, your repo. All of your engineers now sort of, like, have access to, um, you know, BT has an easier way to understand what org and project it should, uh, run an eval and query data from.
- 47:44
Cool. Let's, uh, now go through, uh, actually setting up some of the different things that we'll need as part of, uh, the workshop. Uh, I'm just going to seed a database. Uh, I'm gonna skip over this smoke command, but I'm gonna run make ready. Um, if you don't have, uh, Make or access to, to Make, there are the actual commands right below it that you can also run. Uh, what this is doing right now is just, just trying to understand, hey, did we set up everything correctly? Do we have the, the Braintrust CLI? Do we have a
- 48:14
Braintrust API key? Uh, do we have a default model? Those types of things. So, uh, I think we're all good to go forward now. Uh, then just to kind of give you a sense for maybe the, at least in this case, the, the trivial use case that we are, that we are powering, uh, very, uh, simple chatbot, but really what I wanna understand is, is this sort of wired up in the right way? Are we seeing logs flow into Braintrust? And so I just click the button in our chat. We should now see a log show up over on the Braintrust side with the, uh, the question that
- 48:44
we just asked. So here is an example trace that was just created, uh, based on the, you know, the, the integration that we have with OpenAI Agents SDK. Um, I, I think maybe to that, to that point, uh, just to show you, like, what sort of integration looks like, uh, within Braintrust, this is obviously a, a single example. Um, but all we're, all we're really doing is, is, uh, is this. This is what we're pulling
- 49:14
in from the, the Braintrust SDK, and we are setting a trace processor to a Braintrust tracing processor. Uh, th- there is also an easier way to do this. There is an auto instrument that is a single, um, line change. The, the reason that I did this is because I'm also using our gateway under the hood. Um, maybe I, I buried that a little bit. Uh, Braintrust also has a gateway, so all of those AI providers that you have configured within Braintrust we can now access, uh, from a
- 49:44
single interface, uh, which is what we're doing here. But we have our sort of, uh, tracing set up from our SDK, right, directly for OpenAI agents. We are setting that trace processor, uh, to a Braintrust tracing processor. And again, this means that all of those traces that are emitted, all those spans that are emitted are now, uh, going to Braintrust. Uh, maybe we'll just do it-- Actually, no, we won't do it from here because that's not light mode
- 50:12
. Okay, so I also have, uh, a script in here, or I have, uh, code in here to run evals. And obviously this isn't, like, the, the necessarily the point of the, the workshop, right? We're talking a little bit about observability. But, but I think, uh, in the context of observability, you need to think about the flywheel and, and part of that flywheel is evals. Um, the things that we are able to understand and derive about what happened in production should make their way back to our evals. Uh, and so there's also within the, uh, the repo that you have
- 50:42
a, uh, sort of dummy dataset, right? It has a few, uh, test cases that we wanna now run through our eval. So what we've just done again is used our Braintrust CLI, we've created a dataset, and we've actually loaded in a series of rows via that. Um, I'll now run the eval. Again, this will go, uh, run through that, that dataset. It'll sort of pass each row within that dataset and invoke our agent, and now we will store all of that sort of, um, uh, all that telemetry and the, the sort of results of that, of that eval,
- 51:12
uh, directly within Braintrust. But you could also imagine this is going to be incredibly informative when we do go run that flywheel, when we do go start to understand, uh, and we go make a change. Generally, you want to compare it against something. You wanna be able to understand, did I improve or regress from the last time that I actually ran that eval? Uh, so we have-- Actually, I'll just come out to Braintrust. Uh, so we should now see over here in our experiments page our eval that we have run. You see the, the series of scores that we've defined for this particular use case. Uh, you could
- 51:42
also see, and, and maybe I'll come out to this in just a little bit, but, uh, we have different scores, right? Just like I described earlier, there are, there are scores that are, um, more, more code-based, so required tools called, "Hey, I, I have this question. This is what I expect." Uh, then I have maybe a little bit more of that LLM judge, so support resolution and the tool use quality. So sort of a series or a, uh, mix of scores that can provide the right signal for understanding what's going on with our agent.
- 52:12
Okay. Let's, let's get sort of to the, the meat of this, though. The, the idea here is that we can now, uh, create some type of signal within our, uh, within our logs. And obviously we have a single log right now, uh, and we will need to, to generate some data and put it into Braintrust. But before we do that, let's, um, use those series of scores and push those into Braintrust so that we can now configure an automation so that they can run when logs come into Braintrust. So again, I'm gonna use this, uh, make push scores
- 52:42
command. This is actually going to look for, uh, the series of scores that we've created and push them into Braintrust. Uh, so there is a way in which you can define these in the UI. You could also, like, uh, you see here, define your scores in code and push them into Braintrust, uh, with the idea being that you'd want to run them as part of an online automation. Uh, or perhaps, uh, we also have a playground, and you could give that to another user of the playground, and they can use that same score but have it version-controlled within
- 53:12
your repo. Um, again, here are the series of tools that we just, uh, lo-- excuse me, the series of scores that we just loaded into Braintrust. Uh, you should now see if you go over to your scores tab the different scores that we created. So I have, let's, you know, support resolution as an example, right? This is going to be an LLM judge. A- and, and what you're doing here, again, is defining that criteria in which you want to judge something with the, with, with respect to the, um,
- 53:42
the, the, the user's interaction with that agent. Uh, one thing I'd call out here is that there are different sort of variables you can access and pass to an LLM judge. Uh, this thread is, is sort of what I described earlier. Um, it's, it's what you would see if you come over to our logs page and you click on the thread, which is this icon here. But it's, it's really just going to be that back and forth of that agent, right? This is really helpful, especially if you have a multi-turn conversation and you want to sort of
- 54:12
pass the entirety of that conversation to the LLM, you can use a variable called thread. There, there, there are some other ones, uh, that, that you could also use, like user messages or assistant messages, or simply input and output from a, a particular span because you want to create a score that understands and scores lookup policy, as an example. But there's a few different ways in which you can start to think about, how do I start to assess the quality of my agent across a lot of different dimensions?
- 54:41
So now that we have our scores loaded into Braintrust, the, the thing that we need to do now is actually wire them up to run when traces come into Braintrust. Uh, what, what that looks like on the Braintrust side is an automation. Uh, if you look in your logs page, you'll see up here in the top right- Uh, automations. This is where you can start to create rules that your scores can be invoked on. Um, there are-- If you, if you look at here at the scope, I think this is, is really important,
- 55:11
especially you start to think about, like, the different use cases that you have. Um, you could think I have a multi-turn conversation, but my, uh, turns are actually across different traces. I don't see them nested under each... Um, I don't see each turn nested under a single trace. So what I might wanna do in that case is select this group scope and then provide some type of ID, some type of way to understand how things are related. In my case, it's a conversation ID.
- 55:41
Um, the other thing that you might wanna do is create a score that is invoked for a particular span. Again, you want to assess some type of intermediate output or intermediate step, and you want to target a particular span. When you do this, when you click Span, it'll default to the root. Um, you could also add any type of filter for, I wanna recognize, let's say, my lookup policy, and I wanna say, in this case, something like that. So I want to understand if this is, is root or
- 56:11
the span attribute name is in lookup policy. Uh, there's also, again, like, the trace, which is what we will use for this demo. Uh, this is going to allow us to, uh, essentially access all of the different spans within that trace to produce some score. Uh, so in this case, I'm going to pull out my, uh, communication quality and my support resolution and tool use quality. These are all LLM judges. So i-if you're following along, um, just know that depending on the sampling rate
- 56:41
that you, uh, that you put here, you're going to incur actual model compute, uh, when you go and, and load data into Braintrust via these judges. So just wanted to highlight that. Again, if you're following along and, and perhaps maybe you wanna, uh, lower this a little bit. Uh, I'm just gonna go to one hundred percent just so you kinda get a, a really good signal or a really good idea of what this could start to look like.
- 57:04
So what, with this-- what, what I've just done now is I've said, "Hey, any new sc- any new logs that have come into Braintrust, I want to, uh, run those scores against those, and I wanna run against the out-- the, the, the trace."
- 57:20
Okay. Let's, uh, let's, let's jump into topics now. So there's, there's one side of our signal that we, that we can create. This is the more scoped-down way of understanding a failure mode, right? This is perhaps something that we have-- we're obviously evaling for. We have recognized to be something that, um, you know, we need to understand if we are regressing anything as we go and change something within our agent. Uh, topics, again, is, is more for those unknown unknowns, those, those areas that we are not yet testing for or that we're not yet aware of. Uh, you'll also notice right next
- 57:50
to our automations is an Enable topics. Uh, and actually, let me come back this. I think the, the first thing that we'll do, yeah, enable topics. So you'll see, you'll see a screen that looks like this, um, and you'll see those built-ins, so task, sentiment, and issues. So sort of what es- what I described, what are people or users doing with my agent? What is their sentiment? Are they happy? Are they sad? And then what are sort of like a broad range of issues that they are experiencing, uh, with respect to that interaction of that agent? So I'll enable
- 58:20
this. Um, this will apply to existing traces. Obviously, not a lot of data here, so, um, w-we'll eventually see some data show up here. The thing that a-again, allows you to bring your own way of understanding patterns within your data is where you can actually create your own facet. Um, you'll also notice here, this is where you can also define your own preprocessor. Uh, we're just gonna use the default thread one here today, but just know that if you wanted to bring your own function to Braintrust to
- 58:50
parse something within that trace, uh, you could certainly do that.
- 58:56
So if you look, there's also a file within that repo called Support Workflow Issue Facet. It's a markdown file. Uh, this is the prompt that we will be using for our custom facet. And so back in, uh, Braintrust, I'm gonna grab the, uh, support workflow issue name.
- 59:19
That will be our custom facet. And then I wanna grab the, the prompt here.
- 59:29
Um, the other thing that you'll most likely want to do is to exclude certain facet outputs. So you can see here I've instructed the LLM to output none when the following criteria has been met. Uh, what I don't necessarily want to see is, um, facets that are, that are generated where it's like, "Hey, this-- there's no problem here." Uh, I'm, I'm really only interested in the things that are, that are going wrong. So I will grab, um, this regex also at the top
- 59:59
and I'll paste it here. Uh, you can also choose to apply to existing traces. Um, not super important in this case because we have, uh, we don't have very many. So on, on the topic side, that's, that's really all you need to, need to do to, um, start to understand whether or not you can see some patterns within your, your data. Um, if you come back over to our, our repo here, and you look for, uh, perhaps a different way. Uh, I'm not gonna go through this, but I just wanted to highlight again
- 1:00:29
the, the ability to utilize the Braintrust CLI to do a lot of, um, interesting things. So th-that sort of workflow that I just went through in the UI, I can also do here within, uh, BT. I can actually configure these topics to, to run. I can configure those, those automation settings. Um, so here I've defined my task sentiment issues and my support workflow issues. Uh, the other thing that I did, you can also see some different configurations down there. There is different ways to configure the settings, right? So that
- 1:00:59
pipeline that runs in real time, you can control how it runs or What it runs on. And so by that, I mean you might want to add a filter. You might want to say, like, "I only want to run my topic pipeline on traces that meet this criteria." Uh, you may, again, wanna apply some type of sampling rate based on, uh, some understood metric you have about traffic and, uh, cost this could potentially incur, so on. You also want to understand, uh, the, the topic window. Like, when I generate my topic map,
- 1:01:29
what is the sort of range of, uh, traces that I should use based on some window? Uh, this one I'm going to-- Uh, this is, I think, maybe the, the question you had earlier in the green shirt around, like, how do I, how do I maybe configure this when I have long-running conversations? Uh, part of that is that group by scope, is being able to pull in perhaps other traces that-- other turns that exist across other traces. Uh, this also is another setting you can configure where you want to perhaps, you know, this could be a day,
- 1:01:59
this could be two days. Like, you all understand that the, the conversations that your agents are having or users are having with your agents stretch a very, um, long time period. And so you can configure sort of when the topics pipeline gets, gets invoked. Uh, I'm gonna do something really short here, um, but again, you can configure this based on your understanding of the user's interaction with your, uh, with your data. Or excuse me, user's interaction with your agent. Uh, then again, there's the same sort of scope that, that I highlighted with the scores.
- 1:02:29
Uh, you can have this run at the individual trace or, again, if you have that multi-turn conversation that exists across multiple traces, you would use this group and then use that sort of common, um, identifier, right? In my case, uh, I have a conversation ID that's being logged out as metadata.
- 1:02:46
So I'll click Save.
- 1:02:51
So now we have sort of, like, the, the topics configured. We've enabled our, uh, out-of-the-box ones, and we have also created our own custom facet, and then we have configured our automation settings to, uh, to be run when data comes into Braintrust. Uh, on top of that, so now we can actually start to see what some of this looks like. There is a command here where you can run, um... Essentially, I have some trace data stored out in an S3 bucket. Uh, we are going to just import it.
- 1:03:21
Again, under the hood here, what you're gonna see is BT sync. So we can actually take that data and push it into Braintrust. Um, maybe to your, to your question in the blue shirt about LangSmith earlier, I've actually helped a lot of customers import traces from other providers because they wanted to, like, maintain or still have some of that, uh, data. But just highlighting, like, this is a really great way to, uh, to do this at scale.
- 1:03:49
So, uh, you can see here, we just, we just loaded in, uh, nine hundred and sixty-eight traces, uh, one thousand and fifty-two spans. Or that's the, the rate, excuse me, uh, five thousand two hundred and sixty spans. So pretty, pretty quick way to get data into Braintrust. But again, like, the whole idea here is for us to start to generate some signal and start to understand where we can make some improvements within our agent. Uh, so you'll start to see, like, here's, uh, just some, uh, some examples of traces. Uh, you'll also
- 1:04:19
see that this is the-- This is actually two things you'll see here. Um, each turn is actually represented, uh, as its own trace. So this is the, the newest user question, and here is the, the output for that question. Uh, you could also notice that there are eleven previous messages. And so I might wanna do something where I, I pull in my conversation ID, and so now I can start to pull in other traces that are-- that have that sort of same metadata. But the, the sort of OpenAI agents
- 1:04:49
gives you somewhat of a cheat code where it actually just attaches all of those previous messages, so we don't necessarily have to go through this step. But just wanted to highlight another feature within Braintrust that allows you to group traces by some arbitrary metadata that you've defined within those that they, they, they share across those traces, and you can pull all of those relevant traces into the tree here.
- 1:05:13
Um, really quick, right, the thread view I think is a really great way to understand sort of like the back and forth, uh, of your agent. Uh, there's also, um, the timeline view, which is a great way to understand, like perhaps where bottlenecks exist, or you could even see there's really interesting information for, you know, cache hit, uh, across the different LLM calls. So this information I think can be useful in a lot of different types of, uh, of workflows. The, the other one I'd highlight, and I haven't really talked about this yet, um, this little blue icon in the bottom right is our
- 1:05:43
agent Loop. Um, everybody has an agent, right? Uh, so you can actually use Loop to do a lot of different things across the platform. Uh, you can use Loop, in fact, to create your own UI on top of that trace data. And so this is, I think, a really great way when I was talking about the other scoring method of I want to, uh, allow humans to go and review output of whether it's the agent or the LLM judges themselves and provide some calibration. This is actually something you can prompt Loop with, and it'll vibe code, uh, a
- 1:06:13
UI on top of that trace data. So just a different way to bring in perhaps another persona to sort of inform that flywheel. Um, you, you'll also see as part of this, like we're, we're starting to see some, some automations start to be triggered. Uh, so we have my, uh, judge automation, so communication quality, support resolution, and so on. And what you should also see is some topics starting to be generated as well. So in this case, right, we have a user who wants to check the status of a returned item refund, and so on. Um, we have
- 1:06:43
a negative sentiment. This user expressed frustration and dissatisfaction with the status of a pending transaction. Here's my support workflow, my custom facet here. There's an order record retrieval failure, both specific order ID lookups and broader customer searches failed to locate the order. So just wanted to, you know, two different things here, right? We loaded traces into Braintrust, and based on- The sort of automations and topics we've configured, we are now starting to get that signal in real time, right? So that complex pipeline that generates
- 1:07:13
patterns within your data shows up immediately, right? It shows up, um, based on how you've configured that pipeline, obviously, but this is the, now the signal that we're going to be able to use. Um, why this- while this starts to, uh, populate, I, I'm gonna go out to another project where I, where I already ran this. Um, so I, I showed you some illustrations earlier of what this, uh, perhaps starts to look like, um, within Braintrust. So at the top here, or, or in this page, what you start to see is a,
- 1:07:43
uh, snapshot in time of your, of your topics and the frequency with which, uh, the traces that we have recognized some pattern in, patterns in, where they are being grouped. And so in the first example, I'll expand this, uh, I can understand what users are trying to do with my agent, right? I can see at the top, uh, damage and transit resolution seems to be, uh, the most frequent thing that, again, users are, uh, using for my agent. And this is, I think, going back
- 1:08:13
to what I was describing earlier when I was in the slides, is I can actually use some of this information to understand, like perhaps where my agent can actually go, um, further, or the thing that I don't yet have built within my agent that users are trying to do. Um, all of this taken together will, will, will be used to sort of inform that flywheel. Um, a little bit further down, I have my sentiment topic. Um, my issues did not get generated. The, the reason this didn't get generated is because we, we look for at least 100
- 1:08:42
matches or 100 facets that were generated as part of that pipeline. Um, and this sort of agent that I built and the out-of-the-box prompt that we have for issues don't exactly jive or match, a- and we don't really generate, um, we don't, we don't really generate facets as a result of that, but I think that this is sort of a good thing. Like, we don't wanna just generate bad data. We don't wanna just generate signal that you can't actually go do something with. And so this is where I've created my support workflow issues. And in this case, I'm starting to
- 1:09:12
understand, hey, there's a record lookup failure. There's something going on with my agent, uh, that we haven't again recognized with our evals, but we are understanding that there is a lookup failure as part of that agent going through its process, right? This is the kind of information sort of all taken together, this with the, the context of our online scores, that actually allow me to go do something with it and, and improve my agent in some way.
- 1:09:38
Any questions?
- 1:09:48
Cool.
- 1:09:52
Okay. I wanna show you sort of like a, a workflow that, that I think, you know, I certainly did when I, when I first started at Braintrust, and I think a lot of, a lot of folks, um, they, they did as well, right? We've generated that signal. This is great. Um, now I need to be able to go do something with that. A- and so I think, you know, obviously within the, the Braintrust UI here, there's a lot of different things that I can utilize to start to understand or again, filter through that noise. Um, I might wanna look
- 1:10:22
at, you know, the, the different, um, classifiers or the different labels that were placed on top of these traces, right? That map to those topics that you saw in that other illustration. This I can use, right? To filter this down. Um, probably increase this. No.
- 1:10:41
Uh, this I can use theoretically to filter this down in some way, right? Based on the, the, the feedback that we are getting, uh, and understand the thing that I can do from there, right? So like I filtered my result set down, I'm drilling into a trace, and I'm recognizing by kind of parsing through the trace and the spans within there what went wrong. Um, maybe I'm again using my online scores to filter this result set down again, perhaps even further. But I, but I think, like, the, the workflow that you might go through is you'd
- 1:11:11
use this UI, and you'd start to understand all the different things, like all what's going wrong. I'm gonna take my task X and combine it with this sentiment, and I wanna see what sort of insight I can derive from it. Um, obviously generating the signal, really great. We have a lot of data now that we can actually go make some of these informed decisions. But I think you all can probably recognize the, the problem with that approach as well. It's like I am very reactive. The, the onus is on me as a user of Braintrust to like go through and click buttons and
- 1:11:41
filter and then parse through these traces. Like, there's gotta be a, a better way. Um, I might wanna use Loop. Loop is, uh, right, again, the agent that is embedded here within the platform that can do different things. Um, I asked it in this case, "Help me understand what I should improve in my support agent based on recent logs." So one of the things that you'll start to see that's, I think, again, really interesting from a, a flywheel perspective is how we can start to pull in relevant data to
- 1:12:10
inform that flywheel. Um, another thing that I think is, again, really interesting about Braintrust and specifically Brainstore, the backend that powers it, is that you can write SQL directly against it. All of that trace data, all of those spans can be accessed by writing SQL. Um, you may not wanna write that SQL, but your coding agent can very, very easily write that SQL. And so what you're seeing happen here is Loop. Again, Loop is our, our agent in the platform. I asked it a question, and what Loop did is it started
- 1:12:41
firing off a bunch of SQL queries. It's pulling in context that is relevant for the question that I asked it. What you're probably not able to do, and again, thinking at scale, you're not gonna pull in every single trace searching for signal, right, using your agents, your coding agents. You need that signal attached to your traces in some way, and you need that signal to be real. And then you need a very flexible way to start to query across all of that data so you can only pull in that relevant bit of information. Um, the, the
- 1:13:10
last thing that you, you don't see in this one, but you'll see a little bit later, is once I find those interesting examples based on the signal that was generated, I now wanna go do this. I wanna go investigate that trace. I wanna click into that span. I wanna look at the input. I wanna look at the output, right? This is the kind of analysis that I would probably do, but I'm not able to do at scale. But if I can offload some of that, again, to Loop to inform what I might go change or might, might go fix or might go add, or I could offload to my coding agent,
- 1:13:40
this is the kind of thing that actually I think drives agent quality in a really efficient way.
- 1:13:52
Um, so I'm gonna skip through some of this, but sort of what I described here is u-using topics and scores to understand and, and to, like, actually parse through that, that data based on the signal. Um, what you might also do now is you might actually use that, that filtered set of data. You actually found an interesting example, and you wanna go add that to a dataset, right? The thing that allows us to now go improve that agent, I wanna go bring that example back to my offline evals, right? I could do that here from the UI. I'm gonna add
- 1:14:22
this to my dataset. Now I have this sort of edge case, this thing that I need to go and sort of improve upon or a new feature I need to add based on what I've seen with this user's interaction of my, uh, with my agent. I can go add that back to my dataset. But again, like, the thing that becomes important here is adding the right thing back to the dataset based on the signal that was created within, uh, within our logs.
- 1:14:48
And then, of course, you'd go and you, you guys don't want me to see me code, uh, and, and figure out if I can improve the agent. But once I, once I did, I'd go run my eval again. I'd go compare that to the last time I ran the eval. Did I improve? Did I regress? That new example that I pulled in, did I, um, make it better? Did I actually improve those scores that I ran in an online sense? Uh, the other thing that you might wanna do is, like, you actually recognized these new edge cases. You probably wanna go create a score, right? I have a record lookup failure. That, that seemed to be the, the
- 1:15:18
problem that I encountered the most. I probably wanna go create a score that allows me to capture that before go-- before pushing my, my agent to production. So there are all of these steps, right? These things that you would probably want to do to create that flywheel. Um, and what I think becomes really powerful here is when you can run this in an automated sense. When I can plug a skill or I can plug some type of workflow using all of the primitives that I just described to automate all of that. Um, and
- 1:15:48
so if you look here at, uh, at our workshop, uh, markdown file, there is a GitHub repo for skills that, that we have out there. There's actually a single skill at the moment. Uh, the idea is to give your coding agents the ability to run that automated flywheel, right? We want to actually go through that exact same thing. I want, I want Codex or Claude Code or whatever to write SQL. I want it to view traces. I want it to view spans. I want it to, um, create new scores and run evals and modify something
- 1:16:18
within my agent, and I wanna do that, um, you know, watching my, my screen. Yeah, go for it.
- 1:16:23
What do you mean by get dataset? You look at the traces and you want to expand the training set. What, uh, what that means?
- 1:16:34
Yeah. The, the-- So if you look at my, if you look at my eval that I've created...
- 1:16:53
So this is, uh, this is sort of what an eval looks like in, in Braintrust. Um, your dataset here, uh, is actually, uh, part of, like, the, the commands that we ran at the outset, was to load our sample data, like, our series of test cases into Braintrust. But, but you could also imagine, like, this is, at the outset, this is, like, my sort of scoped-down view of the world in which my agent operates, right? These are the test cases that I want to give to it to understand, "Hey, should I go push this thing to production?"
- 1:17:23
But the, the sort of, like, operation or the thing that I'm describing here is that the-- Where you start to improve the quality the most is, is when you can curate traces, curate examples from production, and bring those back to your, your dataset. So this right here, like, this sort of, like, very sim-simplistic workflow is what I was describing. This pa- this, um, this record here, this row, this trace, it perhaps is indicative of some failure mode, and I
- 1:17:53
wanna bring this back to either an existing dataset or a new dataset that I want to run an, an eval against.
- 1:18:00
Eval dataset, not a training set.
- 1:18:02
Correct.
- 1:18:10
Um, so I think, like, the sort of culmination of, of all of this is, uh, how we can, uh, use all of those, those sort of primitives that I described earlier. I think the topics pipeline, uh, running behind the scenes is, is kind of the, the star of the show. Um, the, the, the importance of adding to the, the dataset, though, is
- 1:18:32
your, your evals should evolve with your agent. And when, when I think about evals, I think about a dataset, I think about my agent, right? So as a function that invokes my agent, and then I think of the series of scores that allow me to understand good or bad with respect to those test cases being passed to my agent. But when I create my eval set, it's a very sort of, like-- It's a way for me to establish a baseline. It is certainly not representative entirely of what users can
- 1:19:02
do with my agent. And so by me adding new records to that dataset, it allows me to capture different per-perhaps failure modes over time and- Go make changes within my agent so that, right, we don't, we don't continually experience those, those issues in the future. But it's, it's more about, like, the, the evolution of your agent and ensuring that all of those sort of like things that come along with your evals evolve with it, um, so that you can,
- 1:19:32
one, capture new things, but also, like, as you go change your code, you're not regressing based-- you're not regressing your agent based on previous interactions that you've sort of curated into that dataset. Um, there's, there's a few different ways to, uh, install this. Uh, I'm gonna choose just the top one.
- 1:19:50
Uh, so if you have the, the GitHub CLI, uh, installed, you can use that. Uh, there's also, uh, npx, and there's just cloning the, the repo locally and, and following the different install, so, like, with respect to the agents that you're using.
- 1:20:17
Okay. So we, we have the, the skill file. We've installed that. Um, so now I might wanna just, like, run this, this flywheel. I wanna see this in action. So I'm gonna undo two different things here. I'm going to run this, and then I'll actually show you, um, a previous example so you're not just sitting here watching Codex run. So we'll do that. We'll let that run. I'm gonna go out to ...
- 1:20:49
another repo.
- 1:20:53
Okay, so same, same prompt. Um, and actually, let me-- I think I sort of skipped over one thing here that I think is probably important to highlight, right? The actual, uh, logs here that were generated or, like, the topics that were, that were generated. So we should now-- We have all this signal, right? This is the thing that I just, um, uh, executed Codex against, this project. It should now be able to pull back all of these different things, right? These tasks, these sentiments, uh, the support workflow issue, all of this and
- 1:21:23
these scores are attached to those traces and are now part of this workflow.
- 1:21:29
So we'll let this kind of go, uh, in, in this, in this terminal over here. All right. Well, we'll come to this one as it's running. So sort of like what you saw in, in the Loop session, uh, what you'll start to see here is, uh, Codex writing SQL commands. So BT SQL, and we're going to pull back all of that information or the relevant information into the session. And so we can flexibly query across this project now and across these traces, and
- 1:21:59
we should be able to also recognize and understand, or Codex should be able to recognize and understand those topics and those, uh, those scores that were run. Um, so we have, you know, you can see here we have nine hundred and sixty-eight support traces. That's exactly what we loaded in. There are scored failures, a visible cluster of escalate to human tool errors. Um, I'm going to inspect the classifier outputs and representative failed traces and turn that into a concrete summary. So again, the ability to flexibly query with that
- 1:22:29
SQL tool, uh, becomes pretty impressive. Uh, it looks like it was able to recognize something with our scores, uh, and then is, you know, continually querying. And then this is what I described earlier of actually being able to drill into the individual trace. So it found something. It found some example that would allow us to, um, improve our agent again in some measurable way. We wanna view that trace, write some more SQL, um, but this is really what allows us to be very comprehensive and, and actually pulling in all
- 1:22:59
of that really rich data, uh, but only pulling in those things that are actually going to drive, uh, some, some change that we can make within our agent.
- 1:23:10
Okay, so let's see what, what else we just did here. Um, looks like we're looking up how to run evals. We are viewing existing datasets. Uh, we just created a, a new dataset. Actually, we're in the help command. Um, but, but hopefully what you're able to see here
- 1:23:28
is this.
- 1:23:32
This is what we're trying to create. This is what the-- This is what, like, the, the primitives in, uh, in Braintrust, our topics, our Braintrust CLI, the back end allowing you to write SQL, uh, having a platform that does both evals and observability. Um, all of these things sort of taken together allow you to, again, create that, that flywheel and then do so in a, in a really meaningful way. Um, more than happy to take some questions. Right now it's, it's really just
- 1:24:02
waiting a little bit on Codex to, to cook, but curious if, if anybody has... Yeah, go for it.
- 1:24:09
Can you self-host this on our infrastructure if we don't want the data to be out of-
- 1:24:15
Yeah, we have a hybrid deployment, so it means the data plane, so where all of the data resides, so your traces, your experiments, your datasets, they would remain in your infrastructure. But the control plane, right, just like the web app that you're looking at, uh, would stay with us.
- 1:24:30
So is it only in the enterprise plan or even the previous COs can take it?
- 1:24:37
The, the hybrid deployment is only available for the enterprise plan. Yeah. Um, to that point, just if you're curious, like, there's two different flavors of that. There is a BYOC one where, again, data plane is deployed in your infrastructure, but we manage it for you. Then there is more hybrid deployed in your data plane, your team manages it. We have sort of a, a Terraform provider or Helm charts based on, uh, the, the, the cloud that you're on. Back of the room? You're gonna have to yell.
- 1:25:06
Can you, can you comment on the-- Can you comment on the most popular, um, models that you see your customers using?
- 1:25:24
Are you making a pattern with this?
- 1:25:27
Um, I, I think for the, for the most part, and I think this is changing a little bit, uh, folks are using more of the, uh, the models from the, the big model providers. Uh, that's generally what I see. Um, I, I think that's, that's starting to shift a little bit. Um, I, I know, like, at the outset, it, it was very, very heavily skewed towards the, the major model providers. Um, I don't know, like, specifically, I haven't seen small models necessarily, but I have seen folks start to, um, perhaps
- 1:25:56
fine-tune their own models and, and see if it can't drive... You know, the whole idea, again, like, beyond, behind, like, evals as, as one example, is to be able to understand if I can swap in and out models without sort of introducing regressions or, uh, deteriorating quality of the agent. But yeah, I'd say largely it's, it's the, the major model providers, perhaps maybe shifting a little bit.
- 1:26:23
Uh, what model would you typically recommend for the topic, um, or sorry, the summary, the summary for the trace? What model would you typically recommend?
- 1:26:35
Sorry, is this, like, for an LLM judge?
- 1:26:38
Um, no, for the topic-
- 1:26:40
Oh
- 1:26:41
... pipeline that you have. There's-- I think you said you can bring your own model to do the summary. I don't know.
- 1:26:48
Yeah, so, uh, important call-out, um, could actually see it probably back at, at this screen when you're looking at the, uh, topics automation settings and Advanced. There, there-- The only option you have for both the summarization and vector embedding step are Braintrust-hosted models. Um, there, there's a couple reasons for that. One is that we, we've trained these models to do specifically this,
- 1:27:18
and, and two, like, the, the complexity of, of the pipeline sort of, um, makes it so that you can't just necessarily plug in any model and, and, and hope to achieve similar results. Uh, we started off by allowing different models to be plugged into that, but the, the topics pipeline or the topics that were generated as a result of that just were not, were not good. And so we have enforced it, so it's only Braintrust-hosted models for this step.
- 1:27:49
When you were showing the scatter plots earlier of the topics and the percentages going up, can you show that as, like, trending over time?
- 1:27:56
Yes, you can. Um, let me go out to...
- 1:28:09
Uh, so because all of that, uh, all of that data exists as metadata, essentially, on top of the trace, you can pull it out for different, uh, visualizations. So I think in this case, I have a saved view. Um, and just, you know, some, some Braintrust stuff here. Uh, if you wanted to, uh, create your own sort of dashboard or your own view, uh, this is how you would do that. Uh, all of these views are like, uh, you can save it, you can create your own custom charts. But, uh,
- 1:28:39
up here at the top, you can see sort of sentiment over time, tasks over time, uh, top sentiments, top tasks, like the things that you've built custom, uh, can be pulled out in here. Uh, you can see that, like in this case, I'm looking for classifications.sentiment is not null, and I'm grouping by that label. So these, these become pretty simple to, to set up. Yeah, I think-- So this, again, grain of salt. Uh, generally speaking, what I've seen though is, like, uh, newer agents or newer things that are just being pushed to production have a higher
- 1:29:09
sampling rate, and over time, those sort of like known unknowns or the things that your scores tend to surface become less and less, and the signal that you derive from it will, will also become less, and so you start to scale back that sampling rate. Uh, I, I think there's also, if you're doing things with an LLM judge that is a little bit more on the classification side or on the taxonomy side, you can look to alter or you can look to swap that out with topics. Um, Dropbox is a customer, a customer that
- 1:29:39
I work with, and they have this, like, very complex series of classification scores that they run to understand some taxonomy, uh, with respect to their traces, and they have experimented with, uh, using topics to do that. Um, one, because-- Actually, the most important reason for that is, is it's a much more cost-effective way to run across the majority or 100% of their traces, where the LLM judge just was not able to scale to that, and you're not able to generate as much signal. So I
- 1:30:09
think there is, to-- maybe to summarize, there is a maturity that I think informs what the sampling rate should look like. It tends to scale back as the maturity of that agent increases. And then there is the types of things that your, that your LLM judge might be producing. If it's a little bit more on that classification or taxonomy side, you can look to swap that out with, with topics. Yeah. So not directly within the UI, but you could create a code-based scorer, right? You could imagine, like, I, uh, invoke scorer one. Based
- 1:30:39
on scorer one's output, I go do something, and I invoke scorer two. The return of that function is an array of scores, right? Is an array of score objects. But it's, yeah, it's certainly possible, but you define it within code. There's nothing within Braintrust that allows you to say, "Hey, I want to run scorer A. Based on the result of that, I want to wire it into scorer B or C." But it is certainly possible just to create it within code.
- 1:31:04
It's not possible to use a filters option to see, uh, scorer results as one of the possible filters?
- 1:31:12
Oh, yeah. You could, you could also just, um, your filter could look specifically for the presence of that scorer span and then only be invoked when, when it, when it's there. Yeah.
- 1:31:33
Or that they become much more efficient at, uh, discovering about their data?
- 1:31:41
Um, I forget the, the customer, but this was surfaced the other day, and it was more on the, on the task side. Um, I don't remember specifically what it was, but it was actually something that they didn't realize that users were doing with their agent. So it was more the use case of, like, uh, feature request or, um, you know, unknown unknowns, but with respect to what users were doing with their agent, not necessarily a failure mode. And I think that's been the-- perhaps the most surprising thing for me is that people are using
- 1:32:11
it to uncover those types of things. I, I think when I was starting to, like, play around with the feature and talk to customers, it was, m- you know, the, the majority of it was, was, for me at least, focused around the types of issues that you would be able to uncover. But I think there's just, like, such a broad range of patterns that you're now able to sort of, uh, attach to your traces and then use those sort of patterns together to inform something about, like, the, the world that you didn't know about your agent. Um, so yeah, it's-- I don't have the
- 1:32:40
exact, uh, thing that the, uh, topics uncovered, but it was much more around, like, "Hey, I didn't realize this, uh, these users were using our agents or trying to use my agent for this."
- 1:32:54
How are the projects intended to be scoped inside of Braintrust?
- 1:32:58
Um, ri- right now, I would say they're, they're, like, use case driven, or you could even imagine, uh, agent specific. So in my case, I have this support agent, and in my project, I am writing my logs for my support agent into my support agent project. I'm also storing the datasets that I'm running my evals for my support agent. I'm also storing the scores for that support agent in that project. So they tend to be a, uh, a, a container
- 1:33:28
for relevant resources for a use cor- u- use case or an agent. Uh, we are introducing very soon a, a sort of, uh, first-class agent object into the project. And so you could imagine maybe there is a multi-agent system, and you want all of the agents' logs to go into the same project but still have some delineation between agent A, B, and C. Uh, so you could imagine doing that as well.
- 1:33:52
Are there any uses in Braintrust that maybe give me a global view but still to my permissions? I might have permissions to five different projects but four more projects.
- 1:34:03
Um, yeah. So the, the sort of monitor page that I showed earlier is, uh, project scoped. Um, if you wanted to create a more, like, organization-scoped view, you'd have to, uh, create queries and pull data out of Braintrust. Um, that'll eventually change. Uh, i- it's more, how do we, how do we allow our customers to do this at scale? Um, you could imagine trying to query across however many different projects and pulling that data into a dashboard, uh,
- 1:34:34
becomes a little bit challenging. So yeah, at the moment, we are, uh, project scoped, but you could imagine that, that, that starting to change over time. I, I mean, I think that's probably fair. Um, I, I think maybe what you'd potentially see is your, your evals changing, um, as often as you are changing the agent. Like, and it's probably, it perhaps isn't even that frequent. Um, but I do think it does have the prospect to change in some way based on what you're able to
- 1:35:04
infer from users interacting with that agent, right? Like, this could be, this could be a prompt, this could be a tool, this could be a new tool, this could be, um, something within the runtime or the architecture that, that needs to change. And to me, like, what I've seen more frequently is that your evals will generally change as a result of something that you are changing. Like, obviously, if it's a simple prompt change, like, you're probably not changing a lot within your evals outside of maybe curating some new examples into that dataset. Um, but,
- 1:35:34
but I do think what I've seen is that agents tend to evolve, um, somewhat frequently, and I think if you are being informed by actual interactions with, uh, with users in a real way, maybe they're actually changing a little bit more than you would expect. Um, but yeah, I do think it's a fair sort of, uh, piece of feedback where, uh, the agent isn't very complex. It doesn't have, or it's not, you know-- It, it doesn't necessitate a lot of changes based on the scope of, of what the agent is doing, and
- 1:36:03
so the evals may not change too much as a result of it. Uh, I do still think there is value in being able to generate that signal in production of, of what's going on. Um, especially if you can do it in a cost-effective way, then, then yeah, it becomes, uh, like that part of it I still think there's, there's value to. Yeah, like I, I have customers that do both. Um, uh, I, I think some of it's preference, some of it, um, sometimes is just not understanding what Braintrust can do. Um,
- 1:36:34
allowing-- So, like, the, the code-based scores I think is a really good example of this, where you may have a really complex code-based score that, uh, uses different Python dependencies or TypeScript dependencies, and you need to run that, or you think you need to run that in your own infrastructure because you have all of these sort of, like, complex dependencies that are powering it. What you can do within Braintrust, when you, when you go push a score, and even when you push a code-based score, you can bundle dependencies within that push. And
- 1:37:04
so the things that that score relies upon can still use those. And I know some customers that I've worked with just didn't understand that they could, they could do that. Uh, to me, the value is, like, not having some of that, like, asynchronous code within your own code base that, um, looks for those triggers or is like, it looks for, "Hey, this trace was generated, and I need to wait sixty seconds from the last event to go sort of be invoked." And so to me, the value is, like, offloading some of that complexity into
- 1:37:34
Braintrust, though you still can do that. Um, one of the, I think, unique ... things about, like, the Braintrust data model is that you can actually update spans in place. So if that span already exists within Braintrust, it is mutable. And so you could add, you can, um, you know, change the metadata, attach scores, attach different metrics. And so if you wanted to do that in code, you had sort of like the, um, you know, the machinery in place on your side, it, it becomes pretty trivial to
- 1:38:04
go and like, "Hey, I have this span ID. I want to go attach these different scores and metrics to it." So I think some of it is just, like, it's preference. Um, some people just want to have a little bit more control. Perfectly fine. Uh, you can certainly do that. So to me, it's, uh, one, it's RBAC probably. Like, you, you still want to, like, be able to write some of that trace data and then lock down who can see it. I think that workflow that I just described is, is, uh, like that flywheel, that automated flywheel,
- 1:38:34
is actually a great way to do that because I can surface insights to people. I can generate a PR based off of all of those insights, but I don't give them access to the underlying traces. I give them access to the analysis of those traces. Here are the SQL queries that I ran. Here is the aggregate data that I pulled back. Um, if you have just, like, no eng-engineers are able to actually view that data, write it to a project that nobody has access to, um, lock it down from that perspective, but then create sort of a service account that
- 1:39:04
is actually able to, like, go run that sort of flywheel on top of it so you don't lose access to all of those insights. Yep.
- 1:39:14
What would you say are the most underrated features of Braintrust for AI engineers?
- 1:39:19
Most underrated features. Um,
- 1:39:24
m-my guess is the Braintrust CLI is very properly rated. I, I think I, I, I tout that, uh, any chance that I can get or I, uh, tell any engineer. Uh, the, the reason I say that is because I think we, we released it in late February. We had a user conference. Um, a-and I still think, like I'm going out to customers today and I'm asking like, "Are you using the Braintrust CLI, CLI? Are you using it?" And some of them still aren't. Um, but to me, most, like I'm-- I would wager 100% of you are using
- 1:39:54
some coding agent, right? Like, this is how you are coding today. Um, having the CLI as part of that, like I, I hope you were able to see some of the, like, types of workflows that you were able to create. Um, to me, it just is, it's become so incredibly compelling. You're already in a terminal, you're already in your IDE, whatever it is. Uh, use the CLI, attach that or, or give your coding agent the skills to understand how to use that, and then use that to augment your workflow. Um,
- 1:40:23
th-the other one that, the other one that could be, like, somewhat compelling, it depends I think a little bit on your, your organization and the different folks contributing to that flywheel. Uh, we have a feature called remote eval. It's a way for you to expose that eval to a playground. This is more like how can I perhaps, uh... Like in, in this scenario, there are perhaps subject matter experts, there are product managers, there are folks sort of like contributing to that flywheel in some way.
- 1:40:53
Like, they are modifying prompts or they are doing some of that hill climbing. And we sort of crafted like all of this, um, you know, the-these applications for them to go plug into in a load co- low-code type of way. The remote eval allows you to expose that eval that I showed to you earlier, that like eval with a task and a dataset. You can expose that to a playground with parameters so that the user of the playground can actually go and start changing the system prompt or a tool description or something like that. Like, why that might be
- 1:41:23
interesting for an AI engineer is that you don't necessarily have to be sort of like the bottleneck in producing applications or, uh, producing something for a product manager to go plug into and allow them to modify a prompt or contribute to that flywheel. So maybe not as, uh, not as impactful as the CLI, but potentially given your organizational structure could be. I think I got most of that. Tell me where, where I'm off. Uh, I, I think, like, the, the thing that I, that I started doing here
- 1:41:53
is exactly that. Uh, like Braintrust itself, at least at the moment, doesn't have access to your code base, right? It doesn't understand, like, what you've written, and based on the information it has, it can't go recommend a change necessarily to specific functions or agent architecture. But when you bring the, um, Braintrust CLI into your sort of coding agent and then attach that, obviously, like I'm, I'm within my repo here. I'm able to now go through, let's see if we've been able to, like, get something here. So in my
- 1:42:23
session, uh, based on what I've been able to sort of understand, we are creating new, new scores. So there is a score change. Um, there's-- we can skip the tests.
- 1:42:37
There is, uh, looks like newer tools that we were able to add to our agent, again, based on some of that insight.
- 1:42:46
So I, I think, like the, the answer to your question is like, this is absolutely possible. What, what I've just sort of demonstrated here is that flywheel of like, hey, figure out what's going on in production. Use my scores or use my topics because I have sort of attached information of s- you know, issues or sentiment or whatever it is. This becomes incredibly compelling now or interesting for that agent to pull in and then go
- 1:43:16
modify your agent in some way. Um, I think eventually as you get further down here, see, oh, so now we have new examples to our dataset. We have, uh, looks like we've created a new dataset with those examples. We're now gonna go run our eval probably a little bit further down here. Uh, you can also see, like we are querying directly from that experiment that was run. So that whole sort of flywheel is captured here within, um, within this agent
- 1:43:46
improvement skill. And I think hopefully I answered your question to some degree. Cool. Um- This to me is, like, one of the most compelling parts of Braintrust right now. Uh, we're all sort of developing our agents so- to some degree with the help of a coding agent. If I can bring the right context, if I can have access to the right primitives that attach the right data to my traces, then this, this flywheel, this, like, automated, automated flywheel becomes pretty interesting.
- 1:44:13
I was gonna ask kind of a follow-up. Like, how-- it seems like you're close to the point where you can start maybe automating some of this flywheel for customers. Like, how... Is that something that's, uh, feasible?
- 1:44:30
So, uh, I, I, I absolutely think it's feasible. Um, if I could show you... I'll show you this example. Um, actually, if I come back here. So within the platform right now, there isn't, like, a button that you hit that says, like, "Hey, go do the, the flywheel." But what I showed you is that you can very easily plug into this using all of the tools that you already have. The things that you're now able to, to do on top of this, right? You could imagine, like, maybe plugging this into a GitHub Action that runs at some
- 1:45:00
particular cadence that, that gets triggered based on something, and we open a PR. Uh, you could also imagine it's, like, triggered off of a Slack message, right? There are lots of different ways in which I think this could be triggered. Under the hood, all of the things are available for you to go sort of opt into this. Um, one example...
- 1:45:21
Here's a, here's an example of, uh, a GitHub Action that, that I, that I created. So under the-- If you look, there's a, a YAML file that actually runs through, uh, Claude Code. It pulls in the Braintrust CLI. It pulls in the skill. But the, the sort of, again, like, the thing that we, we need still here is we need to understand, in this case, like, what changed. And this is, like, my very, very trivial supervisor agent, uh, example. What changed? Why these things changed? Why did, why did, uh, in this case, Claude Code recommend these
- 1:45:50
changes? Here is the, the actual impact, right, based on the evals that I, that I ran. Here are some links out to Braintrust where you can actually go inspect those. Here are these regret- regressions that I pulled in based on that analysis. And then here is your PR. You, you're the reviewer. Go figure out if this is, like, a meaningful thing that we should go and, and, and produce. Um, so I absolutely think this is something that, that you can opt into today. Like, we have customers that are absolutely doing this. Um, one of them that was
- 1:46:20
on a, the, the first slide that I showed is using this to generate five to 10 more PRs per day than they were before. Um, if you have, again, all of that information, if you have, like, the, uh, that really intelligent Codex, Claude Code, or whatever querying very flexibly over your production data and pulling in the things that are, that are interesting, I think i-it becomes pretty compelling from a lot of different use cases.
- 1:46:44
Is the workflow described on the label, like, something that's open source?
- 1:46:48
Yeah. Yeah, that, that Braintrust skills repo, this is open source. Yeah. Th-there's a, there's a YAML file in here as well that, like, that went through the actual, um, like, GitHub-- Like, it did it within a GitHub Action open source as well. Yeah.
- 1:47:05
Can you say more about how you connect the playgrounds to your own agents? Is it something where you can host your agent somewhere-
- 1:47:13
Yeah
- 1:47:13
... and you can provide some kind of interface? Like, what does that interface look like? Does it need to be conforming to some standard, or can it be completely custom?
- 1:47:22
Yeah. So what it looks like is...
- 1:47:32
Um, so I, I, I showed earlier how we write evals, or one of the ways in which you can write evals with Braintrust is by using that eval code. So here's my eval for a different project that I'm taking you outside of. The, the one sort of thing that I call out here that's different that you didn't see in the previous code is the parameters. Um, under the hood here, I am essentially, uh, I have exposed different parameters that, that touch my sort of agentic system. And in this case, it's, you know, something somewhat
- 1:48:02
trivial, right? It's the supervisor agent. But if you look over at the playground itself, um, come over here. So you have different things that you can pull into Braintrust, right? The-- I think, you know, I don't think a lot of people use it very much for prompts anymore, uh, outside of, like, the, the, the very basic stuff. But that remote eval, I can spin up a dev server. So I can be local. I could write, um, eval my filename --dev. It spins up a server, and it, um, automatically looks for it on
- 1:48:32
localhost:8300. Um, you could also expose this server within your own infrastructure. You have sort of, um, created that eval, and you have exposed it, and then you've also exposed different parameters. So in this case, I've exposed a system prompt and the model for that system prompt. And if I scroll down a little bit further, there's also math agent prompt model, research agent prompt and model. But I can sort of interact with all of that complexity, all of the tools that, uh, each one of those agents have directly from the
- 1:49:02
playground. I can simply click Run. It'll go through, and all of the code execution remains on the server, and then the results of that eval are streamed back into the playground.
- 1:49:15
The evaluation
- 1:49:21
on
- 1:49:25
how
- 1:49:27
to...
- 1:49:27
Yes. Braintrust will initiate a POST request, um, to an eval's endpoint that you have exposed on that server. Uh, so there's, there's, like, very, uh... I can show you another example here, um, where you can-- There, there's a sort of, like, create app function that you can use, or you can sort of strip out the things that, uh, that, that you would need to spin up the server. But...
- 1:49:53
That, that remote eval server right now is, is set up on Modal. I'm using this create app. It exposes all of the evals, at least at, the way that I've configured it, in my evals directory, and now those are exposed to the playground, um, for users to go and play with. One of the things that we are doing internally right now is we are-- and, and it's really driven a lot by interactions that we have with our customers. I think I mentioned this at the outset, like, people are thinking very heavily right now about
- 1:50:23
how well their, their coding agent is actually, uh, like how well, how efficient they are with their coding agents. Um, and part of that is actually it's running evals on the skills that they're using, is understanding the sessions that they have with Claude Code and Codex and so on. I think there's still-- I say that with, um, I say that because we are putting a lot of effort behind the scenes into creating some of the different primitives and machinery that our customers can go, like, plug into or use
- 1:50:53
to start to go understand those, those types of things. Um, yeah, just because you see, like, two commits and, and one star there is not indicative of, like, where we, uh, like what we perceive to be as a very, very important problem, uh, for engineers today.
- 1:51:12
Yeah. Cool. Awesome. Thanks, everybody. Appreciate you, uh, coming out.