AI Engineer Europe 2026
Shipping complex AI applications | Braintrust & Trainline
About this talk
Braintrust and Trainline present a hands-on workshop on shipping reliable, production-grade AI applications. Giran Moodley, Oussama Hafferssas, and Mayank Soni discuss Trainline’s agentic systems, decomposing LLM workflows into tool-enabled specialist stages, instrumenting applications with the Braintrust SDK, tracing calls and diagnosing failures, and continuously evaluating production logs through automated online scoring.
Chapters
- 0:00Workshop introduction and speaker introductions
- 4:14Production AI challenges and Trainline agentic systems
- 23:52Workshop resources and multi-stage tool-enabled agents
- 49:01Tracing, failure diagnosis, and SDK instrumentation
- 1:14:26Production automations, resources, and closing remarks
Talk transcript
- 0:00
[upbeat music] Good afternoon, everybody.
- 0:15
Welcome to sunny London. Hey. Um, is everyone's first time here? I think it is, because this is the first conference. Amazing. Amazing. Well, thank you very much for joining today's session.
- 0:26
Um, hopefully you are in the right session, but for those who, uh, need to double-click, uh, this is a hands-on workshop to delivering quality AI applications with Braintrust. And we'll also be partnering with our colleagues at, uh, Trainline, which I'll introduce themselves shortly.
- 0:40
Uh, so probably you're wondering, "Who's this guy? Does he even go here? What's his name?" Introduction to myself is, uh, you can call me Giran, bit like the band Duran Duran with a G.
- 0:48
So hopefully folks are similar one fan. And, uh, I've had a bit of part of helping organizations and enterprises help scale adoption of mission-critical systems, and now moving into the age of AI.
- 0:59
Um, you know, background in, uh, formal mathematics, so I know there's all the rage going from machine learning to data science and now AI engineering, and this comes at a very topical time.
- 1:08
Uh, feel free to connect me with LinkedIn if you wanna stalk or just want to, uh, you know, just do general chitchat around this area. I'm also joined by my two friends over from Trainline.
- 1:17
If you wanna come over and introduce yourselves, team.
- 1:20
Of course. If the mic is working. All good? Yes. Hello, everyone. Uh, thanks for coming to this workshop, especially after that lunch break and, uh, uh, and the sunny time outside.
- 1:31
Thank you very much for coming here. Uh, so my name is, uh, Oussama. I'm a senior AI ML engineer at Trainline. Uh, for those who know me... Yeah, one person.
- 1:41
Uh, yeah, I was a staff platform engineer before, and I have, uh, background in computer vision. Um, and I was doing also mobile apps on the side, so nothing to do, uh, with, with AI at, at some point.
- 1:53
But yeah, now we are doing AI together with Mayan at Trainline. If you go, Mayan.
- 1:58
Hi, everyone. My name is Mayank. I, um, I'm a senior AI engineer at Trainline. I have background, research background in LLMs, but w- my research is in the pre-big LLMs, so like if you've heard of BERT or the older LLMs.
- 2:14
Um, so we are continuing building a state-of-the-art, uh, agentic products at Trainline. If you've come here from abroad, I'm sure you must have used Trainline. If not, please do 'cause it's a, um, it's the state-of-the-art, uh, app for, um, buying tickets- [laughs] -and other things.
- 2:30
So, uh, yeah, welcome, and we're very excited to host you at this workshop. And, uh, yeah, we'll guide you through the, the hands-on, uh, experience.
- 2:38
Fantastic. I'm also joined by some of my colleagues as well from Braintrust coming over the pond. So Phil, Eric, and Rose, if you just show your hands. So don't worry, folks, you're in safe hands if you're ever stuck, so just holler, and they can help.
- 2:51
Fantastic. Uh, just wanna do a little bit of housekeeping as well. If everybody could join into the AI engineer Slack channel or, uh... Well, there's an AI engineer organization, but there's also a specific Slack channel which we'll be using today to help, uh, progress the workshop.
- 3:07
So if you are stuck, we can use the Slack channel to help each other out. Again, we're already on there, but as we progress and get to the hands-on element, and if you are stuck, um, this, this will really help.
- 3:17
We're also providing a, uh, cheat sheet, so if you are encounter any particular hurdle, there's a step-by-step instructions which help you just get to where we need to go to the workshop without, you know, having to, to feel like you- you're, you're falling behind.
- 3:29
Again, a lot of the assets which we share today, uh, you know, is publicly available. Um, and again, we can have any particular follow-ups, um, i- if needed. But we'll give us, uh, a few minutes to make sure you get onto that, and then join the channel AI Engineer Europe 2026 Braintrust dash workshop.
- 3:44
Uh, it's a public-facing channel. All right, team. Um,
- 3:52
guess we can start. Um, yep, just for the people at front and the back,
- 3:58
um, we'll need to you to join the Slack channel. So if you've seen Pierce, then we're good. So, uh, team, thank you, everybody. Um, let's proceed. Okay. So just to kind of help orient today's, uh, workshop, we'll be breaking into kind of [REDACTED:account_number] main sections.
- 4:14
We'll spend a little bit of time just setting the bit of establishing the background, why we're here, and, and why this w- workshop is relevant. Hopefully, it'll set the context for when we go into the workshop, uh, building, uh, the system, talking about this, you know, how do we ship, uh, AI quality.
- 4:29
And then we'll wrap up for, uh, the key takeaways, and we'll be around to answer any kind of, uh, open, uh, questions and answers that we have from the field.
- 4:39
Okay. So hopefully today, um, there's gonna be a lot of people from very different backgrounds. Um, we intended this workshop to be catered for... Again, probably everyone here knows what an LLM is.
- 4:52
I don't want to insult anybody. But hopefully you're starting to, uh, explore your journey in terms of your, um, mature- maturing your, your operations in terms of building AI systems.
- 5:01
So whether you're an AI product engineer, probably an applied team or coming from a traditional machine learning folks, maybe you might be a platform operation infrastructure, um, this is really appropriate for you.
- 5:12
Just kind of a show of hands here, like, who here comes from a, let's say, a traditional data science background, perhaps? Okay. Interesting. Interesting. Who here perhaps comes from maybe software engineering or then pivoting into AI?
- 5:24
Okay. So this is definitely the right room for you. Uh, so hopefully we'll, we'll be able to kind of accelerate things, uh, as we progress. Okay. Um, to set the con- context here, um, I think this is not an uncommon expression, which we are seeing more broadly across the industry.
- 5:40
Um, again, bit of show of hands. Who here has done a, a machine learning or AI POC, let's say, let's say more specifically, a generative AI POC, but then has failed to kind of take that into production?
- 5:51
Okay. Just got a few hands. Yeah. I'd be worried if, uh, [laughs] everyone didn't, but this is kind of a, a key thing which, as, you know, speaking to many of my customers, you know, executives, top-level folks- All the rage.
- 6:02
You know, they're thinking about this new technology. Um, it's not necessarily new, but, uh, it's newer for a lot of folks, especially in kind of more enterprise and regulated industries.
- 6:11
And then they're trying to take this to, uh, delivering value to their customers. But unfortunately, there's a big hurdle between, um, taking what you might develop locally on your machine and then industrializing it and making sure it works in anger.
- 6:24
But the key thing that we've seen from all of the research out there, it's not that the models aren't particularly smart. Like, we've got very sophisticated models, whether you're building something in-house or you're using kind of a top-of-the-shelf, you know, top, uh, commercial LLM providers out there.
- 6:37
What we do see more broadly speaking is the type of operational rigor when it comes to delivering these systems to scale has not kind of kept up. Because traditional software engineering, very deterministic, one plus one equals two, great.
- 6:49
LLM systems, you know, as my [REDACTED:age], or sorry, [REDACTED:age] would say, um, you know, "Two plus two equals ten, Daddy." So it's having to kind of adjust it and make sure that we're delivering that to scale.
- 7:00
And again, holistically, what we're seeing, uh, when shipping these things is the fact that, again, people think a demo state is suitable. Um, but clearly doing two, two to [REDACTED:account_number] to five demos, great.
- 7:11
Putting in production, everything goes awry. Um, treating some of your logs as observability, and then this is really critical to, to how we look at, at Braintrust, is, you know, logs will tell you what has happened, but sometimes you need to go deep into the system and under-understand its behavior.
- 7:26
And this is really where observability comes into play. Um, something as well is, you know, works on my machine, fails in production. I try to patch the prompt. Um, and then again, it's operational until the next issue happens or the next failure mode.
- 7:39
But again, how do we keep track of that? Um, if, especially if you don't have a system in place, um, irrespective of what tooling you use, um, this can, uh, you know, has a categorical effect.
- 7:51
So again, a lot of what we see, again, is not to do with the tooling technology. It's down to more operational workflows, and this is really what we're aiming to help you in today's workshop.
- 8:00
Okay. So again, as I mentioned, it's not the prototype. It's getting to a state where we're knowing exactly what's changed in the system, how do we interrupt with that, and then how do we systematically put a set of rigor so that we can get better and better.
- 8:13
Remember, a hundred perc- uh, our target is not a hundred percent coverage. It's getting as close as possible, uh, while maintaining, um, fixing the gaps that might have existed.
- 8:23
Um, and again, something that we see time and time again is, again, one prompt might work, but as you move and industrialize it, you wanna probably do things like breaking this down into each individual, um, sectors of responsibility.
- 8:35
So again, if you do come from the software engineering background, you know about these. You're breaking down the monolith into microservices. We'll outline a very similar approach here when we talk about building these systems.
- 8:46
Um, again, making sure that we're understanding these changes, uh, putting a set of systems in place. This is really what we wanna do in today's workshop. Okay. In terms of today, um, hopefully, as I mentioned, this is a hands-on workshop, so we're gonna be going into the terminal, we're gonna be going into the UI, and we're gonna
- 9:03
be going through the step-by-step and guiding you along the way. So we'll be doing a staged AI system, uh, with, uh, a multi-stage, uh, tool calling, which really allows us to, um, see this more, uh, let's say, agentic flow.
- 9:18
We, we're gonna use Braintrust to instrument and see how the performance of the application's working. We then also want to then, you know, take a look at creating what we call, um, or identifying failure modes using, uh, a golden set.
- 9:30
So we'll push that through. We also then wanna talk about how do we industrialize this. So then moving from, "Hey, it just works on my machine," to something that you can use in production and have it managed forward with a system in place.
- 9:40
And key thing is identifying, you know, those, those edge cases, because again, you can create a d- a test dataset, but ultimately, um, there's no substitute for real-world data.
- 9:48
So again, we'll be able to show you how we, we, we take those real signals in and evaluate and complete the loop.
- 9:55
Um, so just a bit of introduction here today as well. Who here has heard of Braintrust or played around with Braintrust? Can I get a show of hands? Okay, great.
- 10:03
Fantastic. So, okay, love this. Love getting on in. So just a bit of introduction to Braintrust. We're, we're a company now that's, uh, I think just shy of [REDACTED:account_number] years old.
- 10:13
Uh, we're approximately a s- um, a Series B company. So we just announced that a few months ago we raised eighty million dollars at a eight hundred million dollar valuation.
- 10:21
Um, we have investors such as ICONIQ, a16z, uh, as well as Greylock, uh, to really talk about, you know, helping organizations ship quality AI at scale. So we're the platform for AI observability.
- 10:36
Uh, we've got a heavy u-user base, uh, globally, but we're also expanding our presence heavily in Europe. So again, I'm one of the first engineers to, to join and help build out, uh, our go-to-market function here.
- 10:46
And we're, we're very excited, again, with our customers and our friends over at Trainline to do that. Some of our other local customers include, like, Lovable, Doctor Lib as well, uh, which are really pushing the forefront of these AI systems.
- 10:59
Uh, again, when it comes to using kind of Braintrust, again, I know to-- Again, I wanna kind of get hands-on with it, but where we really distinguish ourself is being able to do this at scale.
- 11:08
So our founder, Ankur Goyal, um, this is actually his third time building Braintrust. Um, so he's an expert when it comes to sort of database systems, and he's built, um, a, a, a company called Impira previously that was acquired by Figma that talks about document extraction.
- 11:24
And he's led the ML, uh, machine learning team out there. And because he realized, you know, building these, these evaluations are hard, understanding production traces are hard. So if we're having this issue, I'm sure there are other organizations out there which are doing the same.
- 11:38
And so he founded Braintrust to really help do this at scale. And as we're doing and understanding this, these, these traces which are coming in and, again, being very highly semi-structured data which changes, uh, he realized that traditional analytical systems weren't fit for purpose.
- 11:53
So we've put-created a, um, a new category of databases called Brainstorm, which really helps identify and accelerate this at scale. With us, again, we're tool work, we're platform agnostic, so irrespective of the agent framework you're using or, or the LLM providers out there, we're, we're, we're intending to help you deliver value, uh, irrespective of that.
- 12:15
Okay. One things I will talk about as we progress, uh, in this workshop is the, is the concept of a flywheel. So, uh, again, if you ever come from agile development, um, you know, perfection is the enemy of good.
- 12:28
We want to start somewhere. So even if the case that it's a new application and you don't know how it's going to be in production, we can start up with an evaluation set.
- 12:35
If you have an existing application to, to instrumenting, that's great. We can pull that information in and identify the failure modes here. So the key thing is get information into the system, identify those modes, remediate, ship it out, and then monitor, and complete this flywheel, uh, again and again till you get to where you need.
- 12:54
All right then. Um, you would have heard, uh, a lot from me. So one of the thing do is I'm going to provide, uh, my colleagues over at Trainline to maybe just share their experiences of, you know, prior to, to Braintrust and, and how they're helping out.
- 13:07
So, uh, Oussama, I'll give it to you.
- 13:10
Thank you very much.
- 13:11
No worries.
- 13:11
Uh, and hello again for people who are joining us, uh, just now. Um, as my, uh, mate, Mayan, uh, introduced Trainline, uh, we are a, a company that's actually helping people, uh, get on the trains.
- 13:25
Um, trains are different than planes, uh, if you don't know. Uh, there is this w- like worldwide system, central system for all planes around the world. It's not the case for trains.
- 13:35
Uh, and in Europe and the UK, it's very hard, I would say, if you would like to install an app for each carrier, it will take the whole space on, on your ph- mobile phone for sure.
- 13:44
So Trainline is actually being that platform, uh, basically, uh, to help you book tickets, uh, mobile app agnostic, platform agnostic, carrier agnostic. You do it on one app, and you can book a train from Paris to London, uh, from, from, uh, from Leon to, to, uh, to Milano whenever you like, basically, every carrier in the EU.
- 14:07
Uh, we sell like, uh, almo- almost like six point [REDACTED:account_number] billions, uh, of tickets on the trains, twenty-seven millions active users and counting. And the other interesting probably one, uh, for, uh, for this, uh, conference is how many, uh, AI conversations actually that we have from our-- with our travel assistant.
- 14:24
So we do have a travel assistant that is exposed to people, and it's not just a chatbot. Uh, it's actually a multi-agent system that actually can handle refunds for you, uh, can handle changing trains for you.
- 14:38
So it is very, I would say, a proactive, um, agent system, agentic system and not just a, a chatbot. Probably Mayan would like to add more about that.
- 14:47
Yeah. So one of the benefits of having twenty-seven active customers is you've got a huge space of how you can serve agentic applications live to the customers. One of the examples of that is what Oussama talked about is a travel assistant, which you can get to from a ticket window in the application.
- 15:04
Um, it's an agentic system, which is something we want to talk about, uh, a little bit later in the slide.
- 15:09
Yeah.
- 15:10
Go ahead.
- 15:12
Awesome. Which brings us, I would say, to the next, uh, point. Uh, I will keep it short. Uh, selling train tickets and being like a, a, a train, uh, tickets company, how come that we are doing, uh, machine learning?
- 15:24
Of course, we, we can. Uh, there's so much things, uh, to do and to help people, uh, in terms of their journeys, getting their tickets, getting their trains, getting back home, basically.
- 15:35
Uh, we build-- we do two things. So the classic ML part, which is actually building models, we do that. We build ML models inside of Trainline from scratch, from data to model.
- 15:46
This is something we do. And we also do the multi-agentic, um, agentic AI systems that we are now familiar with on top of LLMs that, uh, we, uh, love and cherish, all the toolings and our context engineering, all of that.
- 15:59
We do that. So we do these both sides of the story, uh, at Trainline.
- 16:04
And these are two examples of, uh, what users are actually, uh, using, uh, on top of those, um, systems. On the left side, so this is, uh, you can, you can think of it as your weather application, but for train disruptions.
- 16:20
So basically, you have a ticket for a train. You get there. We know if this train would be, uh, would be disrupted or not, if it will be, uh, pro-probably late or not.
- 16:30
We know, we know that, but based on huge data that we have on top of it that was a machine learning model that was trained and can actually, uh, predict, uh, train, uh, disruptions, uh, being late and all of that.
- 16:43
So this is the classic ML part. The other one, that's the travel assistant that I told you about. It is ve-- as I said, it's very, I would say, advanced, uh, multi-agent system.
- 16:53
It can show you alternative trains if your train is canceled or something wrong with it. Good luck doing that yourself, even with ChatGPT. And, uh, the, um, the other one is handle refunds.
- 17:04
So you can actually get for a refund, um, uh, on your ticket if your train is late or all of that, and it actually can give you all of that, uh, without a handover, and also can, uh, can do the handover to actual, uh, uh, human customer support, uh, in our lovely customer support team.
- 17:21
Um, which means that if you are doing this at production, in production level and that scale, at, uh, Trainline scale, it means that, uh, uh, or you can ask the question, are we breaking things at Trainline?
- 17:33
Uh, of course, we don't want to, to do that. Um, and we are moving fast because technology is definitely is moving fast. And, and this is why we are here.
- 17:42
We are, we are here to show you, like, what are we doing to move fast without breaking things, uh, in terms, of course, uh, of AI. And we can do that on top of handling the complex, uh, software systems that we, uh, we love and cherish, cherish from APIs at scale and, uh, serving millions of users.
- 18:02
Um, so how we would like to think about it is this scale. So we know for sure that whatever we have our software systems, this is deterministic side of the story that we have.
- 18:12
And on the other side, building the ML models, this is the non-deterministic side of the story. And we know for sure that the agentic systems are in between. There are parts of them that are deterministic.
- 18:24
There are parts of them that are not definitely deterministic. Uh, uh- And this is how we, uh, the framework of thinking that we have. In terms of quality, how are we handling that?
- 18:34
That's the question. For the ML models, we do, we do care about the quality of data and, and the, that we use for training the models. But also we have on top of it, of course, uh, ML machine learning evaluations, uh, whether offline or online.
- 18:50
Offline, it means like before going to production, you, you need to do your evaluations. And online is the data from production, you get it and evaluate your model if it's doing, uh, well or not.
- 19:00
In our case, for instance, for that for- weather forecast, uh, is it pre- predicting the state of the train, uh, if it's disrupted or not? Is it, is the model correct in its prediction?
- 19:10
So that's one, uh, one example. On the other hand, for people who are, who are familiar with software engineering, we do all [chuckles] I would say quality checks and the diagrams or whatever, and the tooling that we have for handling, uh, quality at scale, a very large scale at Trainline.
- 19:26
Uh, and for those systems, I think you already guessed it, it's basically a combination of both. It's not one without another, for sure. We do everything that we do quality-wise in terms of deterministic systems, but we also use the, the, the eval- evaluation, um, uh, side in, um, from the non-deterministic systems.
- 19:50
Uh, so yeah, that's I would say how the, I would say the framework of thinking that we have at Trainline, thinking about these systems and, um, definitely, uh, Braintrust, uh, is helping with that.
- 20:01
And by the way, disclaimer, we are not I would say, we are here just because we are convinced that it-- Braintrust works for us. We, we are not paid, so that's one hundred percent.
- 20:11
Uh, we, we are happy customers. We, we have been with Braintrust for a long time, uh, and we use it in different, um, some of, uh, Braintrust features, we actually use them here.
- 20:21
So for instance, for the ML AI evaluations, we do that, and we, we, we follow the scoring of travel assistant on many, many, many levels from, uh, from tone of voice to a-actual helpfulness, uh, when it comes to tickets.
- 20:36
And tickets are really, really, I would say, complex, uh, in terms of reasoning and what you should get, what you should not get. It depends on the, if the train is late or not, uh, the type of ticket.
- 20:46
Is it a return ticket or is advance? So many complex cases. Uh, but yeah, we do follow that, uh, evaluation side. Probably Mayan would like to add more about that.
- 20:54
Yeah. So can I get a, a raise of hand from people who have struggled with LLM costs, number of tokens, switching models, problems like this? Right. So yeah, it's nice to see.
- 21:06
So it was a problem with Trainline as well because we do this at scale, and the amount of OpenAI, Anthropic amount that maybe pays just like this are. So we have to keep switching models, which is like the best model for our use case, cheaper models, uh, to-to models with efficient tokens.
- 21:23
Now, anytime you wanna switch models, you wanna make sure that it is performing at least at the same level as your current model, right? Before Braintrust, we had no specific way of doing it because we didn't have scores set up.
- 21:36
We didn't have, you know, sorts of like the evaluation. So with the usage of Braintrust, what it has enabled us to do is to simulate how the performance of the lower model would look like.
- 21:46
And we've used Braintrust extensively to run offline evaluation and see what the effect is going to be, and also online evaluation to see that the intended effect that we observed in offline is, is as expected.
- 21:58
So that's, I think, one of the use cases. The other one is more generic, which is like it would have taken us like a lot long to evaluate a new feature shipping into travel assistant.
- 22:07
But with Braintrust, we have been able to make sure that we are assured that the new experience for the user is going to be good. So Braintrust has helped us a lot, uh, in shipping fast.
- 22:17
Cool. And also the other part is observability for sure. Uh, we do, we do use Braintrust for observability if this one is working. Yes.
- 22:25
Yeah.
- 22:25
Um, so yeah, this is a true example from, uh, from, uh, Braintrust. We are spoiling basically the workshop with whatever you are going to see there. But yeah, we can, we can track everything regarding tool calling, the other agent, um, the in, in terms of number, in terms of quality.
- 22:41
That, those kind of insights, you will need them to get from proof of concept to actual system in production or in your company or for something actually, um, uh, that's being used there out there.
- 22:54
Uh, and, uh, users find it, I would say, true working product and not, I would say, just any, uh, slope, uh, AI-generated system. No. We are, we need, we will need that for definitely for, for those cases.
- 23:07
Uh, would you like to add something here?
- 23:08
I was just gonna say that Braintrust enables you to look inside complex agentic workflows up to like tool call level, token level, which is like very insightful and helps you debug a lot of things in production and before you deploy it.
- 23:22
Yeah. And one last thing probably is the cross-functional, uh-
- 23:26
Yeah
- 23:26
... uh, friendly, uh, point. Uh, that's something that we have discovered, uh, along the way. Uh, because we are building those systems from travel assistant to the models, at some point, we need to have-- We are a big company, and we have a pro-- people from product, people with non-technical backgrounds, and we need to communicate, and we
- 23:43
need to share many things. And they also need to self-serve some of those things. We cannot, I would say, babysit any, uh, people and say, "Hey, you should do this.
- 23:52
Should... Let me get you the logs, download it, and send them that." That does not work. At scale, uh, we need like a way like where we can help, I would say, uh, we can work cross-functional way, um, and let people be, I would say, um, uh, uh, free to do whatever they want and self-serve with, with
- 24:10
data and insights. So this is what we have also discovered, uh, along the way. And Braintrust helps us, I would say, uh, with many requests as well that we had, but yeah, really appreciate it, uh, for sure.
- 24:21
And of course, more. We are using just a part of, uh, uh, that system. We have our own things. Uh, Braintrust, I think, uh, you are building even more, more, more toolings.
- 24:31
Yeah.
- 24:31
So, so definitely, uh, more to come. I think the workshop will definitely help you discover all those, uh, things. And, um, of course Have fun building, uh, during the workshop.
- 24:42
I'd say laptops out, and if you are interested, Trainline is hiring, uh, in the AI engineering side, of course. And, uh, give it to you, Giran.
- 24:50
Perfect. Well, thank you so much for that, team. Um, yeah. And just to say again, an impartial thank you, uh, and extension of my wife, we thank you so much for the split saver functionality. [laughs] [laughs]
- 25:00
So, uh, give your kudos to that. All right then, team.
- 25:05
Right, let's proceed with the, the setup. So again, this is going to be a hands-on workshop. Uh, better dust off your Bash skills, team. Jokes. Thus, I've got a nice raft of commands to help out with that.
- 25:15
So hopefully everybody, or at least most folks, have joined the Slack, uh, organization for AI Engineer, and then also joined the, the channel. I've put a link to this, uh, but there's also a QR code to the repository which we'll be using sort of today.
- 25:31
So I'll give it a few minutes. Um, re- really key things is, um, signing up to a free Braintrust, uh, account. Um, if you use Gmail, you can use a little plus sign trick to just create your account.
- 25:46
If you, um, you know, keep it... If you are an existing Braintrust, uh, user today, uh, so hopefully it doesn't pollute your, uh, existing one. You will also need access to an OpenAI API key.
- 25:58
Um, this should be easy for you to generate. If for some reason you're not able to generate a key, then please let my colleagues know. We'll be able to send it and DM you on, on Slack to use for this specific session today.
- 26:12
Um, and I think obviously as an engineering conference, and especially in AI, uh, if you are using AI coding assistance, feel free to use that within your ID or in the terminal to guide you along the way, ex- ask questions around the code base and so forth.
- 26:26
Um, I've tried to simplify a lot of the scaffolding in place, so I'm using Mize to, uh, manage, uh, Node and pnpm, um, on, on the machine. But alternatively, you can download a Node v22 in the specific version there.
- 26:43
Um, it should be fine. Uh, but again, I just want to caveat, especially for this workshop, I fixed that specific, uh, let's say, runtime. I'm using Make to, uh, just again, some bit of syntactic sugar to wrap the commands.
- 26:57
Um, if for a reason you don't have Make installed, if you're on a Windows machine, you can then just run the, um, pnpm commands, which are then just wrappers for the package.json file-
- 27:07
Hello
- 27:07
... here there today. So yeah, we'll give it a few minutes-
- 27:11
You mentioned the time
- 27:12
... uh, for folks to, um, to scan and clone. And just worth pointing out as well, um, I have a, uh... In the sheet today, I can see lots of folks have already, um, gone through it.
- 27:25
But, um, yeah, we provide a step-by-step logo. So as mentioned, as we're proceeding with the r- um, workshop today, if you are stuck, you can please ask questions on the Slack channel, but also you refer to this.
- 27:38
Um, as mentioned, this will stay. This is a public-facing asset. So, uh, even if you, uh, outside of this workshop, you're stuck, you can kind of go through it at your own, uh, leisure.
- 27:47
That's good. It wasn't muted. It wasn't muted.
- 27:54
Okay. Just kind of maybe a show of hands, um, who here, um, is not able to get an open API key?
- 28:12
OpenAI, sorry, API key. Hopefully, folks will be able to provision that. It's, uh, really critical to this.
- 28:22
Yeah.
- 28:23
There are some people who joined late who don't have the QR code.
- 28:26
For the Slack?
- 28:27
For the Slack, yeah.
- 28:28
I can do that. Yep. Okay, folks. I think we did have a few late starters, so, um, yeah, I don't want to spend too long with this. But if you did join late, please scan that QR code, uh, which will give you access to the AI Slack engineering, uh, organization, and then join the AI
- 28:53
channel, uh, AI Engineer Europe 2026 Braintrust Workshop. It's a public-facing channel. Through that channel, you'll be able to get access to the pinned, um, uh, content. So the repository as well as a cheat sheet, which we'll be following as well.
- 29:13
Just, yeah, just a bit of a sense check, folks. Um, hopefully, most of you have been able to join the Slack channel and are able to clone the repository, have access to that.
- 29:25
All right then. Um, uh, as I mentioned at the start of the, um, the agenda as well, what we'll be doing is creating a, um, an example, let's say support triage agent.
- 29:36
So, you know, hopefully, this kind of, uh, exemplifies, um, a lot of, uh, systems that folks maybe have already built or, or trying to build. Um, it is, it is a fictitious, uh, um, application, uh, designed specifically for this workshop.
- 29:49
Please do not use it in production. It's just really around to just teach us, um, you know, how, how do we build these AI patterns into scale. So the idea there is, um, given, uh, a ticket might be raised in the particular system, uh, we have a set of agent that goes through a, a pipeline process.
- 30:06
I wouldn't say pipeline, but a stage process with tool calling, uh, that produces, um, uh, a set of information that can be emitted into downstream systems, uh, from that perspective.
- 30:16
Okay. Um, just to kind of help visualize what, uh, the system does. Again, mentioning we're taking the ticket input. Uh, we have a first step, which is collecting the context.
- 30:25
It's quite a deterministic way of extracting information.
- 30:29
We then proceed with the agentic type of portion where we've broken it out into [REDACTED:account_number] stages. So we have an LLM and tool calls to, to triage that particular issue.
- 30:39
Um, we then want to do a policy review to make sure the output is, is correct. We then wanna create and draft, let's say, a customer-facing reply with Reply Writer.
- 30:50
And then we wanna package this together. And de-depending on, um, how severe this particular ticket, we can invoke another tool call to whether we need to escalate, uh, this to, let's say, a human in the loop, uh, and then drop the final result.
- 31:04
So, uh, again, not too, too complex, but just give you an idea of, of the, the system that we'll be trying to build and, and operationalize, uh, in this session.
- 31:13
One thing to point out is, you know, as we begin to use Braintrust, it's like w-really where does it fit into helping you, you know, deploy and, and manage this at scale?
- 31:22
I've just isolated, uh, this here where, you know, everything we're gonna be doing towards the end, especially as we progress, we'll trace it end to end, um, as my colleagues have talked about.
- 31:32
Um, we'll use managed mode for prompts or the offloading that what you'll do locally into a secure environment. Tool calls as well will be managed. So again, uh, how do you-- you'll talk about getting to external systems.
- 31:45
And lastly, around evaluations and scores, which will be run through, uh, our Braintrust infrastructure.
- 31:53
Um, it's worth pointing out as well, if you do go into the GitHub p-- uh, repository, uh, there's a page-- GitHub Pages version of this as well. So again, the slides will be readily available for you to use and consume.
- 32:07
Okay. Key thing as well, uh, with this is I've tried to help you, um, do each phase a-and checkpoint this. So if you are ever stuck, the idea is just use Git, Git checkout to a specific tag or the branch.
- 32:21
But in this case, each individual tag or branch is fully runnable at that stage. So if you're ever stuck again, check-- Git checkout to that branch. So you make install, make set up, run the commands, and then you should be able to get, uh, near identical output every, every, every time.
- 32:37
So again, this is all documented in the, the README, but it's also included in the cheat sheet as well, uh, for you to progress.
- 32:44
Okay. So yeah, just to kind of give you the sequence of events, uh, we'll be doing kind of the scaffold and, and setup. Uh, I'll talk about building a s-- basic agent.
- 32:53
Again, this workshop isn't really designed to talk about building agents. It's about, okay, once you do have an agent, how do you then operationalize that? So, but just for brevity, we wanted to kind of do it step by step.
- 33:04
And then we'll kind of get into the, the core of it, where we'll add the tracing, talk about the evaluations, you know, talking about that golden set and identifying where we can, we can improve.
- 33:14
And then again, tying it all together and using kind of managed infrastructure within Braintrust to, to help operationalize this out. And again, talk about that collaboration that my colleagues at Trainline talked about.
- 33:24
Also, kind of the key thing, which is like once you do identify a production failure, it's then how do we apply a fix to this and then complete that flywheel and give you a finalized asset.
- 33:34
Okay. So in step one, we're gonna be talking about building the agent. So in this checkpoint here, um, again, we need to start somewhere, right? So what do we do?
- 33:45
Uh, we will take an initial call to an LLM. We'll set up a prompt, one shot, one in, one out, and get an output. Uh, so again, if you're doing a proof of concept, um, this might be great, uh, as part of that initial spike.
- 34:00
But again, we know there's, there's work to be done. But yeah, for context, we're just gonna be stepping through this. But as I mentioned, just because it works in the demo doesn't mean it's necessarily gonna work in production.
- 34:10
Um, I've put some pseudocode here just to articulate here, but you can see here as, as, as with all, uh, these, um, you know, language models, we provide a, a set of system prompt.
- 34:21
We provide the user and especially the text, and we pass that output back into our application. So one function, one model call, uh, we have the structured output, which we wanna receive as part of the, the output.
- 34:33
Okay. So what we'll do as well, uh, when we do the, the scaffolding and checking out to the first, uh, branch, uh, we can run a set of prebuilt tickets.
- 34:42
So it's, uh, it's a case that's kind of already done in code, or you can use the command line. So I'll just demonstrate that shortly, what it would look like, uh, with the outputs.
- 34:51
And you'll get something similar to that in, uh, JSON format. Okay.
- 35:01
So hopefully everyone can see this at the moment. Um, I have the, the application here.
- 35:10
Okay. So I'm just going to go to basic checkout.
- 35:16
Git. Okay. And, uh, what's this? Let's move it. So if I take a look here, so I'm now checked out into, uh, that first tag, so building the basic agent.
- 35:32
Um, I've got the application here, which is again, very similar to the, the prompt code. Um, we're creating the client, uh, calling OpenAI SDK. Again, for brevity, I've, I've just-- I didn't use any agent, agent SDKs, but we fully support that as well, uh, as we proceed with the, the workshop.
- 35:52
Um, so this is readily available, and you'll see again from make file. Uh, if you don't have make installed, then you can just execute, uh, pnpm scripts or, uh, I don't know if you, if you use npm or Yarn, that's also possible as well.
- 36:06
I haven't tested it, but theoretically it should work 'cause it's just using a package.json.
- 36:11
So in this case, let's say you want to do something like, you know, make ticket.
- 36:17
So I want to give it things here, you know, my password
- 36:21
ne-needs to be reset. Um, in this case, I've provided some defaults, so I'm just like enter, enter, enter. And now it's just making a call to OpenAI. I'm just using, um, you probably would have seen from the environment variable file, I'm using GPT-5 mini.
- 36:37
You can switch it if you want to, but, uh, just for the purpose of this, we'll keep it simple. Um, so you can see here I've got the ticket and it's provided some, um, some output here as well.
- 36:46
But again, that's a singular shot. Yeah. I think we've... One moment.
- 36:53
Okay. So, uh, you know, based on this output, you know, it looks fairly plausible. Um, but again, it's not gonna account for a lot of the edge cases that we want, especially if you've got a lot of nuance to the organization which we're trying to build into, uh, the logic here.
- 37:20
Uh, I can even do things like make, uh, demo.
- 37:26
So make demo for the scripts. Um, yeah, it's the same thing. I'm just calling the same function, but I've got it codified as JSON here.
- 37:44
So yeah, the JSON fields are available to see.
- 37:58
Awesome. Okay. So think it's fairly straightforward. Uh, I don't wanna dwell on that too much. So the next thing I want to do then is talk about, um, adding in kind of local tools.
- 38:13
So we probably wanna say, "Look, let's try to make this a bit more deterministic." Even though our prompt might be very well structured, I may wanna bring in, uh, different ways of how it might operate.
- 38:23
So, um, in this case, uh, I'm calling [REDACTED:account_number] different tools to look at relevant, um, help desk articles around that. Uh, this could be both internal and external. I may wanna look at certain things that have happened to the account.
- 38:36
So let's say a customer might have done a certain migration, and that might have an impact on backend systems. And there's more-- probably a reason why I wanna, you know, create an escalation here.
- 38:44
In, in the case, what I've done is I've made it a bit more deterministic. For the purposes of this workshop, again, this is kind of treated as code. But, uh, in reality, you are probably are gonna be interfacing with external systems like a vector search, MCP, CLI, and, and other types of, uh, uh, interfaces to, to build
- 39:01
out, uh, this capability. And again, key thing here is, uh, the more things that you add, the number of ways that it can fail will also increase. So again, this is why tracing, as we begin to go, will become more important.
- 39:17
Right. So just one here. So again, readme here. Next thing we'll do is go to add local tools.
- 39:28
So what it... Git check all. All right.
- 39:39
So in this case, yeah, the tools that I created,
- 39:44
uh, are then available here. But again, they just kind of checked in as code for simplicity's sake. And again, I can do the same thing where you can make a ticket,
- 39:58
say, "Password needs resetting. Account locked."
- 40:28
Okay. Now, so yeah, it's provided a little bit more information. You can see it's a little bit more verbose, uh, because we've introduced, uh, tool calls into this, and it's giving more context to the, the LLM.
- 40:50
Um, also worth pointing out, if you are feeling stuck, folks, that a lot of these work, uh, or the, the tags of the workshop branches are built sequentially. So if you-- let's go into, let's say, Git tag number six, it's gonna include everything that's part of that.
- 41:04
So don't feel like you have to go through each one. If you are feeling stuck and you wanna skip, you can do that as well.
- 41:17
Okay. Let's get on to tools. So again, we already-- I've already showed the code anyway, so, but this is an idea through pseudocode what it would look like.
- 41:31
Um, stages. So I think this is the next thing where, again, we're breaking down that monolithic call from LLM. We're now introducing tools, and the next thing is even further drilling down to special stages of how the LLM should behave.
- 41:42
So you would have seen from that sequence, uh, diagram we have done effectively five stages for this, where I'm now setting up things to collect the context, triage, determining if it's meeting brand policies, providing a customer-friendly reply, and also something internally for our systems, and then finalizing the result for downstream systems.
- 42:01
Um, again, it's just coming from, you know, traditional software engineering. You're breaking down your problem. You can see the-- exactly where something's going wrong in the stack, and you'll be able to remediate.
- 42:12
So again, where possible, try to be more explicit and break it down into challenges that you can work on.
- 42:21
So yeah, just a bit of pseudocode. Um, this is kind of like what it would look like. Um, probably gonna put a debugger, um, within your ID and, and take a look at that.
- 42:36
Okay. So I'm gonna be doing this as-- in, in part, but, uh, we've already started with a st-- um, starter point, and then I'm gonna talk about, you know, doing the, the specialist stage here.
- 42:47
So let's take a look at the next part of the readme.
- 42:54
So git checkout Special stages here. So if I take a look now, um, in the source folder, it's, it's got the individual, uh, functions which are, are, are pieced out.
- 43:09
Uh, the prompts associated with that, uh, is being loose. So let's say here, so triage from what we're using. And if I go down to the application, you can see, you know, uh, it's, it's down here.
- 43:22
So I'm using a, uh, asynchronous functions to, um, execute.
- 43:29
Yep. And similarly as well, if I do something like, you know, make ticket, maybe I'll do something a bit different. What's another classical problem? Oh, it's just like, I need to upgrade my plan from Pro to Enterprise, but the website
- 43:53
is not working. Getting five hundred errors. You know? Uh, so in this case, I'm in customer
- 44:04
tier two. Uh, we're talking about, um, billing here, and my account is actually account number [REDACTED:account_number]. So again, just to show you that we are-- this is live. It's not doing something that's hard-coded in the system.
- 44:41
Um, worth pointing out as well, this stage is slightly slow because we've broken down the individual one-shot LLM into sequential calls. So it is expected to take a little bit longer, but again, that's, that's part of building out this agentic flow.
- 44:55
And now you can see again, it's a bit more, um, verbose, but it's-- there's a lot more thought into, um, this, this agent here. So I guess you can see again, we, we talked about a billing issue.
- 45:07
Um, you can see that it, it believes that it's quite high because a customer wants to up- do that upgrade from a different tier. But there's obviously an impact to this from a revenue perspective.
- 45:16
So in this case, we should escalate to the appropriate, uh, people on our side. So you can see here what the escalation region should be, um, and, uh, both the internal and the, uh, the customer-facing reply as well.
- 45:32
So we've included a, a confidence score, um, to say, you know, if this is true issue. But again, um, as mentioned, um, tool calls will be able to pull information which is happening across different systems to provide a, a, a greater level of confidence.
- 45:51
Okay. Perfect. All right. With that in mind, um, so, you know, again, we've now hopefully shown how we can build an-- take an agent, break it down, build something that's multi-stage, introducing tool calling, uh, to give us where we need.
- 46:11
The next thing is then to provide that information and, and start tracing it, so we can actually see what's happening, you know, down to the individual details. And this is where, you know, observability piece comes into play.
- 46:23
So what we wanna do in this kind of section is break down the full execution path. We know that the stuff is very s-- uh, nested in structure, so there's tool calls, there's additional function calls behind it.
- 46:36
Um, we do wanna track some of the key things, and I think, um, you know, Mayan pointed out, you know, the early struggles that they had to talk about, you know, laten-latency, cost, uh, tokens count.
- 46:46
Especially f- time to first token is a very important metrics that we see many of our customers trying to identify. Um, also, you know, what were the inputs? What were the outputs?
- 46:56
Metadata associated with this. And then also including additional, uh, types of fields so that, again, when we talk about monitoring and observability, we can query it, uh, within the UI on the fly and set up alerting as n- if needed.
- 47:09
So again, having an output is not enough. We need to understand, uh, the full execution path, and that's what tracing allows us to do.
- 47:19
Um, yeah, just to give you an idea of the context, and I know it's a very contentious topic, topic to some of them, because again, I come from an, a, a, a background in full stack development.
- 47:29
So, you know, I split both coins in terms of both, you know, Python, Go, as well as TypeScript. But, um, yeah, we-- our SDKs is, is multilingual, so Ruby, Go, um, I think even .NET, we've got some folks who are using it, and, uh, that's, that's all good.
- 47:45
But yeah, I've just kept TypeScript for simplicity here today. But yes, our SDKs do cover a range of, um, different languages out there that you can start tracing the application with.
- 47:57
So yeah, uh, a real key thing to this and, uh, one of my challenges I do see with our customers broadly, is they will do an individual interaction, uh, against a, a singular ch-- uh, parent span.
- 48:10
And this even might work with, let's say, multi-turn, multi-conversational agents. What we wanna be able to do is trace that into a nested structure, so you can see everything, but in one interaction, let's say one conversation, uh, in a, in a parent span.
- 48:24
So it's really critical as we start instrumenting our application to make sure that we're getting the, the right structures in place. Otherwise, you're not, you're not gonna be able to see the full effect of where things might go wrong in the application.
- 48:38
Okay. Um, again, th-this will just come down to reading the trace. Hopefully, folks, um, especially as I know some of the-- you've got more software engineers in the room, uh, that are using kind of the traditional observability tools out there.
- 48:49
This is not too dissimilar. Uh, but for folks who may be new to this, uh, screen, a trace just allows us to, uh, not find out what has happened, but what's currently happening with your application Uh, in, in, in real time.
- 49:01
So, uh, again, tracking every single call, bringing that metadata. And again, depending where the failure mode is, we can then identify, okay, what might be the, a course of remediation that we need to take.
- 49:14
Okay, so in the next stage, what we're gonna do is add tracing to this. Um, and then we're actually gonna run it, and then go into Braintrust and hopefully we'll see it happening.
- 49:24
But one thing I do wanna bear in mind before I do that, so...
- 49:31
So let's go to our README. That's telling us to add tracing. [typing]
- 49:43
If I can spell. Okay, now we're great.
- 49:54
And I... I'm gonna keep it there. Okay, so I've introduced a, uh, helper script here called tracing. So, um,
- 50:04
what we do quite well, uh, within our SDK is, again, we wanna not introduce more complexity where it's needed. So if you are using, again, the standard, um, LLM provider's SDK, you can simply just wrap that function.
- 50:17
We provide that out of the box. But then when it comes down to the individual, uh, calls, again, we can set up, uh, some, some helper, uh, scripts here to help just kind of wrap up, uh, as needed.
- 50:28
If you are using something like Python, then you can also have a, a decorator function which helps as well. In this case, you know, I, I've got a nice little, um, function which helps with the parent and child.
- 50:40
And then when it comes down to the actual, um, application tracing, uh, you can see here, like I've got a, um, a child span which has been,
- 50:52
um, executed throughout this. Uh, again, I don't wanna do too many code diffs at this point because the co- code is, um, expanding pretty heavily. Uh, but yeah, just kind of wanna walk through the kind of key concepts for this.
- 51:09
Uh, okay. Um, before I run the, uh, application, I do wanna go over into the, the Braintrust UI. Sometimes it might create a, uh, if you've already signed up for a free trial account, so hopefully you would have done that earlier today.
- 51:21
It'll create like a, a test project. That's totally fine. We will create a new project as we execute this. Uh, a really key thing, uh, as well, um, if you haven't done already, top left-hand corner, you go to your profile.
- 51:35
Um, oh, okay, let me just create a run, so, okay.
- 51:40
Let's create that just for now. Uh, go to your profile, and then where it says, um, API keys,
- 51:49
um, you know, enter your, uh, a name for the key, generate that in, and then use that within your environment, uh, your, your .env file, or secure that, uh, in a key vault if you have that already.
- 52:02
Um, there's also the OpenAI key, which we're hoping you generated. Um, that should also be used, uh, in
- 52:11
s- AI providers. So I've set this at the organization level. You can also do this per project, um, you know, depending if you wanna segregate it by a particular team or environment.
- 52:22
Uh, that's also fully supported. Uh, but in this case, um, we'll need this key for later when it comes down to the managed and online scoring. So, but just bit of a tidbit, um, that key that you generated, make sure you put it into your, uh, Braintrust, um, organization here to be used later.
- 52:39
Okay. Right. With that in mind, I'm going to run the demo.
- 52:50
So here, so... And, uh, just point of reference, if you are ever stuck with some, you should just use make setup, make sure everything's in place. In this case, then we want to do kind of make demo,
- 53:04
and just we're gonna run and execute those tickets. I can even do it, uh, using the, uh, make ticket command as well.
- 53:17
Okay. So while that's running in the background... [microphone feedback] Oh, sorry. While that's running in the background, um, if you go back to, uh,
- 53:26
your application, you would see, uh, there's a project called helper workshop. So that's the one that's will be created as part of this, uh, workshop here today. Um, if you navigate to the Logs tab, you'll start to see this, uh, coming through in real time.
- 53:46
So worth pointing out, again, thanks to our capabilities of Brainstore, uh, we have near instantaneous write into our system and a read available shortly after. So especially it's, it's a non-blocking function, so a lot of our customers, uh, especially the more sophisticated ones, are really using Braintrust at scale and really pushing the envelope.
- 54:06
So it's not the case of, "Hey, I've got a thousand traces," you know. They're pushing, you know, tens of millions of traces at a, at a time across a very short period, and they wanna be able to, again, aggregate this.
- 54:16
And again, as you build more sophistication in your application, you're sending this out to more users, that's gonna grow up pretty quickly, and you need to be able to, um, have a system that can handle this at scale.
- 54:30
So once you start to see the logs that are coming in, um, come in reverse order. I'm just gonna take a look at the first one here, and we'll start to see that impla- instrumenta- application globally traced.
- 54:42
Um, there's a particular button here that allows you to view this in full screen, which I think is quite helpful. So, uh, as I mentioned, because we've got this nested structure in place where we're going through each, uh, in this sort of tree, um, this interaction at the top, um, is the, is the, is the demo ticket.
- 54:59
You can see at the very top level, we're saying how long it took to actually run that invocation. Uh- Prompt tokens in, out, cost and latency associated with that as well.
- 55:10
I've also included some metadata because I wanna be able to kind of extract and filter that out as needed, and I'll show you how that work. Um, and everything's available here to view.
- 55:22
Um, metadata is there. If I also wanna take a, a different look at this, we can even look at individual steps. So in this case, I wanna look at what's happened with, uh, triage specialist down to the actual, um, invocation to the LLM.
- 55:36
So again, SDK, uh, provides a lot of, um, flexibility around this, so you can see what was put in,
- 55:43
uh, the reasoning behind it, and the information that we set output. And then coming down to the last, you know, should we escalate or not?
- 55:53
Um, another views that we will see as part of this is, uh, taking a look at the timeline. So just gives you an idea, like a, a waterfall methodology to see, okay, was there a particular step which is taking longer?
- 56:03
Do we need to remediate? Uh, it's all possible to do.
- 56:08
Um, all right. So if I take a look at logs page, let's get a back out of that. Refresh. So I think it was kind of four tickets that were pushed.
- 56:21
Yeah, they should come through right now. Okay. Yeah. So that's-- it's two tickets at the moment. So again, that same information that we replayed is also available in the console.
- 56:35
Okay. Yep. So we've covered tracing. Let's talk about the evaluation portion for this. Okay. It's quite interesting, uh, again, depending where you are in your journey of building this application.
- 56:53
Uh, a lot of customers already build an application already, have it sort of monitored and pull that in. But what happens when you have, uh, effectively a cold start problem where you don't know what you're building, right?
- 57:03
So what's-- what does good look like? Um, and effectively, what does good, good enough ship to mean? Um, in the case of the support application that we're developing today, um,
- 57:16
kinda key non-negotiables. You know, have we categorized the support case? Have we made sure that there's no severity or low severity outcome for, uh, these issues which are blocking?
- 57:26
Um, d- does the escalation stay in policy with, with SLAs? Um, does the structure look, uh, look sound? And if we're making any particular changes, does it actually improve, uh, the way we look-- without actually breaking or having a regression, uh, to, to the application?
- 57:42
Um, we can do this using evaluation. So an evaluation, um, for those... Hopefully ma- many folks will know what they are, but for those folks in, in the room, think of it as a way you have your, your dataset, your input, you have a task, and then you have an outcome which you want to evaluate against or
- 57:59
scoring function. Uh, and this is kind of a little bit different, uh, comparing, say, traditional software development, uh, into working with AI systems because of its non-deterministic nature.
- 58:10
Okay. So in this particular, uh, portion, what I'm gonna be talking about here is
- 58:17
creating what we call a golden dataset. So in this support application, I've been kind of testing anecdotally, but I wanna create a, a set of edge cases where I think this is really gonna help give us a at least an initial level of confidence to the business that what we're releasing out into production, uh, you know, is,
- 58:35
is sort of f- for purpose. We al-- There's always room for improvement, but I wanna be functional. I, I, I wanna just think, "Look, it's not just me releasing this application based on vibes."
- 58:44
I have a concrete way of saying, "Okay, this is how it's, it's performed over time." Um, to do this, again, we use kind of two main types of scoring functions.
- 58:54
The first of which is deterministic, so I think a lot of folks may have already started using this. U- again, coming from traditional dev software engineering, uh, you know, unit tests, uh, I wouldn't say analogous, but they are quite similar in nature where, um, they're very easy to run, uh, cost-effective.
- 59:11
You're not actually using a model at this point. The secondary type, which is a little bit more sophisticated, which is then using an LLM as a judge, so another AI system.
- 59:20
And this is really helpful when it comes to systems where there's nuance which can't really be de- determined on determined-- deterministic systems alone. So again, creating one to talk about, you know, branding style.
- 59:32
Is this meeting, you know, customer satisfaction, uh, and so forth? So these-- Most important thing is if you cannot write it in a deterministic way, you wanna be using LLM as a judge where possible.
- 59:45
So... And why this is more important, again, it's just making sure that any change that we make is a safe to change, uh, as we progress this.
- 59:55
Okay. So I'm gonna pivot into, uh, the IDE again.
- 1:00:03
Let's take a look at the, uh, README file,
- 1:00:07
and then we're going to do... Well, looks the dictionary. Where you coming from?
- 1:00:16
Okay. Control C. Let's get checkout. There we go. Okay.
- 1:00:26
And then based on this here, so make demo. Oh, make setup. We should be fine. Okay. So one thing you-- I'd like you to do as well is then to run the seed dataset command.
- 1:00:39
So we do make seed_dataset. Okay. So what this does, it's gonna upload our evaluation, uh, test cases into, into the Braintrust UI.
- 1:00:55
So if I pivot into, uh, my datasets now, you'll see it's called helper seed dataset. Uh, again, just for simplicity, I've created ten, um, inputs. Um, I've also categorized them, um, around there so- Your input, what we expect, and some metadata associated with that.
- 1:01:16
So kind of the core structure to, uh, creating an evaluation.
- 1:01:27
So, um, as part of this as well, um, and the deterministic and non-deterministic, I've created some scoring functions, uh, which we'll be using as part of this. So it's available here.
- 1:01:45
Uh, and then, yep, scoring functions. So again, checking the category, uh, is a schema in place? Um,
- 1:01:55
is there an escalation reason when, when needed? Again, very easy to run, uh, and codify.
- 1:02:04
Uh, yeah, we have a bit of more sophistication with the, the customer rubric. So this is more of the LLM as a judge use case here.
- 1:02:13
All right. Um, then what we can do is then
- 1:02:18
go back to the README. So we don't need to do the demo. We've already pushed that through.
- 1:02:29
Uh, let's go ahead and then do, um,
- 1:02:34
make eval. Okay. So this is then going to run
- 1:02:40
an evaluation. Uh, run this out here. Um, the seed data as well. Uh, again, you don't necessarily have to put it into code, but, um, whether that's coming from a database or in this case, I've just used a flat file JSON for this, uh, just to give you some context.
- 1:03:02
So we are running this evaluation. And we should have, if we go to the
- 1:03:27
UI, the place for experiments. So I'll just double check in here. Yep. So that experiment ran against all the JSONs associated with it.
- 1:03:42
So we should-- you should get a, an output like this, at least in the terminal.
- 1:03:46
But if we go back to the UI,
- 1:03:48
uh, again, you can-- we, we're starting to track the-- our application against the inputs and outputs for this particular dataset.
- 1:04:01
Um, and quite similar to the, um, the tracing view you would have seen from the online, uh, traces, you get to see a very similar view here, uh, and seeing how it works across, um, your experimentation as well.
- 1:04:26
Yeah. And I think, yeah, pointing out, uh, as we progress through the workshop, uh, we'll start to see how do we then improve it using the, the set of difference functions and track that across the UI.
- 1:04:41
If you need a bit more real estate as well, um, you can collapse the menu, uh, which does help.
- 1:04:52
Okay. Time. Perfect. Okay. Um, all right. Let's proceed with the next sec- checkpoint, uh, around deploying and managing this.
- 1:05:12
So, um, a-as I mentioned, I think a lot of
- 1:05:17
folks, or at least anecdotally when I, when I've seen and spoken to customers, is again, things will work really well on my machine. Okay. Now I'm checking that into code.
- 1:05:26
I kind of want to take this into a, a place where, um, I can start to collaborate a little bit better. It's in a-- there's a point of reference.
- 1:05:34
Um, I've got versioning history. You can start to identify again, who's made what change. And you need a way to be able to bring these users together. So I think as Oussama talked about a, a collaboration.
- 1:05:46
Um, what's interesting as well is again, changing the prompt on your machine and then trying to ship that code to a repository, and then maybe somebody who's, let's say a non-technical SME, a product manager perhaps, they wanna update the prompts.
- 1:06:01
They can't do it. They have to tap you on the shoulder. M-maybe that's happened to some of the folks in the room here today. Uh, I know it's happened to me, uh, a few times before.
- 1:06:09
Um, and again, I understand it can get really frustrating. So we actually just want a way to be able to pull that together. A really key thing, especially for those who work in very regulated industries, like reproducibility is a very big k-key thing.
- 1:06:22
And I've worked in, you know, better part of over a decade in, in both banking and, and capital markets. I know obviously with, uh, you know, the regulatory out there, especially, uh, things like right to be forgotten, understanding, you know, who's made a change, especially in a stress exit scenario, how can we put this into a new
- 1:06:39
system? This is really key to, to helping unlock that. And interesting when it comes to know-- to identifying changes before you do that. So again, we just wanna... We don't wanna just be making changes, pushing out production and asking us what's happened.
- 1:06:55
We, we need to be able to kind of get that i-in place.
- 1:06:58
Um, and so for us, again, uh, we're introducing some capabilities now. So what you've been running at the moment, you've been running the tools, you've been running the prompts, everything on your local machine.
- 1:07:09
What we wanna do is then offload that capability into Braintrust. So when your application is running in a secure environment, it can refer to Braintrust to, uh, pull that information, uh, and then- Uh, help with the path of execution, which again, follows the tracing mechanisms that we have done.
- 1:07:25
So, um, by default, when you're running the make commands, um, the runtime mode was set to kind of local. If you want to use managed mode, just use the prefix managed and whatever the remaining make commands, uh, will happen there as well.
- 1:07:40
So what I wanna do then is, um, pivot into, um, the IDE.
- 1:07:45
So going back to make, uh, my README.
- 1:07:56
Okay. It's a little tricky. Let's see. Ooh, ooh, ooh, ooh, ooh.
- 1:08:09
Okay. So git checkout. Mm. Okay. So just take a look here.
- 1:08:23
We'll do make setup. Right. And the key thing I do wanna emphasize at this point, because we're managing this in, in Braintrust, use the setup,
- 1:08:34
um, Braintrust, um, command here. So what that's gonna do, it's gonna package up those scoring functions, those tools, those prompts, and push it onto the secured, uh, infrastructure.
- 1:08:48
So you should get an output like this
- 1:08:51
and what that would look like, uh, in the UI.
- 1:08:56
So I'll just go back to overview to see. So on the left-hand side, um, if we take a look at our prompts, you'll start to see the, the [REDACTED:account_number] prompts that we created as, as part of that, that workflow.
- 1:09:09
Um, give you an idea here. If we go to the, uh, the treehouse specialist.
- 1:09:14
So, um, uh, uh, you can define a slug. So let's say, uh, an immutable ID, which you can then refer to it in the code. Um, this can also be generated if needed.
- 1:09:26
Um, the prompt, um, is again, treated as code. Uh, we also use some interpolation there if you wanna parameterize this, and I'll show what that looks like shortly.
- 1:09:38
If I go to the scoring function, uh, so let's click on scores.
- 1:09:44
Uh, okay, no, it'll pop-- come up in a second. Um, take a look parameters. So I'm just gonna take a look here, and I've created a parameter specifically just to simplify things, changing the baseline model.
- 1:09:57
So maybe just kind of a show of hands, um, you know, these models get released so quickly. Has anybody had like a PM or let's say, uh, let's say a non-engineering SME say, "Hey, by the way, can I change, uh, the prompt-- Can I change the model or the prompt and see what that would look like?"
- 1:10:13
Has, has that kind of happened to you folks? I think we've got a few hands here. Great. Okay. So the good thing with this now, using the, the managed parameters, those non, uh, technical SMEs, they can come into Braintrust and change the prompt here, write a comment to say, "Look, um, let's use, um, a different model."
- 1:10:34
Let me use something like four mini, uh, and just say, "Testing a new model."
- 1:10:40
I'm going to save this version here and write the comment.
- 1:10:48
And what I'm gonna do as well, just to give you an ending claim, I'm gonna do is, you know, make set up Braintrust. So every time you change something, um, you can run this command, but I'm just doing it for brevity here.
- 1:11:01
Um, that's more to keep, um, this model in sync. So if I go to prompts,
- 1:11:07
uh, you know, uh, that's kind of the sync. But if I take a look at, um,
- 1:11:12
the, uh, sorry, the parameters in place. Yeah, it should be there. And so what I want to do is
- 1:11:24
if I do something like, um, uh... I think this is also runtime model.
- 1:11:55
Account unlock. So in this case, by running the ma-managed mode, I'm just changing the course of the execution, uh, not to run the model locally, but to follow the path o-of what Braintrust is, uh, set for the model.
- 1:12:18
Okay. And, uh, if everything goes well... So you can see it's loading just now.
- 1:12:32
You can see that I've changed the model here. So again, I didn't have to do any code changes. All I had to do was go into UI, change the model I wanted to or any other parameter, run that, and have that, uh, used as a, uh, establishing baseline for the evaluations.
- 1:12:56
Um, again, if you want to, uh, you can also run the, the demo script to push in the demo tickets. I'm just skipping that for, uh, the workshop today.
- 1:13:08
Um, I think I do see this with many of our different customers is saying, "Oh, you know," um, there can be not necessarily a cause for concern. It's saying, "Hey, by the way, Braintrust is now having access control to these certain parameters."
- 1:13:21
Um, what I would say is Braintrust is not really intended to replace that rigor. You probably still want to use things like version control systems anyway to track that.
- 1:13:30
What we're just saying is when it comes to operationalize it and making sure that other users are able to Work on a shared system, this is a recommended path that we would take, uh, to help out with that.
- 1:13:41
So again, you would probably still have to have your prompts, your tooling, your parameters in a kind of centralized way, but then provide automation in place for you to synchronize that and, and work.
- 1:13:50
And that's the best way we've seen customers take advantage of this.
- 1:13:58
Okay. Uh, this next portion we're gonna talk about online scoring. So now that we have those evaluations in, in place, um, what we're gonna do is then apply those, um, scoring to actual live production logs that are coming through, uh, in the application.
- 1:14:16
So, okay, it's great that we've done our test cases. We, we've got some level of confidence that it's working, but again, there's no substitute for production data, right? We all know this.
- 1:14:26
So what we're gonna be doing is, uh, creating-- uh, again, moving that logic, uh, into Braintrust and setting what we call automations that would then track and, uh, evaluate this as logs are coming in, in, in real time.
- 1:14:39
Uh, worth pointing out, when you start your journey, um, it's probably, especially if you're using, uh, LLM as a judge, you wanna start with a, let's say, a higher sampling rate.
- 1:14:49
So again, as logs are coming in, uh, where possible, you, you wanna make sure you identify a baseline. But again, there's a trade-off when these calls can be quite expensive, especially if you're using more sophisticated models and you need a higher ra-rate of reasoning.
- 1:15:04
At that point, when you do want to, uh, are happy with, um, the output, you really wanna reduce the, that sampling rate down to five to ten percent. So again, you, you managing your, your cost effectively.
- 1:15:16
Um, deterministic scores, again, they're cheap. Recommend running them all, all the time.
- 1:15:21
Sorry, what, what are automation rules here?
- 1:15:24
S-sorry, what was that question?
- 1:15:25
What was, what are automation rules here?
- 1:15:28
Yes. So I'll bring that up in the UI. I'll show you. Yeah. Okay. So if I go back into the, uh, IDE,
- 1:15:36
I'll go into README here. Um, ah, so that's why I missed out the manage tools.
- 1:15:44
So let's just say git checkout. Okay. There. And as I mentioned, they all build upon each other, so, uh, skipping this is totally fine.
- 1:15:58
So git checkout. All right. Now, if I do, um,
- 1:16:05
it'll make setup. Everything should be fine. Make, uh, setup
- 1:16:14
Braintrust. Um, so in this case, I'm actually taking the, the tools as well, um, for the production. Again, this is just to really help accelerate things where possible.
- 1:16:29
Um, coming down to, um... If I hit refresh, um...
- 1:16:37
Right. So our scorers. The scoring functions that you saw earlier, uh, in code, they're now managed in Braintrust. So you can see it's, it's available here. Um, the ones that I wanna call out is the, the triage.
- 1:16:50
So the-- I've got, uh, LLM, uh, as a judge, um, which has been applied.
- 1:16:57
So putting output, taking, uh, this. And there's an automation rule in place. So rule quality online.
- 1:17:08
Um, so what I've done here, again, I've automated the setup, uh, as, as code, so to, to bootstrap the project. But a-again, uh, what we're able to do here is to say, look, depending if it's an individual span, we can, uh, run the execution or the entire trace.
- 1:17:23
And this is why metadata is so important because we, we may only wanna trace maybe specific failures that might happen within the code. Uh, again, depending on the use case.
- 1:17:34
Um, my sampling rate, as I mentioned, is set to, to a hundred, but again, for more expensive calls, we wanna taper that as well. So this is what the automation, uh, or how we would do that in, in Braintrust day-to-day.
- 1:17:46
It's like packaging the, it's like packaging the scorers and-
- 1:17:51
No. So automations is more like the execution against, uh, incoming logs. So it's, um, with that.
- 1:18:01
Yeah. But I'm, in more, more realistically, I've applied automation to setting up this environment and scaffolding. Yeah. So that, that's-- just wanna delineate that here.
- 1:18:13
Sorry, there's a question here?
- 1:18:14
Yeah. What kind of things are you scoring if you don't have the related data?
- 1:18:18
Um, could you expand on that, please? Yeah.
- 1:18:20
So the automations are like LLM as a judge, right?
- 1:18:23
Yes, correct.
- 1:18:23
What are you having a judge for properties?
- 1:18:26
In-
- 1:18:27
If there's like no ground truth generally-
- 1:18:30
Okay
- 1:18:30
... like, how can you perform online scoring if you don't have ground truth?
- 1:18:33
Yeah. Well, we probably wanna take that as an edge case, push that into a dataset, identify it, and then move that back. So that would be the approach that I would take for this, um, if you don't have any ground truth already, right?
- 1:18:46
So this is why I said depending where you start, it's better to have some kind of data and then begin your flywheel around from that.
- 1:18:53
Oh. Well, so if you have some data, how do you apply that to the LLM as a judge automations?
- 1:18:59
Oh, that case. Um, we can probably put that into a dataset and then replay that through the playground. That's, that's gonna be the way we would do that. Yeah.
- 1:19:06
Yeah. All right. Um, so we checked out online scoring. Coming to the end of time. Okay. Now to the remediation portion. So hopefully seeing the delta or, again, why we, we're here today.
- 1:19:23
So hopefully this might help you, uh, with that particular question. So again, here's something that might happen as a, as a plausible input to our agentic system is, you know, customer, uh, user might say, "Hey, you know, this is an urgent, but our CFO can't export the invoices, um, before, uh, the board meeting."
- 1:19:43
The model says, "Look, hey, this looks, this looks okay for me." Someone says, "Not urgent."
- 1:19:48
Comme ci, comme ça. But the business is very different, right? This probably does need, uh, immediate attention. Uh, your CFO, I'm sure there's an end of quarter report that needs to be done.
- 1:20:00
And this is the difference between, uh, what we're doing here today, is trying to identify, um, what is a proper failure mode and then remediate that where possible. Um, so in this case, again, I can run, um, um, this particular mode here.
- 1:20:14
I've got this in, in this, in this dataset.
- 1:20:18
So, uh, just to kind of give you a, a play of scene where we wanna replay the failure. We wanna set up a suc- uh, evaluation against this. We wanna tighten the prompt, run it again, and see, uh, what that looks like.
- 1:20:30
We probably wanna do suc- against not just one particular test case, but then across our entire test cases as well, just to see if, if it does work as intended and we haven't regressed on something else, uh, from that perspective.
- 1:20:45
Okay. So let's go ahead and have two... We have two separate branches for this. So I'll split out branches A. One's gonna have the, the failure rates in mind.
- 1:20:54
We're gonna go through the UI and view that. Um, and then we're gonna talk about the, the remediation path, uh, there as well.
- 1:21:01
Okay. Lastly. [whispering]
- 1:21:21
Should just come here. So many products for this.
- 1:21:37
Let's see. [keyboard clacking] Okay. So, um, make setup. [keyboard clacking]
- 1:21:57
Okay. So we can even do runtime one.
- 1:22:17
So runtime mode managed. And we'll make, uh, replay failure.
- 1:22:33
Okay. So I've got the set of five cases, which are, you know, the regression of failure modes, in this case.
- 1:22:42
So that first one that you saw in the example ticket, you know, that's, that's done as a JSON file here.
- 1:22:55
Mm-hmm. So if I take a look here...
- 1:23:16
Okay, so if I take a look at the, the failure mode here,
- 1:23:20
I can drill into this. See the, uh... And you'll notice as well, now that we set up the online, uh, the managed tools as well as the online scoring, the trace becomes, um, even more sophisticated in the fact that we're executing this against the secure, uh, Braintrust, uh, environment.
- 1:23:41
So again, moving from local to, to managed.
- 1:24:04
All right. I'm just gonna go here to, uh, the make file.
- 1:24:21
Sorry. Package.json. Let's go to, uh... Here goes
- 1:24:43
scenario. To manage scenario here.
- 1:25:01
Com. Loading high impact. Let's run an evaluation here.
- 1:25:15
So I'm just doing an evaluation against a specific, um, scenario.
- 1:25:21
All right.
- 1:25:53
Monitor some latency or things like that
- 1:25:56
Yeah. We'll do that.
- 1:25:57
Just
- 1:25:58
Okay. Okay. Um, yep. So as we can see, as we're progressing with, uh, the application, the experiment, a-again, Viewer just allows us to see the progress of our changes.
- 1:26:17
Um, so you can see we've run the, the latest set. We noticed some, you know, degradation here, which just allows us to track it. Open that up. So the managed data set that I, I had in mind, uh, the fail-failure rates are being captured.
- 1:26:30
And you can see, you know, I have the ability then to compare it against existing, um, experiments to see-- to track the progress and remediate where possible.
- 1:26:48
Okay. So forgive me with the next thing, is I'm going to go to the README
- 1:26:54
and then proceed with the, uh, remediation. Okay.
- 1:27:04
So do, "Git checkout initiation." Make sure. Make.
- 1:27:21
Trust. Okay. So in this remediation, I've changed, um, the, the prompt which I've used.
- 1:27:40
Uh, so, uh, one way to, to view that is if I take a look at the, the prompt and the, uh, the change.
- 1:27:50
So let's see. See any, any differences there.
- 1:28:12
Some place. Okay. Come down to the, uh,
- 1:28:22
README file here. Let's say I do... Do wanna say...
- 1:28:36
Say collect zero. Let's give this a say.
- 1:28:48
Run the, run the Braintrust command using contours.
- 1:29:08
Yeah, there's a specific flag, I think. I don't know if you wanna just double-check with that.
- 1:29:25
Yeah. That's it if, if it fails. So if you wanna just put that if, if exists,
- 1:29:50
replace, um, in, in the chat. Yeah. Yeah.
- 1:29:56
Do you want me to just do this?
- 1:29:59
Huh?
- 1:29:59
Do you want me to just do this?
- 1:30:00
Yeah. No, no, it's fine. So yeah, just folks, uh, if you wanna... In the part of the remediation script, if you set the environment variable Braintrust_if_exists, you set that to replace, um, it's gonna push in the, the updated changes into the, into the environment.
- 1:30:28
So then for take... Oh, sorry for that.
- 1:30:39
Okay. Um, and then you'll see as well, now that we've updated the prompt for your code and pushed that up, um, so it's available here as well.
- 1:30:57
Um, so as well in the UI, um, again, part of the operation, operationalization, uh, we can see, you know, who's changed what, but actually what has been changed to that particular prompt, uh, as in this, in this case.
- 1:31:13
Take a look here. So including tool facts.
- 1:31:16
Print that out. Okay. So coming back to here,
- 1:31:23
let's say we didn't do a, um... Run this evaluation again.
- 1:31:35
So we made the change to the prompt, and now we're running, um, the remediated version to see how that performs.
- 1:32:05
Perfect. Okay, so that's run the experiment here with the new changes, hopefully.
- 1:32:11
And I'll pivot into the, the UI and we go to experiments here.
- 1:32:16
And it's barely there, but you can just see, you know, kind of taper back up. Just, you know, any improvement going up is an improvement. But yeah, just to give you an idea here that, um, you know, now that we've done the evaluation, uh, actually what I can do is, um, do a diff.
- 1:32:36
Uh, and we can do-- so do a, a comparison in the delta. Yeah.
- 1:32:51
Yeah. So yeah, that's come up, uh, and improved over time,
- 1:32:55
um, uh, which were, which were, uh, the intended outcome. Yeah.
- 1:33:03
Okay. So, I think we're approaching the end of the, the content. Um, so I know it's, uh, been a number of steps, so I really wanna thank you for your time and attention to kind of walk through that.
- 1:33:20
Uh, as mentioned, the, the artifacts are public. We've got the cheat sheet there. We've got the Slack channel to help if you have any, any questions. But, uh, hopefully just to give you a summary of what you've, uh, accomplished today, uh, in, in this order, is you went from taking that single short prompt into building a five-stage,
- 1:33:37
uh, AI agentic workflow using, uh, tool calls. What we're able to then do is then inspect how this works, pretty much end to end, diving into this by adding Braintrust tracing, making sure everything is recorded.
- 1:33:50
Um, we also then wanna talk about, you know, how do we then evaluate the system from a, um, you know, when it's not online, when it's something new. Uh, creating those, those, uh, effectively a golden set with those test cases which we wanna execute against.
- 1:34:03
We then deployed, uh, that, uh, those managed prompts, those tools and parameters into the Braintrust secure architecture, uh, to be able to use. And we've also added online scoring to then evaluate the system, uh, as it unfolds.
- 1:34:16
And then we picked a particular, uh, production failure. We looked at the trace. Uh, we modified the prompt in our, uh, case, and we saw the delta there. And running the evaluation again, we saw it, uh, achieve, uh, back up to, to where it needed to be.
- 1:34:31
And in this case, completing the full evaluation of, you know, building, observing it, deploying it, and taking note of that moving forward.
- 1:34:40
Um, so yeah, just to kind of, um, call it, um, and, and bring it home. So again, hopefully this is not uncommon, but again, what might work in production is not really gonna work in prototype.
- 1:34:50
Uh, we really need to break this down, identify the failure modes and, and move forward. And that's where, again, explicit stages become really, really important, right? More, um, again, this does introduce more, uh, areas of, of where things could go wrong, but it's easier to debug if that's the case.
- 1:35:07
Um, again, there's no substitute for diving into the code and tracking everything. So I would say it's, it's observability's table stakes at this point. If you've got a production AI application and you're not tracing it, you need to go back to the drawing board and get that done operational.
- 1:35:22
And hopefully, again, we can show you how to, to be able to do that, um, using Braintrust.
- 1:35:27
Um, again, no substitution for production logs, but better to start somewhere from nowhere. If you have an idea of what an issue might happen, these are your perfect ways to, to set up an evaluation.
- 1:35:38
So using those, uh, failure modes as your test cases. And as I mentioned, this is a continuous process, right? Nothing's ever done. If you've ever worked in agile development, constant feedback is, is important.
- 1:35:49
And again, we're bringing this, this operation model, but with, uh, a newer surface of, of, of operating.
- 1:35:57
Um, and yeah, just to hopefully bring it home to your teams here today. So, um, you know, my encouragement to you is to, if, if this is something of, of, of interest, is pick something that's already operational today.
- 1:36:09
It d- doesn't have to be the entire suite. Maybe start off with something that's maybe a bit more, um, uh, that, I guess, more critical that you really want to improve, um, the operational modes.
- 1:36:20
Add the tracing, collect your edge cases from that, that mode, um, build those determined scores, and then route everything back as possible. So again, the faster feedback loop that you have, uh, you have more insight and the more that you can, again, improve the overall, uh, delivery, uh, operations of, of your system.
- 1:36:40
Um, yeah. And just kind of call to action. So again, I know we've thrown a lot of content at you. It's, uh, we- we'll obviously try to get a bit of feedback.
- 1:36:48
We're trying to put some tables in place, so I appreciate everyone's kind of juggling everything to do this. But, um, ag- as mentioned, uh, uh, you have to start somewhere.
- 1:36:55
Let's just try to accelerate you. Uh, we have a list of documentation that's provided. Um, that's, uh-- You can also use our AI agent on the agent to, to search if you have any questions.
- 1:37:07
But we also have a cookbook available. So I tend to throw the cookbook directly into, you know, Cu- uh, Cursor, Codex or whatever, and say, "Here, based on this, take the SDK, uh, and start tracing my application."
- 1:37:19
It does it pretty effective. We even do have, uh, we've actually announced a CLI, um, for our, the Braintrust, um, application. So that allows you to even do things as auto instrumentation.
- 1:37:30
I'm just gonna plug my colleague Eric, who's do- doing some fantastic work, and please check out his booth [laughs] 'cause it's, uh, it's, it's amazing. Um, and again, if you-- this is something that's interested you and you wanna explore more, then, you know, please reach out to your account team at Braintrust.
- 1:37:44
Again, we're happy here to s- to support where- where possible. Uh, and if you are on Discord, you know, feel free to, to join. Happy to answer your questions there, uh, from that.
- 1:37:54
And, uh, yeah. Again, really wanna thank you for your time, attention, energy. I know it's a really sunny day, and I, I don't wanna keep people in here. I want you to get some fresh air.
- 1:38:01
But, uh, yeah, just on behalf of Braintrust, Trainline, and myself, like, thank you so much for your time and attention. It's, it's been an honor, and look forward to seeing you out there tracing and, and gaining value from delivering AI in production.
- 1:38:13
Thank you. [clapping] [outro jingle]