← All AI Engineer talks

AI Engineer World's Fair 2025

How to build world-class AI products — Sarah Sachs (Notion) and Carlos Esteban (Braintrust)

About this talk

In this hands-on workshop, Notion AI lead Sarah Sachs explains how observability, curated datasets, LLM-as-a-judge evaluations, and separate retrieval assessments support reliable AI products. Braintrust presenters then demonstrate evaluation scoring, Playgrounds and Experiments, OpenAI-based implementation, production logging, human review, and feedback loops that help technical and nontechnical teams iterate together.

Chapters

  1. 0:00Introductions: Notion AI, Braintrust, and workshop goals
  2. 1:43Notion AI: observability, enterprise search, and evaluation fundamentals
  3. 16:44Audience questions: LLM judges, datasets, and RAG evaluation
  4. 34:28Braintrust scoring, Playgrounds, Experiments, and implementation
  5. 1:02:54Hands-on activities, settings, production logs, and human review
  6. 1:43:33Workshop closing

Talk transcript

  1. 0:00

    [upbeat music] Wow, look at this turnout.

  2. 0:16

    It's kinda crazy how many people are excited to, to listen to Sarah and hear what Notion AI has been building. Uh, so just gonna do some quick introductions before handing it over to her.

  3. 0:26

    My name is Carlos Esteban. I'm a solutions engineer here at Braintrust. Uh, so we're gonna... Doug and I are gonna share a bit about Braintrust after Sarah presents. Um, just wanted to say hi.

  4. 0:38

    You'll see me talk a bit more [laughs] today. Uh, I've been at Braintrust for six weeks. Previously I was at an infra company. [laughs]

  5. 0:47

    Yeah, the, the six weeks may be funny, but also I'm the, the veteran on the team, so Doug here, it's his third week.

  6. 0:52

    Third week. [laughs]

  7. 0:53

    So we're, we're up here in front of you and gonna teach you about Braintrust and, and evals. Um, yeah, if you wanna go ahead.

  8. 0:59

    Yeah, sure. Uh, solutions engineer also, alongside Carlos. Like he mentioned, been here a full three weeks at this point. Actually, not even full. It's, like, my- [laughs] ... I don't know, 11th or ninth day.

  9. 1:10

    Um, but yeah, this is, this is incredibly exciting to, to be here. Obviously a lot of interest in, in Sarah's talk, uh, and hopefully we can teach you a little bit about Braintrust in the process.

  10. 1:20

    So yeah, just wanna present Sarah. She's the lead of Notion AI. Uh, they've been customers pretty much since the, the beginning of Braintrust. Definitely pioneered the platform. And yeah, super exciting to hear what she has to say about their journey.

  11. 1:32

    Awesome. Thanks. I've been using Braintrust for 10 months, so I might be the most experienced Braintrust user on the panel. [laughs] Just kidding. Um, no, I mean, I think, uh...

  12. 1:43

    I wanna save time for questions, too, before we dig into the workshop. I think at a higher level, um, what I tell the team and what I tell people who are also building is, like, all of the rigor and excellence that comes from building great AI products comes from observability and good evals.

  13. 1:58

    And that's how you scale an engineering team. That's how you build good product. And ultimately, we spend maybe 10% of our time prompting and 90% of our time looking at evals and iterating on our evals and looking at our usage in Braintrust.

  14. 2:12

    And that, I believe, is the right balance of work in order to know that you're not just shipping something that worked well in a demo with your VP, or that you'd got working, you know, um, on the Caltrain to work and finally it did what you wanted that one time, but actually worked consistently and for the, for

  15. 2:30

    the users you were curious about. So I'll, I'll speak a little high level, save some time for questions, and then I hope everyone, um, sees value in the workshop as well.

  16. 2:37

    Because, um, it's really a tool that is accessible to all different levels. Like, you know, I have tech leads on our team that, you know, themselves are true experts in Braintrust.

  17. 2:46

    Someone like me that might not be necessarily executing experiments, but I'm in the platform every day. And then a whole host of data labelers and specialists that work out of the platform as well, which we'll talk about.

  18. 2:56

    So, um, who here has used Notion? Okay, love it. Who here has used Notion AI? Okay. Some sales opportunities. [laughs]

  19. 3:09

    Um, so what is Notion AI? Oh, that's me. Um, I've been at Notion for about 10 months, as I said. Um, Notion, this idea, for those of you that didn't raise your hand, is a connected workspace.

  20. 3:21

    What are we connecting? We're everything from, um, workplace management, asynchronous work, documents, um, but we also connect to third-party tools like Slack and Jira and Google Drive. Um, we have 100 million users, um, over 100 million users now.

  21. 3:36

    And, uh, something that's really important to note about Notion AI is we offer a free trial of all of our AI products, almost all of them. Um, and what does that mean?

  22. 3:44

    It means that the scale that we build for has to support that scale. So if you looked at the balance of who raised their hands, right? Maybe not every user is an active Notion AI user, but the scale in which we support Notion AI, we wanna offer it.

  23. 3:57

    So a new feature, for instance, like concurrently might have far more users than paid enterprise plan users or people on our business plan. And I like to think we're known for our exceptional customer experience.

  24. 4:08

    Um, we're certainly a design-driven company, um, which is to say that we care deeply about our product, and there's a lot of polish that's associated with the brand. And polish is not often associated with gen AI experiences. [laughs]

  25. 4:20

    So how do you add that level of polish and care into what you're building, um, while still building at the speed and rate of acceleration that exists in the industry, right?

  26. 4:31

    Um, for instance, we're very proud that we partner with a lot of foundation model providers. We also fine-tune our own models. Anytime a new model is released, within, um, usually less than a day, we're able to give that to all of our users in production.

  27. 4:45

    Um, in order to move at that pacing, um, but still have a polished product, you need to have evals and you need to have integration with a, with a product like Braintrust.

  28. 4:54

    So here's, um, our latest launch. This was, uh, two weeks ago, a suite of projects that we just launched with AI. [jazzy music]

  29. 5:03

    Okay. No audio, but I'll give you a little voiceover. Um, so AI meeting notes is kind of our first...

  30. 5:16

    So we now have text to speech... Or speech to text, excuse me, along with transcription, AI-generated summaries, and fairly soon you'll be able to see this also interact with your task database, your whole workspace, workspace awareness when talking about action items.

  31. 5:32

    Um, w- we have an enterprise search product that when you're searching, can search across everything. I use this constantly, particularly in standups when they're talking about Slack threads that are, like, 50 messages deep that I stopped following at 10:00 PM, right?

  32. 5:45

    I can use Notion AI. And for much deeper searches, we have a deep research tool. Um, that deep research executes parallel searches, serves a lot of our fine tune agentic capabilities, um, and is kind of our first transition out of workflows into agents, where rather than having, like, a set list of, um, tasks or flows that your

  33. 6:04

    AI, um- Program is running through, uh, now we're giving it the reasoning capability to decide on different tools and, and spend longer on the work that it's doing. So that's Notion AI for work.

  34. 6:16

    Um, that's the latest suite that we built. Um, and I'll talk a little bit about what we've learned, um, as we've built, um, this latest suite and as well older generation Notion AI products.

  35. 6:27

    Let's see. [buzzer] No, we already watched that. Okay. So believe it or not, Notion AI actually came out before ChatGPT. Um, that's kind of its own story that none of you are here to hear about, but we had early access to the generative models and always believed that content generation was core to how Notion worked.

  36. 6:49

    So our first product launched right around the same time. It was called the AI Writer. It allowed you to generate just in line, you know, write a sentence about X, Y, Z, right?

  37. 6:59

    From there, we started building Autofill. Autofill is our AI kind of agent that lives in every database property. Notion allows you to have databases. Now all of a sudden the AI is not just acting on the page, but is acting across the database and is triggered frequently and is doing things like translating things so that every column

  38. 7:16

    is in a different language. This is where we started seeing a lot of usage that was more unpredicted from our users, and it's around the same time that we also started building, um, a core AI team.

  39. 7:27

    Then we built kind of a natural RAG solution. This was kind of the, um, uh, I should say like the first time that we had like a full data platform AI full collaboration, and this is where we started offering, for instance, Q&A to like free users.

  40. 7:43

    So we have to have embeddings for everyone. We have to think about multilingual workspaces, things like that.

  41. 7:49

    And then now we get to the era where we started working with Braintrust. So you saw you could search across all apps. We launched the idea to have attachments.

  42. 7:57

    Um, so a lot of people use Notion as file storage or they'll upload things to Notion. We can search over those attachments. And finally, the things that we just talked about.

  43. 8:04

    So one thing I wanna note is that we didn't start with what I just showed you, and I think that would've been quite naive. Obviously, with the technology we have today, that would've been quite easy, but we also knew that we didn't have the technology to build those things yet, and we worked really with where the models

  44. 8:18

    were most capable. Um, what makes it hard to evaluate? So number one, um, these are exceptionally large datasets. Notion's really lucky because we use Notion constantly when we work on Notion, and so we can generate a lot of training data or evaluation data just from our own dogfooding.

  45. 8:37

    Um, I think that's actually one of the unique advantages that allows us to be one of the more fast-paced enterprise AI solutions. Um, similarly, our human evaluators were just exceptionally overwhelmed.

  46. 8:46

    Like, how do they look at... I think we were, like they were working in Google Sheets. I think we hired them even before we onboarded Notion. We have data specialists, and they were looking at this like dump in Google Sheets and trying to figure out how to parse the prompt with like a Google Sheet formula, right, and

  47. 9:01

    figure out what did the user actually say? How do I play with a few shots? It's very, um, involved, and I think anyone that utilizes human labelers well, um, and it's also been proven in research, is that particularly when it comes to fine-tuning but also iteration, quality is much more important than quantity, um, in terms of the

  48. 9:18

    insights and things that you extract. Um, and so we definitely needed a scalable, efficient solution to support how we were looking at our data and also keeping track of user feedback, all of the thumbs downs that we were getting in the development usage.

  49. 9:31

    So this is kind of what our iteration cycle looks like. So let's say that we wanna decide on an improvement. So for instance, let's say that we wanted to launch a Jira connector in our universal search product, right?

  50. 9:44

    What does it mean to query on Jira? How do I have my AI even figure out what's a task, what's a sprint, um, how to query from the Jira workspace?

  51. 9:53

    Um, we then curate targeted datasets from that workspace. Um, you'll hear me keep mentioning these data specialists. Uh, for much smaller enterprises, that might be your PM. That might...

  52. 10:03

    I would highly encourage it to also be your engineer. It should be the people that are closest to the data. We're at a scale now where we have a specialty.

  53. 10:11

    That's kind of like an LLM trainer that's a mix of a PM and a data analyst and a data annotator. Those are our data specialists. But they are creating just from logs and using a product and prototype handcrafted datasets.

  54. 10:24

    Um, those can be 10 things, right? But I would actually... Let's pause. Make it 10 things first and make sure that it's formatted the way that you want. We have definitely done the bad thing of creating lots and lots and lots of dummy data, and it's not structured the way we want, and that's a real pain.

  55. 10:39

    So first is like making sure everything is structured and your data flywheel is set up successfully to get insights that you might want. Then we tie them to scoring functions.

  56. 10:49

    Um, scoring functions should come after you've looked at the data for a long time. There are a lot of out-of-the-box scoring functions. We don't use them frequently. Um, we tend to use things that are specific to the product.

  57. 11:00

    Um, I'll talk a bit about the LLM-as-a-judge process that we use. That's probably what we rely on the most. But there are also a lot of things that are heuristic-based.

  58. 11:07

    So for instance, if I have, um... If I wanna make sure that my reasoning model triggers querying from Jira appropriately, everything in that dataset, that tool call should have in its query must be Jira, right?

  59. 11:21

    And so we can also make deterministic functions. Um, similarly, we have a dataset of places where there are multiple languages happening, and we know what the output language should be.

  60. 11:31

    So for instance, Toyota's a large com- customer of ours. Toyota might be working in Japanese and English. Someone might be asking questions in Japanese, but the output might be like, "Write me a paragraph about X, Y, Z."

  61. 11:41

    They're asking in Japanese, and like the lucky job of the product [chuckles] team is to figure out what language they actually want the response to be, right? It's a hard problem to have.

  62. 11:50

    We'll have a dataset of those as a curated Braintrust dataset, so we can, before we ship anything, run that eval and make sure that we don't break the multilingual language context switching experience, right?

  63. 12:02

    So for certain experiences, we actually run the eval as kind of like an ad hoc CI- Um, experience. I think there are some companies that actually integrate NCI. We don't do that currently, but you could.

  64. 12:12

    Um, that's more just to do with how long it takes to run the evals and how CI's built at Notion. Everyone submitting code would run evals, and that'd be a lot.

  65. 12:20

    Um, we, we inspect the results, and then we keep working. Um, the other thing I'll say is it's not just engineers that are on Braintrust. Because of that, we often will have, like, our PMs in Braintrust working very closely.

  66. 12:33

    So for instance, for that research mode product, we found that people were trying to draft reports much more than they were just trying to do research, and we found that by looking through all of that thumbs down data that was propagated into Braintrust, um, from internal dev usage.

  67. 12:46

    That's something that we could escalate to our, um, PM, but we also want the product thinkers and the designers to be thinking about that as, like, a first-class citizen.

  68. 12:54

    You can think of this as your version of UXR, right? Um, is to actually see what the model's doing well and how people end up wanting the model to do different things for them.

  69. 13:04

    Um, and then in terms of that feedback by development, I want to talk a little bit about the LLM-as-a-judge system. There's kind of two components that I've seen, uh, pretty prominently in industry.

  70. 13:14

    One is LLM-as-a-judge in which you have one prompt which judges everything in your dataset. So for instance, uh, is this information concise? Um, is this information faithful? Those are very common types of LLM-as-a-judge prompts.

  71. 13:28

    There's another version of it, um, which is a little bit more laborious, um, from a creation standpoint, but I find to be far more insightful, which is for every single element in your dataset or trace or whatever it is that you're, um, evaluating, um, you have a particular prompt.

  72. 13:42

    So for instance, I might write a prompt that says, "This answer should be in Japanese. The bullets should be formatted this way, and it should answer XYZ and point to page A," right?

  73. 13:53

    This is because obviously Levenshtein distance isn't gonna do a very good job of capturing the expected value to what we want, but we actually know exactly what we want the output to look like, and we don't wanna have too conservative of a prompt that will always break.

  74. 14:05

    For instance, just having a golden piece of data saying, "This is what the output should look like." We actually just say, like, what are the rules for what we want?

  75. 14:12

    This works really well with search evals actually, um, because our index is always changing. Um, when we wanna reevaluate how search retrieval goes, we say the first result should be the most recent element in the list about the Q1 offsite, right?

  76. 14:24

    And maybe since we redid that-- maybe since we created that element in our dataset, there's a new document about the Q1 offsite that wasn't what was in our golden set.

  77. 14:32

    This allows your golden set to be much more up to date and current. Um, yeah, and so I highly suggest that. Um, and that's actually what our, um, data specialists run.

  78. 14:42

    And what's really nice about it, as I mentioned earlier, our ability to switch to new models, um, this can output a score that's far more reliable. Um, and from that score, we can quickly understand if a new model has any serious regressions.

  79. 14:55

    Um, and this allows us to, um, given, like, our modular infrastructure of different prompts for different tasks and different models for different prompts, we can really quickly say, "Okay, Nano came out."

  80. 15:06

    Um, Nano was a fairly effective model that's very fast, maybe lower reasoning, and very cheap. What are kind of the high-frequency use cases where we don't need advanced reasoning, but we do want it to be able to run quickly?

  81. 15:18

    Let's gather those and, with one button, run all of the evals on them, see which prompts it does or doesn't do well on, and then quickly just change to those prompts, you know.

  82. 15:27

    And, and that cycle can become very fast. Um, and that allows us to stay at the top of the frontier in a way that I think has really been advantageous for us as a customer and, more importantly, um, for our users.

  83. 15:38

    Um, and the outcomes have been crazy. Um, I don't think that we could exist without Braintrust or a similar type of software today. It is critical to our iteration flow, and it actually is our IP.

  84. 15:49

    So everything that we build, the IP comes from how we evaluate it and how we build it, and that comes from where we build on top of Braintrust. Obviously, we have prompts that are like a series of strings, and obviously we have code that navigates through them, but how we decide if those work well is just as

  85. 16:04

    critical in how we decide on model selection, and all of that lives inside of Braintrust. Um, similarly, our AI product quality has skyrocketed because now we have observability. Sixty percent of Notion Enterprise users aren't speaking English.

  86. 16:17

    I will tell you a hundred percent of Notion AI engineers are speaking English. [laughs] So how do we build something that works for a majority non-English speakers? We have to use rigorous evaluation metrics that understand that multilingual experience.

  87. 16:31

    I think that was the last slide I prepared. I wanna save some time to answer any questions that you guys might have. Um, yeah, I think there's a mic if you wanna-- or you can just scream it, and I'll repeat it.

  88. 16:41

    That might be faster. Uh, you had a question?

  89. 16:44

    So for the LLM-as-a-judge, so what you're saying is actually you guys have multiple judges and each responsible for a finer or tiny thing to check?

  90. 16:54

    Like-

  91. 16:54

    Yeah. The question was, um, for LLMs as a judge, do we, we have-- she wanted to clarify that we have multiple judges each responsible for a smaller thing. The answer is yes.

  92. 17:03

    Uh, let's think about it in terms of, um, scope of what they're evaluating and magnitude of number of samples that they're evaluating. Um, usually we have a variety. Sometimes we have like a one-to-one mapping, where for a single thumbs down output, we just have a prompt that is shared with no one else, and that's what we engineer.

  93. 17:22

    We have some layers on top of Braintrust that make that easier for our data specialists to write that work with the Braintrust, um, SDK. But we have that, and then we also have LLMs as a judge that will operate on top of everything, um, that's in that dataset.

  94. 17:37

    You'll learn more in the workshop the distinctions with like datasets themselves and logs. But datasets are hand-curated, and it can be just one aspect of your trace. So in an AI interaction, you can have like five, five LLM calls that happen between the user and, like, the Notion AI response.

  95. 17:52

    We can extract just one of those and put it in a dataset. Yeah. Yeah.

  96. 17:57

    Have you tried doing any like, uh, automated prompt optimization, or is it generally like you, you'll have to go manually through a person to look at the results and then update them?

  97. 18:06

    The question was about automated prompt optimization. The answer is yes. We have played with it. I'm not sure that a majority of our problems are as solved by that as they could be in other workplace contexts.

  98. 18:16

    I've been to like dinners and events where I've heard massive success from it. Maybe we haven't cracked it yet. Um, a lot of people have, but we've played with it.

  99. 18:24

    Yeah

  100. 18:25

    One quick question, question

  101. 18:26

    Yeah

  102. 18:27

    Uh, what's your process of aligning the thumbs up, thumbs down with the scoring function?

  103. 18:32

    Yeah. The question was about aligning thumbs up, thumbs down with scoring functions. Um, they don't... I mean, so a majority of things in our data set were things that are either thumbs up or thumbs down.

  104. 18:43

    Sometimes it's just legally how it works because we have a process with some alpha users where that's the only way they give us permission to look at their data for evaluation purposes.

  105. 18:52

    Uh, we don't really rely on thumbs up for anything except for maybe internal Notino's thumbs upping things as, like, good golden data for fine-tuning, but we don't really rely...

  106. 19:01

    There's, like, no consistency in what makes someone thumbs up something, so we don't look at that that closely. And for thumbs down data, um, it's more just that this is a functionality that we know we didn't do our best work on.

  107. 19:11

    Um, but that thumbs down could have been given in September of 2023. Um, and so we don't perform how we did in September 2023, so it doesn't necessarily align with the LLM-as-a-Judge because that's judging a particular experiment, not, um, what the production user experience was at that time.

  108. 19:28

    And that's what makes this really powerful because we don't just need to look at what our output was in September 2023, right? Our data can be far more robust and last much longer because really what we're getting from the thumbs down is the natural language request from the user.

  109. 19:41

    Everything else we can re-modify. We know the state of their workspace and what their request was. Everything else, you know, we, as our server code changes, changes the output of the LLM and the engine.

  110. 19:55

    Any other questions? Yeah.

  111. 19:57

    Yeah. Uh, for the LLM-as-a-Judge, uh, is the scoring more, like, binary or is it on a scale?

  112. 20:01

    Yeah. The question was if LLM-as-a-Judge scoring is binary or a scale. I'll tell you, like, the honest answer in practice. I believe that we do it as a scale, and I don't think that scale is very well, uh, calibrated, and I believe that we just grab everything that's less than a certain score and look at them as

  113. 20:16

    all equal. Um, that's just, like, how we do it. I'm sure that there are technical ways that... I've read about people that run it multiple times and calibrate their LLMs-as-a-Judge.

  114. 20:25

    For us, we still have humans look through those outputs. We actually will take the failures and pass them into another LLM and ask them to summarize what the failures are, so that we can write a quick report for the engineer working on it to see, because we have thousands of samples.

  115. 20:38

    So then we actually look at those deltas and pass it to another LLM to tell us what the biggest themes and difference were. Um, yeah, that's how we do it.

  116. 20:46

    I actually think there's been a lot of academic research on that calibration, but it has not been necessary for us from an investment perspective, um, so we haven't worked on it much.

  117. 20:55

    Okay. A follow-up question for that. Uh, so when you change models or when you change prompts, there, uh, there are two types of e-evaluation. One is you can subjectively say, "Hey, is this, uh, safe?

  118. 21:04

    Is this trustworthy?" Or, um, or, like, binary yes or no. Uh, another thing is when you change prompt, uh, the results you can compare in pairwise, like, is this result better or-

  119. 21:14

    Mm

  120. 21:14

    ... is this result better? Do you use a combination of both?

  121. 21:16

    Yeah. The question was about pairwise combinations, um, versus just saying yes or no. Um, it depends on the experiment that we're doing. Um, oftentimes if we're experimenting between two methodologies, yes, or if the control is particularly important, like this is what's in prod and I don't want to break prod, um, then we will use that kind of

  122. 21:36

    AB, um, setup. If it's the development cycle that I talked about earlier, where we're still in dev or we're still just with alpha customers and we don't know what the golden experience is and we're far more comfortable breaking things, we're much more...

  123. 21:48

    we're much less likely to use that type of setup. But Braintrust actually has a great UI for letting you do both.

  124. 21:54

    Okay, thanks.

  125. 21:54

    Yeah. Yeah.

  126. 21:56

    Curious if you have a, a process around figuring out what criteria to judge-

  127. 22:01

    Mm

  128. 22:01

    ... for a given feature. Um, like, is there any best practices, or is it thinking-

  129. 22:06

    The question was about criteria to judge as a feature. Um, I mean, there... I'm sure there are best practices. I think for us, the reminder... So, like, you know, there's thousands of these, and then they live in this ether, and it's maintained by the people that wrote them, and then they take on a lot of power, right?

  130. 22:24

    And very few people actually investigate the losses, and that's really risky for building a robust enterprise application. Um, so things that we've learned, lessons that we've learned, because that...

  131. 22:35

    And I, I see some people nodding. You can imagine how that's a little bad sometimes. For instance, we used to use them just to catch specific regressions. Like, we have a special markdown format that works with Notion.

  132. 22:46

    We just used them for formatting at first. Um, and we'd be like, uh, "Has to, you know, be consistent with our markdown and use bullet points, and has to talk about this topic."

  133. 22:56

    All of a sudden, we weren't catching that it would, like, switch languages, because that's true in that prompt, right? And so then when you start relying on it to catch everything or to be, like, a yes for something, then you have a problem where you're over-reliant on it and you're not catching regressions.

  134. 23:11

    So there's kind of two approaches that I would suggest that make that successful. One actually goes back to your very intuitive question, which is, um, having it for a particular task.

  135. 23:21

    So we have a particular task of markdown formatting. We have a particular task of language following. Um, we don't assume that they can do everything. Or you have a very small set and commit yourself to looking at the losses.

  136. 23:33

    The problem is you don't commit to the losses, and you know that they're lossy. Um, I think that's, like, a trap that we certainly, um, fell into when we were first developing this concept about nine months ago, um, before they were called LLMs-as-a-Judge.

  137. 23:46

    Um, we just... LLM graders, right? Um, I would say that's, like, kind of the two biggest pitfalls. Yeah. Um, do we have time for more, or it's up to you?

  138. 23:55

    Yeah, we can do a couple more.

  139. 23:56

    Okay.

  140. 23:56

    Yeah.

  141. 23:56

    Um, yes.

  142. 23:58

    Um, how do you isolate RAG, uh, versus generation? How do you, like-

  143. 24:04

    How do we isolate RAG versus generation is a very good question. Um, are you asking about... Well, do you mean, like, in terms of how we evaluate it, how we evaluate changes?

  144. 24:14

    Yeah. How do you-

  145. 24:15

    Yeah. So, um- This has more to do with, like, an... What you're experimenting against. So for instance, there are times where, um, the index itself changes when you're doing an evaluation, and things like recent docs.

  146. 24:28

    Particularly for retrieval, freshness of the index or the index at time of evaluation are exceptionally important and can definitely change the downstream results, which can be unobservable and confusing to the rater or to anyone looking at the evaluation.

  147. 24:41

    So when is our next offsite? You know, maybe I did that before our Q3 offsite, and it's, it's Q1 right now, right? So, like, the answer would change. One approach is to create a technology that freezes that index.

  148. 24:53

    That's actually really expensive because you don't know what elements of that index you wanna freeze. You don't know what the query is. It's a lot of storage. What we have found is that we freeze the retrieval and take the retrieved results and operate on, like, did this retrieval actually retrieve the right thing?

  149. 25:07

    We operate on that as its own evaluation that has its own evaluation framework that has much more to do with, like, our vector databases and Elasticsearch setups. And then we have a retrieval that s- says, like, presume that this is what was retrieved, execute everything else, but, like, what was actually returned as frozen.

  150. 25:24

    And then the point at which the index... The point at which, like, the actual answer wasn't in that index and the eval can't be done, we just get rid of that sample.

  151. 25:33

    I see.

  152. 25:33

    So we either assume that retrieval worked or just do the re-retrieval. We don't really merge them very much for that exact reason. Yeah.

  153. 25:41

    Uh, just one follow-up.

  154. 25:43

    Yeah.

  155. 25:43

    Um, so do you, do you s- uh... Do you isolate the retrieval at the time of creating the golden dataset?

  156. 25:51

    Yeah. Do we isolate retrieval at time of creating? Um, it kinda depends on what we're building and what we have, like, access to. The answer is the majority of the time we try really hard not to find what the index looks like at that period of time.

  157. 26:02

    That's very technically difficult and brings up privacy questions. Um, for instance, the permissions of that object could have changed. Um, we try not to do that. Yeah. But we also get enough feedback where it doesn't become as frequent of a problem, again, because we're, we're lucky in that we use Notion so much that we get a lot

  158. 26:19

    of natural, particularly for retrieval and natural use cases. Let's do maybe one more, two more.

  159. 26:25

    Yeah.

  160. 26:25

    One more?

  161. 26:26

    One more. Sure.

  162. 26:26

    Let's do someone from the back. Yes.

  163. 26:29

    Um, how many prompts do you have, and how do you, um, manage their dependencies?

  164. 26:35

    How many prompts do we have? Um, I don't think this is private. Um, let me- [laughs] ... think about if it would be. We've over 100, like hundreds. Um,

  165. 26:46

    how do we manage dependencies? Um, I mean, some of them... Well, we have our evals for them, and that helps a lot. A lot of them are different variations of things.

  166. 26:56

    Um, for instance, like there's some model-specific variations on our prompts. Um, we have a large team. Uh, it, it hasn't really come up. I think the only time that managing a large number of prompts has been a problem is when a particular, um, model provider has an outage, and we need to switch traffic over to other model

  167. 27:13

    providers. It's not like all 4o mini traffic can just go to Sonnet 3.7, um, for a variety of reasons. Obviously, that works for a lot of it. It's like, okay, I'm willing to pay more for a small amount of time until OpenAI comes back up.

  168. 27:27

    But for some of them, like our database autofill, uh, you know, it's 12 times more expensive. We'd have like millions and millions of dollars of, of debt for that decision in that s- in that period of time.

  169. 27:37

    Um, and we can't make that. And so I think it comes more to, like, ownership over prompts and making sure that a fallback mechanism is up to date, um, so that whoever's on call, they can just say, like, you know, uh, 4o is down, and we can automatically, for the owner of that product or that prompt, they've

  170. 27:53

    already determined what to switch to instead in like a code-driven or config-based mechanism. Um, I would say that's been the most laborious aspect, is like getting everyone together and creating that system.

  171. 28:05

    I wish we did that from day one. We did it more recently. Um, and that was kind of like a tribal of elders. Everyone that's ever written a prompt, please come and fill out this config.

  172. 28:14

    Um, but otherwise, it hasn't been a problem yet.

  173. 28:18

    Yeah.

  174. 28:18

    Great.

  175. 28:19

    Okay.

  176. 28:19

    Thank you.

  177. 28:20

    Thank you, everyone.

  178. 28:20

    Thank you, Sarah Sachs. [audience applauding] Yeah, Notion AI is my lifeline. It's great to search on there. It gives me all the answers I need.

  179. 28:35

    She did not endorse this. [laughs] Okay, great. So we are the Braintrust speakers, uh, the, the two left here on the, on the podium. And, uh, we'll be going over some lectures, then jumping into the platform.

  180. 28:50

    You'll do an activity, then we'll come back,

  181. 28:53

    uh, cover a new topic, and, uh, so on. So we're gonna start with just covering why eval, what is an eval, what is an eval, excuse me. Then we'll talk about running the same thing via the, the code, via the SDK, in this case, in TypeScript.

  182. 29:09

    Then we'll move into production. So day two, how are you logging? How are you, uh, looking at the data points that you're collecting in your production application? And then, uh, finally covering human-in-the-loop.

  183. 29:21

    So how do you incorp- incorporate user feedback or, uh, actual human annotators to improve the quality of your prompts, improves, improve the quality of your dataset? And if we have time at the end, we'll do some remote evals, stuff which extends what the UI is capable of in the playground specifically.

  184. 29:41

    I think we have a poll in the Slack that's currently pinned. If you're part of the Slack channel, feel free to add your response there.

  185. 29:52

    So jumping into evals and getting started. These are some of the curated, uh, tweets that we've seen notable people, uh, talk about online. So just something to, to think about why it may be so, so useful for them.

  186. 30:10

    Well, it helps you answer questions. So again, you know, is my change regressing, uh, the performance in production? Am I using the best model for my use case, the cheapest model for my use case?

  187. 30:23

    Does it, like, have brand consistency in its responses? Am I learning from the data that I'm capturing from the logs? Uh, and am I able to debug and troubleshoot some of the responses that- Are underperforming.

  188. 30:41

    This is a trend across the whole industry. Everybody's dealing with hallucinations and, uh, performance degradation. And so we're here to, to try to set up a system using statistical analysis to catch these mistakes, uh, proactively and also with online evals reactively on that online traffic.

  189. 31:01

    This can help, help your business by allowing you to move faster, helping you reduce costs and, uh, scale teams. It's really helpful to have non-technical people collaborate with technical people, uh, to build the best AI apps possible by, uh, using the playground for non-technical people and then having that SDK, like compatibility and everything centrally managed in one

  190. 31:22

    place. Moving to the core concepts of Braintrust, so you can think of it in these three slivers. Uh, starting off with prompt engineering, right? You want to keep track of the versions of the prompts that you're, uh, iterating on.

  191. 31:37

    Uh, you want to do this rapidly. So having a playground or a place where you can quickly go through these changes and identify improvements or regressions is, is really crucial.

  192. 31:47

    Uh, then you want to have evals that are automated ideally, right? And this can be kicked off via the SDK. Uh, you'll have a score that is from zero to one as a percentage that will give you a signal of, is this going up?

  193. 32:02

    Is this going down? What needs attention? What doesn't? And then finally, observability. What's happening in production? Are users, uh, thumbs-upping or thumbs down? Are they providing, uh, ideal output that I can then use in my dataset to keep improving the evals or keep improving my AI feature?

  194. 32:22

    So jumping into, you know, what, what even is an eval? Well, the definition that we came up with is that it's a structured test that checks how well your AI system is performing.

  195. 32:31

    It helps you measure quality, reliability, and correctness across scenarios. Again, the, the criteria is really up to you. As we heard Sarah say, uh, the LLM-as-a-judge can be across whatever dimension you want to, to test.

  196. 32:47

    There are three components that go into an eval. Uh, so you have your task, which is the thing that you wanna test. Uh, this is the code or prompt, uh, that you're evaluating.

  197. 32:56

    Uh, it can be a simple, you know, one-time call to an LLM or a full agentic workflow. The complexity is really up to you. Uh, we'll see some simpler versions in the UI, and then in the SDK, it's really, you know, the complexity is up to you.

  198. 33:10

    You can, you can go crazy there. Uh, the next piece is the dataset. So this is the set of real-world examples, the test cases that you're gonna be throwing at the prompt and seeing how it's performing based on the score output, right?

  199. 33:23

    So the score is the logic behind the eval. This is outputting from zero to one hundred a score. Uh, you can use LLM as a judge or full heuristic functions.

  200. 33:34

    Uh, so it's, it's great to use a combination of the two and try to meet in the middle.

  201. 33:40

    There are two mental models to think through. Offline evals, which are evals in development. As you're iterating on the prompt, figuring out what works best or what model to go with, uh, you're writing these structured tests and running them through these predefined datasets, right?

  202. 33:56

    You're not using live traffic in this scenario, uh, but on the online eval side, you are. So this is real-time tracing. You're monitoring that live production application, and you're also getting scores from those real outputs.

  203. 34:11

    And that will allow you to diagnose problems, monitor performance, and j-- and capture that user feedback, so you can keep improving and, and close that, that feedback loop.

  204. 34:21

    So we'll show how you can use both of these types of, of evals, offline and online.

  205. 34:28

    This matrix here is really helpful to understand what should I even be improving, you know? The-- And it-- You have to look at two things, right? W- The score that's being output by your evals and using your own eyes, looking at the output of the LLM, do I think this is a good or bad output?

  206. 34:46

    If both match, you look at the prompt, it's good, high score, great. You-- That's, that's where you wanna be. But if not, you have to decide, okay, do I wanna improve my evals, the actual score, or do I wanna improve my AI app?

  207. 35:00

    And this is not very trivial. You-- There, there does require, uh, some thinking of, of what, what requires, um, my focus, improving the eval, improving the, the actual app.

  208. 35:13

    So now zooming into each of these components. So s-- uh, the task, this is the thing that you wanna test. Uh, so starting off with a prompt, right? This is, uh, a screenshot from the UI.

  209. 35:23

    It's the input that you're going to give to the LLM, right? Uh, you can set up a system prompt in Braintrust and then pass, uh, user prompts as mustache temp-templating, and that's something that we'll see in the activity shortly.

  210. 35:38

    Uh, so, you know, this is, um, great for getting started. You may have some multi-turn use cases where you're having a whole conversation and you wanna evaluate that, uh, with a tool like Braintrust.

  211. 35:51

    So that's where extra messages comes in, and you can provide the, the system prompt followed by the user message, the assistant response, the user response, the system response, maybe a tool call thrown in there, right?

  212. 36:03

    All of that prepackaged and given at once to your eval system, uh, that will then output a score.

  213. 36:12

    Tools are also used, uh, with Braintrust, and they should be for, you know, RAG applications or for agentic applications. Uh, so this can live in Braintrust. You can push your tools to the Braintrust library, and they'll become accessible to anybody playing in the playground or running experiments.

  214. 36:29

    Um, and the last piece here on-- in the task are agents in the UI. So now you can have prompts chained together where the output of the first prompt becomes the input of the next.

  215. 36:45

    This is gonna keep, uh, increasing in, in scope within Braintrust. You'll-- We'll have branching and more capabilities. Uh, this is currently a beta feature. Uh, but really exciting to be able to do end-to-end evals on prompts being chained, uh, as many as you'd like.

  216. 37:03

    So now moving into datasets. So there are three fields, right, three columns in a dataset, one of which is required, the input field. Uh, so this is what's, like, gonna be passed into the prompt.

  217. 37:14

    This is the, the user, uh, input into your AI feature application, right? Uh, the other two columns are optional. You can have your expected column, which would be that anticipated output or that ideal response.

  218. 37:28

    Uh, this is difficult to create initially, uh, and we see a lot of customers leave this blank at first, and then start to fill it in with, uh, human annotation or user feedback.

  219. 37:39

    Um, if you're doing a RAG bot, it's common to have, like, assertions. So what do you want to be included in the output? Providing that in the expected column will help, uh, make sure that, that it's using that as the ground truth.

  220. 37:54

    And then metadata is for any additional information that you wanna track at that specific, uh, dataset row, so with the specific user prompt.

  221. 38:05

    So some tips here. Uh, we recommend for you to just get started. The-- You don't need to create two hundred rows in a dataset to run your first eval.

  222. 38:13

    You know, five, ten rows is great. Just start small and get some feedback. And, you know, the next thing is don't stop iterating. Keep adding rows or tweaking rows, using logs as that, that source of truth, right?

  223. 38:30

    How are your users interacting with the feature that you're developing?

  224. 38:35

    Of course, at first, th- you may not be capturing logs, uh, but there's still that, that process, whether it's via synthetic data or with internal testing, that you can, uh, help yourself improve the dataset.

  225. 38:47

    Human reviews is another way of establishing ground truth. Very much needed in certain industries. Uh, if you're dealing in the medical space, you need doctors to look at the output, or you need lawyers or people with highly specialized skills.

  226. 39:03

    Cool. So now moving into the score types. Uh, so this is the last piece. So we covered tasks, we covered datasets, and now scores. Uh, so there's, uh, two types, right?

  227. 39:12

    We covered the LLM-as-a-judge a little bit. So this is subjective, non-deterministic, more of a qualitative assessment, uh, that will be done on the output. And then on the other side, there's a full code base heuristic score.

  228. 39:26

    This is very exact, deterministic, objective. Uh, so you wanna try to use a combination of the two, uh, and like I said, meet in the middle, right? Uh,

  229. 39:37

    there are some organizations that will choose to go one way or the other, but we found that most will typically incorporate both as much as they can. Uh, I know Sarah mentioned that LLM-as-a-judge has been crucial for her.

  230. 39:52

    So some tips here. So we recommend for you to use a higher model in the LLM-as-a-judge to evaluate the smaller ones. Uh, just something that we've noticed could be useful.

  231. 40:03

    Uh, then you wanna make sure that the, the judge is scoped to a specific criteria. Uh, we want it to, to be focused and not be a broad, uh, decision that it has to make across five or six different criteria.

  232. 40:20

    You wanna make sure that you're evaling the score. This is probably the, the most crucial piece, but you want to put the score prompt, the LLM-as-a-judge prompt, through the same rigor that your other prompts are receiving, right?

  233. 40:31

    So you wanna make sure that you're testing it on a human-annotated feedback. So what would a human think looking at the same outputs? Is it matching the LLM-as-a-judge? If not, probably needs some improvement.

  234. 40:48

    This is a little, uh, tidbit about the Braintrust UI. So we have two views, Playgrounds and Experiments. They look very similar. Uh, this wasn't always the case historically, but they've now grown to be very similar because of customer requests.

  235. 41:03

    Uh, the Playground, a way to think of it is a quick iteration. It's more ephemeral, right? So what you do in the Playground won't necessarily stick around and be part of the historical view.

  236. 41:13

    Uh, but you can always save a snapshot of it to the Experiments view and, and sort of bring it over, and that will then kick off the more traditional experiment, uh, which is the same thing that happens when you run an experiment via the, the SDK, right?

  237. 41:26

    So you run an eval via the SDK or from your CI pipeline. That will end up in the Experiments view. And the benefit of that is that you have everything in one place that you can review and compare across weeks and months.

  238. 41:39

    Uh, so you can compare your performance with the new model that dropped today with, uh, a prompt from three weeks ago in the Experiments view and see the scores change over time.

  239. 41:52

    Cool. So th- that was probably a lot, but now you understand the components and ingredients behind, uh, running an eval. So now we're gonna jump into the activity. There should be a document pinned in the Slack channel.

  240. 42:06

    If not, we can also pull up the, um, uh, QR code quickly. But if you can, uh, pull that up and try to follow along. I know the Wi-Fi has been a little spotty.

  241. 42:17

    If you can't follow along, no, no worries at all. Uh, Doug here will be leading through, uh, those steps. [clears throat]

  242. 42:27

    All right, let's have a little fun here, um, actually getting our hands dirty with the, the Braintrust platform. As Carlos mentioned, we have, uh... Sorry, wrong button. Uh, the, the, uh, the document guide out here that'll, uh, walk you through the setup of, uh, the, the workshop.

  243. 42:46

    Um, I'll, again, I'll walk through it. It'll give you, sort of enumerate, uh, the different requirements that you'll have to have at least for how it's set up today, just as like a Uh, call out here, we're using OpenAI here under the hood.

  244. 42:58

    That certainly doesn't mean that Braintrust only works with OpenAI. You can use any sort of cus- um, AI provider, uh, custom AI provider. Uh, you can sort of, uh, bring your own here as well.

  245. 43:07

    But just wanted to call out we're gonna use OpenAI here to, to do this. Another thing to call out, uh, somebody ran into this in our last workshop, is the particular version of Node that you're using.

  246. 43:17

    Uh, just don't be on version 23. Uh, 22, 20 are, are generally pretty good ones to, to use. And then, uh, let's get going. Uh, so you can, you can start here at the top by installing Node, Git, and TypeScript.

  247. 43:31

    I've already done this here on my machine, but, uh, you can get a sense for, uh, for what this looks like. Um, obviously we, we need a Braintrust, uh, uh, organization and a project here to actually go and play with that playground, uh, create some evals, um, run some experiments.

  248. 43:48

    So if you go to www.braintrust.dev, this is where you can create your, your organization and your project.

  249. 43:58

    Uh, we're calling our project Unreleased AI today. You could certainly call it something different. You would just see multiple projects be created inside of your account, and I'll show you why in just a second.

  250. 44:07

    Uh, but this is where all of this activity is going to, uh, to be, uh, happening underneath. You can almost think of as a project as, like, a particular feature, right?

  251. 44:16

    So you have feature A, B, C, and they, they may each have their own, uh, unique projects. To speak very high level about the use case here, Unreleased AI, what we're trying to do is build sort of an application that looks at a, a GitHub URL and looks at the, uh, the commits that have happened since the

  252. 44:32

    most recent release. And the idea is to inform us as developers what's coming in maybe a, a subsequent release. [sniffs]

  253. 44:41

    Once we've created our account, uh, I'll come out to the Braintrust platform here.

  254. 44:48

    So I'm gonna kind of follow along with all y'all. Uh, I'll create a project, right? Unreleased

  255. 44:55

    AI. Uh, then you're gonna need to configure your AI provider, right? And so this will be the, the OpenAI API key that you'll configure. Uh, again, here are the different AI providers that, that you can, uh, you can configure within the platform.

  256. 45:09

    You also can use, uh, default cloud providers and then, like I mentioned, bring in your own custom providers to the platform as well. Uh, configure that, and then we'll come back to the guide.

  257. 45:25

    Uh, this is where the repo exists, right? It's a, it's a public repo in our Braintrust org. You just clone this locally.

  258. 45:36

    Let's zoom in a little bit. Um, obviously I have it here, uh, locally. The, the next thing that we'll do is, uh, we're gonna copy this, uh, .env.local.example file to a .env.local.

  259. 45:52

    Uh, you're gonna replace the Braintrust API key and OpenAI API key with your specific keys. Optionally, you can, uh, include this GitHub token. This will allow you to not hit sort of, like, rate limits on the GitHub side.

  260. 46:05

    Not super important for, for this. We're probably not gonna hit that here with, uh, just running a couple examples, but that's essentially what you do, right? We're just gonna run that CP command, and you, you have that there.

  261. 46:16

    Fill in your, uh, your API keys. [sniffs] Okay. Let's do this now. Uh, so the first thing that I'll do is run Pnpm install, right? This will install all of the, the dependencies for this application.

  262. 46:32

    The other thing to call out here is it's gonna run this command, uh, after installation. So this is going to actually create some of those resources that, that Carlos was just talking about.

  263. 46:42

    Uh, it's gonna create a couple different prompts. It's gonna create some scores and actually create a dataset within that project that we just created. Uh, so this is, uh, I think just to highlight here, one of the really unique things about Braintrust is being able to connect the things that we're doing within our code to what's happening

  264. 46:58

    within the platform itself. Couple different ways to actually utilize or use the, uh, the Braintrust platform, whether you are kind of maybe strictly code-based using our SDKs, or you are using a lot more of the platform itself, you can do a lot of these things and, and share code and, and so on.

  265. 47:15

    But just to kind of scroll through here, this is how I'm creating my prompt. Maybe you scroll down a little bit further, here's another prompt. But when I run this, this is gonna actually go and push all of these resources into that project that we just created.

  266. 47:42

    Okay, cool. So we have our prompts, we have our datasets, uh, we even have some, some scores here down below. So let's just kind of investigate some of these different things that we created.

  267. 47:53

    So here's my prompt. Uh, as Carlos mentioned, we have that sort of, uh, mustache syntax here that allows us to, uh, inject certain things based on, uh, you know, whatever our dataset looks like.

  268. 48:05

    Maybe to, to back up, thinking a little bit about what Sarah said in her talk, thinking more about, like, that, that data structure that you, that you want, and then this is how we will then create that, that underlying dataset.

  269. 48:16

    But this is really gonna map to, uh, the dataset that we've created within that repo, right? So again, what we're trying to do is we're trying to get that list of commits, summarize them, and give them, uh, give sort of that, that summary to that developer of what's coming in that next release.

  270. 48:33

    Uh, the other thing t-to highlight here are our scores. A-and luckily, um, I was able to follow most of Sarah's, uh, best practices, being very targeted in what we're doing with these individual scores.

  271. 48:45

    Uh, to highlight a couple, I'll, I'll, I'll just open up my, uh, completeness score

  272. 48:51

    So this is really, uh, just being able to understand, did I-- did the LLM do a, a good enough job, uh, when it looked at the commits, is the summary complete, right?

  273. 48:59

    When it-- did it summarize all of the, the breaking changes? Did it summarize all of the net new features? Right. So this is where we can actually start to, uh, create that criteria that, uh, that, that meets the expectations that we have of what is excellent, what is good, and so on.

  274. 49:15

    And then this then maps to individual scores between zero and one. Uh, the other one here I'll highlight here is this formatting score, right? This is actually just, um, you know, more of that heuristic, uh, just being able to understand is it following a particular format that we've laid out.

  275. 49:32

    Uh, and this is a little bit more binary, just a zero or one. But just to give you a sense of the, the different scores that we can add to the project.

  276. 49:42

    Uh, last one I'll highlight is just the, the dataset. Uh, if you, if you look through here, maybe I'll zoom in a little bit further, right, you should start to see some of these things actually map, uh, what we have within, within that prompt.

  277. 49:54

    And so this is how we're gonna be able to pull those things in and actually run our evals.

  278. 50:03

    All right. Come back here, make sure I haven't, uh, skipped any steps.

  279. 50:09

    So now that we have a sense for what we just put within the platform, let's actually start, uh, playing around with it. Uh, so we're gonna go into our, uh, eval playground.

  280. 50:20

    So this is where we can actually create sort of evals on the fly. Um, I'll create an initial playground here.

  281. 50:28

    Feel free to name it something better than that.

  282. 50:31

    And, uh, here's what we can do. So let's, let's load in our two prompts that we've created. So we have our, uh, changelog one and our changelog two. Uh, thinking back to what Carlos was talking about of, like, the, the different sort of components of, uh, of an eval, right?

  283. 50:46

    We need our task, right? In this case, it's our, our prompts. Uh, we need a dataset, and then we need a, an ability to score it. And so we can actually include all of these different scores here.

  284. 50:58

    And now we have all of the different components to run an eval here within the, the platform. Uh, so I'll click Run, and then you'll actually start to see all of these different rows within this dataset actually run in parallel, and then we're gonna, uh, generate those scores against the output that the LLM, um, created for those

  285. 51:17

    individual prompts. And we should start to see some of these scores come back, uh, pretty soon. Lots of different ways to actually go in and, you know, you can, uh, highlight each one of these different things.

  286. 51:27

    You're able to actually, if-- once this starts to come back, look at the, the rationale that the LLM gave to providing the score that it, that it gave. But this allows you to, like, you know, and then you can even look at, at a diff, again, once this, uh, some of this stuff comes back.

  287. 51:43

    Or we can maybe get a more of a summary type layout, and I can look at, like, okay, so my, uh, accuracy score for my base prompt up here or my base task, how does it compare relative to my comparison task?

  288. 51:55

    So very, very quickly I got some, um, some scores here that give me a sense for how well these particular prompts fare against the scores that I've defined against, a-against them.

  289. 52:07

    Um, maybe another really quick thing to, to highlight here is, is up here at the top, right? We were talking a little bit about, like, maybe understanding over time how are these things faring as I start to change these prompts.

  290. 52:19

    That's where, uh, experiments come into play. Uh, so I can create an experiment, again, very quickly. It's gonna use what I've configured within the playground. This will actually create, uh, again, yeah, we'll back up a little bit here to experiments.

  291. 52:33

    But now I can start to track over time these scores. Uh, I can change the model of these. There's a lot of different ways to start to view, uh, this data.

  292. 52:41

    Uh, maybe I wanna understand what does cost look like, uh, for these particular tasks over time. I'm able to look at that. If I change the underlying model, I'm a- I'm able to group it by that model and see how those models affect that score.

  293. 52:54

    But this becomes a very trivial thing now.

  294. 52:58

    All right, going back to that playground, maybe I wa- I wanna test here with GPT four one, how does this, uh, impact my scores, and then how does it impact my scores relative to the cost maybe that I'm current- incurring to do that?

  295. 53:12

    So run, run that experiment. Again, now I can track these things over time. Um, maybe, maybe really quickly as sort of a, a bonus to this, um, going a little bit

  296. 53:24

    off grid here, uh, just because I've actually heard this numerous times at this conference, and we just had a question over here was like, how can we, like, automatically optimize prompts?

  297. 53:33

    This is something that we are thinking about internally, uh, and hopefully within a, a couple weeks or so we can release this, this feature to our customers. But this loop here allows you to, uh, either optimize prompts or generate net new dataset rows.

  298. 53:46

    And then the kind of unique thing here is that i-it-it has all of the context of this playground. What are the, uh, the evaluation results the last time it ran?

  299. 53:55

    How can we, uh, change this prompt to ensure we beat those results? And so you can see it's-- y- I'm guessing most of you all are familiar with, uh, something like Cursor.

  300. 54:04

    So sort of a similar type of interface here, but it will actually, again, use those results, start to change the prompt. Uh, pretty soon it'll, it'll, uh, give me like a diff here.

  301. 54:13

    Here is maybe a, a better prompt that we can accept.

  302. 54:17

    And then it'll again, it'll run those scores for that new task, and it'll, it'll give us a sense for whether or not we are better or worse because of that.

  303. 54:27

    But now this becomes a little bit easier to do that iteration when we have sort of AI layered here, um, in the mix. And, and another question that we got in the, in the previous, uh, workshop is, are we using Braintrust to actually evaluate this?

  304. 54:40

    And, and we certainly are.

  305. 54:46

    Yeah, this is really cool. This can also improve the dataset, right? You could ask it to enhance- The dataset, add another row, maybe even highlight some of the dataset rows and be like, "Can you improve the prompt based on these specific scores?"

  306. 55:02

    Yep, absolutely.

  307. 55:03

    Yeah.

  308. 55:04

    Yeah.

  309. 55:04

    Okay, cool. Any questions that you guys may have? Yeah, we can open it up.

  310. 55:11

    Yeah. How do we add... If, let's say, if we have a homegrown model, how do we add it in the list in here?

  311. 55:18

    Yeah, that's a good question. Uh, custom AI models being-

  312. 55:22

    Yeah

  313. 55:22

    ... used in Braintrust. In the settings here, in the AI providers, you can add your own custom model. Uh, so if it speaks OpenAI, you should be able to add it in here or choose whichever one you want if you have it in Inference.

  314. 55:50

    Yeah.

  315. 55:50

    What if your score does not have, like, upper bound that you can con- convert into percentages? So how do you calculate score then?

  316. 56:00

    So what, what kind of score would you be thinking about here?

  317. 56:02

    So for example, I'm building fi- finance, uh, ticker, right? Offering, like, buy some kind of stocks, and I use score as how much money I made. So that could range, you know, from $10 to, like, $10,000, right?

  318. 56:18

    How can I convert this, like, in a percentage score?

  319. 56:22

    Is that value relative to something that's expected, though? Like, you're, you're talking about, like, an unbounded range. That's not necessarily what I think these scores are designed to do.

  320. 56:33

    Right. But I can engineer my prompt and try to maximize some kind of score that I can get, right? But that does not mean that I have a range in which I try to operate.

  321. 56:43

    I'm just trying to optimize a function, some generic that I am, you know, defining my score.

  322. 56:51

    Hm. Yeah, it's definitely something we've heard before. Um, I think that's part of the, the struggle of writing scores, is trying to normalize them and, and bring them into that range of zero to one.

  323. 57:02

    So up to you of, of what you decide the floor and ceiling are.

  324. 57:08

    Yeah.

  325. 57:08

    Um, do you have anything for evaluating multi-turn conversations?

  326. 57:13

    And as part of that, would you be able to review the agent's features?

  327. 57:19

    I, I think you kinda went through that earlier, right?

  328. 57:22

    Yeah. Yeah. In the playground, we can just...

  329. 57:27

    So here at the bottom, you can add more messages. So the idea is that you could provide a whole back and forth in, in context, and then evaluate that multi-turn conversation at once.

  330. 57:39

    You could do it as well in the SDK. And then for the,

  331. 57:44

    the agents here, uh, the idea is that you can chain multiple prompts together.

  332. 57:52

    Uh, so here we could just grab the two. But the idea is that the output of the initial prompt will become the input of the next, and so on, and then you can evaluate them as a unit, right?

  333. 58:02

    And as opposed to the, the multi-turn, right, you're providing all the context at once, whereas here it's, you know, making multiple LLM calls.

  334. 58:12

    And just piggybacking on this question. For a multi-turn scenario, do we score each turn versus the entire conversation? Also, how about if the conversation has, let's say, 100 messages in between the agent and the user, um, what about the memory?

  335. 58:28

    Do, do we retain the context for the message one versus the message 90th message? Or how does this all play out? Do we score the entire conversation versus each turn?

  336. 58:38

    With the multi-turn extra messages approach, you're providing everything at once to the LLM, so it's one LLM call. So yeah, you would need to comply with the context window of the model that you're working with.

  337. 58:49

    Uh, the agent feature, though, each prompt is its own LLM call.

  338. 58:55

    I think maybe just speaking also more generally, um, if I come back over here, I just wanted to highlight something 'cause I think it's relevant to the question.

  339. 59:06

    Um, do, do you, do you evaluate sort of like the, the individual turns? Um, to me, like, it, it's always, like, somewhat dependent, but if I look at these logs here, uh, this is sort of like this, uh, this application where there's multiple steps that have to be taken for the user to get the, the output that

  340. 59:24

    it needs. So you can see here, uh, that the first step is actually rephrasing a question, right? So there's actually this chat history that is needed and then the, the most recent input, and we need to be able to rephrase that question into something that the user actually means.

  341. 59:38

    And if the, uh, the rephraser here, the task that we have for this, falls down, it means the rest of the, the application is gonna fall down. So I think one of the, the things that becomes really powerful here is the ability to actually score these individual calls as they're happening.

  342. 59:53

    So I have a rephrased question. From there, I'm gonna determine the intent. But if I scroll down a little bit further, I can actually create scores for those individual spans, so I can understand the...

  343. 1:00:05

    Like, the, the application has all of these different steps. I wanna make sure that, like, if something does not work, if the output doesn't match what I would expect, I need to be able to go back through there and figure out where it fell down.

  344. 1:00:17

    I don't know if that helps at all, but-

  345. 1:00:18

    Almost like a session trace then, right?

  346. 1:00:20

    Yeah.

  347. 1:00:20

    You're going back to the session and seeing what happened and where does this trail through the tracks.

  348. 1:00:24

    Yeah. Yeah. And then you're able to configure scores for these individual spans that are created.

  349. 1:00:32

    Um, does Braintrust support, uh, evaluating speech-to-speech models, like for real-time or Gemini multimodal, like, multimodal evaluations, voice, etc.?

  350. 1:00:46

    Yeah, that's a good question. We have a cookbook about evaluating a voice agent that you should check out. Uh, it would be through the SDK.

  351. 1:00:54

    Yeah. Any questions?

  352. 1:01:00

    Cool.

  353. 1:01:01

    Are you guys able to follow along in the activity? Like, the internet's working, everything's okay?

  354. 1:01:06

    Sorta, kinda.

  355. 1:01:07

    Sort of. [laughs] Yeah.

  356. 1:01:12

    Do you, do you guys offer Braintrust as for on-prem installation?

  357. 1:01:18

    We have a hybrid deployment, so we would manage the control plane, you would manage the data plane. Everything meets in the browser, similar to the Databricks model. So yes, our response is, is you would have the data fully in your VPC.

  358. 1:01:45

    Yes, question over here?

  359. 1:01:46

    There's a lot of similarities with LangSmith. Uh, how do you compete or not compete, but how are you different from LangSmith? I know you have more features it seems.

  360. 1:01:57

    Yeah, that's a good question. So the, uh, the question is how do we compare to LangSmith since we're both in the eval space? And, you know, I think you're right in that we're do very similar things, and maybe if you're looking from a distance, it looks like it's the same.

  361. 1:02:12

    But up close, what we hear is that our UI/UX is more intuitive, cleaner, easier to work with, and, uh, crucially, the, the performance and scale of Braintrust is unmatched due to our underlying, uh, infrastructure.

  362. 1:02:26

    Brainstore was actually developed in-house, and it was only possible because of the technical team, um, especially Ankur and the founding engineers that have, you know, deep database knowledge from working at SingleStore, MemSQL and having built many databases before.

  363. 1:02:42

    Uh, so they were able to increase the, the full, uh, tech search capabilities and yeah.

  364. 1:02:54

    Cool. So finished with activity one. Again, this will be available for you guys to, to do at your own pace at home whenever you want. Uh, there's also a, a free tier available for, uh, Braintrust, so you'll-- you have still a bunch of credits to use.

  365. 1:03:09

    So moving into the second lecture. So now doing the same thing that we did, uh, via the, the SDK. So we got to see a little bit of the code, right?

  366. 1:03:18

    We cloned the repo. Uh, we s- uh, pushed the prompts, the scores into the Braintrust, so we could use them in the playground. Uh, but here we'll actually do an experiment via code.

  367. 1:03:30

    Uh, so just wanted to quickly go over high level how this works. Uh, so you know, the, the top row we went through. Uh, you didn't really get to see the assets defined in code, uh, but we talked through the Braintrust push and then how they ended up in the Braintrust library.

  368. 1:03:46

    So the second row is what we'll go over now, where you're defining the evals in the code. You're running a Braintrust eval command. This can occur via the CLI or in a CI pipeline.

  369. 1:03:58

    That's very common, right? So when you're trying to, to open a PR and merge it into main, you would trigger this eval process, and then all the experiments would appear in Braintrust.

  370. 1:04:10

    Uh, and you know, Sarah was mentioning that they have their own process that looks at those experiment scores, summarizes them with an LLM, and then writes a report. Uh, so you can get creative with this pipeline.

  371. 1:04:24

    So pushing to Braintrust, very similar, right? The, the key here is defining the assets as code. Uh, so you would just follow the pattern, uh, to the right here, projects.prompts.create, or it would be project.scores.create.

  372. 1:04:38

    And then you would just define, uh, what would appear in Braintrust and, and this could be used in production. Uh, you could, you know, a lot of people will either push their prompts into Braintrust or pull them.

  373. 1:04:49

    Uh, really, the, the prompt management style is up to you.

  374. 1:04:53

    But it's, it is really helpful to have that, uh, source-controlled prompt versioning in place. Uh, depends if you wanna go push or pull.

  375. 1:05:06

    Running an eval via the S-SDK is very similar to what we saw in the UI. You need the three things again, right? So you need a task, a dataset, and one or more scores.

  376. 1:05:15

    Uh, so we'll show you what that looks like in the actual code, and then we'll, we'll run the command and we'll, and we'll see what that looks like in the, in the UI.

  377. 1:05:24

    And then we'll come back and, and do some of the logging day two, uh, moving into production stuff.

  378. 1:05:34

    Cool. Um, back to the guide here. Uh, if you're looking on activity two, uh, we'll start again with just kind of understanding what we have within the code base, and then we'll actually run some evals.

  379. 1:05:45

    You'll, you'll, again, like Carlos mentioned, you'll see some similarities with what we did within the platform, but now we're just kind of strictly within the code. Again, just trying to like hammer home here there, there's like multiple ways to actually, uh, interact with the Braintrust platform, uh, trying to meet you all where you want to be.

  380. 1:06:01

    Uh, but if I come over to the code base, uh, again, this is where the resources.ts, this is where all of this will be defined, all of our pro- uh, prompts, scores and such.

  381. 1:06:13

    Uh, I have this eval folder. Uh, I'm gonna open up my changelog.eval.ts. One thing to kind of highlight here, um, if you, uh, if you name your files with that .eval.ts and then you pass just a folder, we can pick up all of those evals a-and run those within the platform.

  382. 1:06:31

    But, uh, let me open this up a little bit. So this is really all we're doing, right? So we have this eval, right, which is coming from the Braintrust SDK.

  383. 1:06:41

    Here's our dataset that we've created. Uh, the task that, um, you know, that lives over, uh, here from, uh, from Braintrust. And this is really like, again, all we're doing from the, the repo side to run evals within the platform.

  384. 1:06:58

    So I'll clear this and then run pnpm eval Uh, this is again just another command that is, uh, unique to this repo. All we're doing under the hood is running this Braintrust eval and then giving it the name of this file.

  385. 1:07:12

    But you can actually see the experiments, uh, running right now within the platform. Here is the link that you can go out to, and this should look very similar to what you just saw when we ran the experiments within the, the UI specifically.

  386. 1:07:26

    But, but now again, this, this becomes, uh, part of the, uh, the history of our experiments for this particular feature and can track this over time, but now we are kind of strictly within the code while we are, while we are running this.

  387. 1:07:47

    Um, maybe just a-another thing to kind of highlight, I, I think I did last time, but I didn't really, um, pull anything in here, is, is different ways in which we can start to, um, compare this.

  388. 1:07:59

    So, uh, maybe we're, we're really worried about sort of like the, the duration or we are worried about sort of the, the number of, um, completion tokens. There are different ways in which we can start to view the, the data here.

  389. 1:08:13

    Uh, again, I didn't change-- Or actually, I did change the model, right? This is another way in which we can start to, uh, understand the, the scores for the particular model.

  390. 1:08:21

    But again, like, this all just happened via, via code, uh, some simple commands that, that we just, uh, wrote on our command line a-and pushed that into the Braintrust platform.

  391. 1:08:36

    Cool. Any questions about running evals with the SDK and that Eval, uh, class that we showed and the Braintrust Eval command?

  392. 1:08:47

    Yeah.

  393. 1:08:48

    Just wanted to confirm with you, uh, like there is, uh, image like, like vision models and stuff. Does that work with the UI or is that only SDK?

  394. 1:08:59

    Yeah, that should work with the UI images, yeah.

  395. 1:09:01

    And can I do non-LLM like models? Like if I have like a mo- like a grounding model as part of my pipeline?

  396. 1:09:11

    Yeah, that's a good question. Um, I'm not sure if it will let you add it as a custom model. Um,

  397. 1:09:20

    yeah, I can't-- I, I would say go try it and see and, and if there's any issues, we have a Discord channel and, you know, we could add it to see if more people want that feature.

  398. 1:09:30

    Sure.

  399. 1:09:30

    Yeah.

  400. 1:09:31

    Thank you.

  401. 1:09:36

    Yeah. Yes.

  402. 1:09:36

    Python-based SDK?

  403. 1:09:38

    Yeah, we have a Python and TypeScript-based SDK, and then we have some additional ones in, in other languages like, like Java or Kotlin that, um, Ruby that are created with Stainless, so they're just wrappers of the API.

  404. 1:09:54

    We're, uh, creating a Go SDK as well.

  405. 1:10:03

    Great. Any other questions about activity two? Uh, ready to move into, uh, production and we can talk about logging.

  406. 1:10:14

    All right. Let's do it. So why should you even set up logging? Well, you know, there's a few reasons here. Uh, the main reason is to measure the quality of your live traffic.

  407. 1:10:24

    Uh, so this will allow you to, to set up alerts, if you'd like, of when it starts to dip, right? If you're putting out less than your best work and users are noticing, maybe they're providing feedback or, you know, maybe they, um, you know, the scores that you have o-on those specific logs are just dipping, uh, you

  408. 1:10:45

    can notify the right teams. You can debug, troubleshoot a lot faster by having that visibility into all of the functions that are running, all of the, the LLM calls going back and forth.

  409. 1:10:56

    Um, and crucially, you can close the feedback loop. So you can capture the feedback and, you know, we hear a lot of customers that are doing that, uh, but they're not implementing it into their, uh, iteration process.

  410. 1:11:08

    They're not changing the prompts or, or, uh, adding it to datasets and, and closing that loop, so this can really help you do that.

  411. 1:11:17

    So what does logging into Braintrust, uh, look like? Uh, we have a very easy path, which is wrapping your LLM client. Uh, so y- if you're using OpenAI, uh, you can just use the, the wrap OpenAI function around it.

  412. 1:11:31

    Uh, if you're using the Vercel AI SDK, we have a very similar thing. So this is essentially like one line of code and then now everything is, is getting tracked, all your metrics, right, the, the amount of tokens, your latency, uh, all the cost, all of that is getting populated in, in Braintrust.

  413. 1:11:48

    Uh, you can also trace arbitrary functions. We have the same sort of wrapping approach where you, you d- use a trace decorator or wrap trace, and it will, uh, track everything occurring within that function.

  414. 1:12:00

    If you wanna be more specific or more granular or have more flexibility, you can also do that. We integrate with Otel, and we have a Span.log support, which can allow you to add any additional information, metadata that you wanna track, uh, and add it to, to that interaction.

  415. 1:12:17

    And initializing the logger, right? That's important 'cause it tracks which project you want this to be synced to, and it authenticates you into Braintrust. So that's part of the process.

  416. 1:12:25

    And, uh, in the code repo that you clone, you'll see that if you go to the, uh, generate/route.ts, you'll be able to see all of these, uh, steps, uh, in the code and, and understand what's needed to, to get it into Braintrust.

  417. 1:12:43

    So now that you have the logging and real user traffic entering Braintrust, you wanna have that, uh, assurance of quality or you wanna have that visibility into, into quality of responses, and a great way of doing that is through online scoring.

  418. 1:12:58

    Uh, so this is going to use the scores from your playground or from your experiments, right, from your offline evals and bringing them into that production environment, into that live traffic and, uh, giving you scores that you can then use to, to filter down the logs, grab the, you know, underperforming use cases and add them to a

  419. 1:13:18

    dataset, keep iterating, keep improving.

  420. 1:13:22

    You can also do some A/B testing, right? You, you tag, uh, certain, uh, logs in different ways, and you can compare the online scores, uh, for each one of those.

  421. 1:13:34

    So we'll show you in the activity how you can, uh, set up s- online scoring rules and, uh, which spans you can, uh, enforce them to be on and, uh, the sample size.

  422. 1:13:45

    So you could choose to, to online score 100% of the logs or 1%. It's up to you. We do recommend for you to start with a low sampling rate, uh, and then work your way up as you trust the metrics more and more.

  423. 1:14:00

    Uh, so this is a bit of the process, and then Doug here will show you in the platform how you can go about creating the, the rule.

  424. 1:14:13

    So once you have online scoring set up, it's really helpful to have one-click lenses of what matters to you or to your team, so you can think-- You know, if you've used Notion, you've probably realized that they have, uh, views, right?

  425. 1:14:27

    You can customize the filters, save it to a view. Same idea in Braintrust. Uh, you can filter, sort, do whatever you want, save it to the view. It'll, uh, retain that, those settings, and then your team can just use that.

  426. 1:14:39

    So it's a great way of moving quickly, defining, uh, the, the filters that matter to you.

  427. 1:14:47

    Great. So now back to the platform, back to the repo, and go over logging and online scoring.

  428. 1:14:54

    Awesome. Let's, uh, let's configure some, uh, some online scoring. [clears throat]

  429. 1:15:01

    All right. Uh, before, before we do that, obviously, we need to, like, get some, some logs into our Braintrust project. Uh, to do that, we'll come back to our application, and we're gonna spin this up.

  430. 1:15:10

    So I'm gonna run Pnpm dev, and this is gonna run this, uh, localhost three thousand. So again, the, the use case that we're building for is that, that change log generator.

  431. 1:15:21

    Uh, the idea here is you, you provided that GitHub URL, and we're gonna give you, um, you know, the, the LLM will give you the summary of those commits and then maybe categorize them into different, uh, different categories.

  432. 1:15:31

    So, uh, I'll click that. This should go through the process. Actually, while it does that, let me come back here just to kind of, um, highlight where this is all happening.

  433. 1:15:40

    So, uh, just kind of connecting what, what Carlos was talking about, like how we can start to instrument our application. There's a few different things here, right? We have our, uh, Braintrust SDK from-- or for TypeScript, obviously, and we're gonna wrap the AI SDK model.

  434. 1:15:56

    Uh, so this is really gonna give us a lot of, uh, goodness out of the box, uh, being able to understand, like, uh, the, the metrics, right? What are the, the completion tokens?

  435. 1:16:04

    What is the cost? All of that stuff happens simply by just wrapping your, uh, the, the LLM client with that. Uh, that's again, like, really the, the easy way to do it.

  436. 1:16:13

    Then there's, uh, being able to, like, m- uh, you know, instrument your application with a little bit more granularity. And so, uh, maybe you wanna use this trace function from the logger to actually be able to specify what that input is, what that output should be, and then add some metadata.

  437. 1:16:29

    The reason you might want to do something like this is now this maps to the, the data structure, right? That becomes really powerful to set up those scores against, right?

  438. 1:16:37

    So we start with something that's very basic, where we get something very easy out of the box, and then we can go down to the, the actual like, you know, the individual tool calls, if you will, uh, and then actually create the inputs and outputs for that so that we can then create scores on top of those.

  439. 1:16:54

    Uh, back to the application here. Right here is the, you know, essentially the response from the large language model. Uh, coming back into Braintrust,

  440. 1:17:04

    we should now see a, a log. And then this is essentially what just happened, right? Because of what we configured with that Wrap AI SDK, uh, as well as the, the different logs that we had set up, right?

  441. 1:17:15

    So there's this top level. Uh, think of this as a, as a trace, and all of these, uh, sort of under the hood or underneath it are the individual spans that we actually wanna understand.

  442. 1:17:24

    This a- this allows us the visibility into the multiple steps that the application will take and allows us, uh, like I s- like I, uh, s- showed you in that previous example, to create scores for these individual spans as well.

  443. 1:17:37

    Uh, but you can see here, right at the top level, Generate request. Here is my input, right? What are my commits for this repo? Here's the URL, and then here is the output.

  444. 1:17:47

    But then being able to drill down into these individual things, right? So what are my commits? What are my-- what is my latest relea-- excuse me, release. Uh, and then again, being able to, uh, to drill down into these becomes really beneficial as we start to need to understand how our application is performing.

  445. 1:18:04

    From there, right, just gonna come back here. Uh, we now want to s- maybe set up some online scoring. So, uh, we walked through both, uh, within the platform running experiments.

  446. 1:18:13

    Wa-walked through the SD-SDK running experiments as well. This is pre-production, right? This is, uh, now when we have our logs running in production, we wanna understand how our application is performing, uh, relevant or, excuse me, uh, relative to the scores that we've already configured.

  447. 1:18:30

    So I'm gonna come back into Braintrust, click on the Settings, and then the...

  448. 1:18:39

    Where am I? Excuse me. Just Configuration over here on the left, and then Online scoring.

  449. 1:18:55

    So this is where we can now create the rule. Uh, let's say this is...

  450. 1:19:03

    Um, this will allow you s-to select the different scores that you've configured within your project. Now, uh, I could select all of these, or maybe I have, like in that example, uh, that I showed to you previously, I really just wanna apply this to a particular, uh, span, right?

  451. 1:19:18

    And this allows you to get very granular. Uh, right now we'll just configure all of these for the, the overall trace, and we'll do it for 100% of the, uh, the samples.

  452. 1:19:29

    Oh. Okay. So now when we come back to our application,

  453. 1:19:45

    and when we generate this, we'll have the same sort of workflow, right? We're gonna generate this. Uh, the, the logs will look the exact same as they did previously, but now we are gonna score this.

  454. 1:19:54

    Uh, we're gonna understand the things that are happening with- within production. How does this output fare to the, the scores that we've, um, uh, configured for this particular task?

  455. 1:20:05

    This allows us to understand, we could even create automation rules here, so maybe a score drops below fifty percent. Whatever it is, we can now create sort of an automation that allows us to understand, uh, when something like that happens.

  456. 1:20:20

    So you could see here, uh, we have, uh, our scores ran. Again, you can, you can drill down into these things, uh, and you can see, you know, what is the rationale for that particular score in that instance.

  457. 1:20:37

    Um, maybe the other thing here, and then because we had, uh, some good outputs, but this is where, like, we would wanna think about, um, how do we give users the ability to, to traverse through these logs in a really, uh, scoped-down way.

  458. 1:20:52

    Uh, this is where what Carlos was talking about with, with views. And so maybe it's being able to understand, um, let's-- if I open this up and I create a filter on my completeness score, and I'm just gonna modify this slightly.

  459. 1:21:02

    I'll say less than, uh, fifty percent. This now can become a view that I save, uh, to

  460. 1:21:12

    this project. And now it becomes really easy for somebody to come through and look through these logs, and they can start to understand the inputs and the outputs. And if, if it makes sense, we can add these spans to the dataset.

  461. 1:21:26

    That's another thing that, uh, I think when you think about, like, the, the, the, the sort of feedback loop, uh, that we need here, we were- we're kind of like, uh, in pre-prod, right?

  462. 1:21:35

    Where we're developing, uh, evals in the UI or, uh, in our code base. We push it into production. How do we create that flywheel effect from those two different places?

  463. 1:21:44

    This is part of it, right? Being able to, uh, enable sort of like the, the humans, uh, looking at these things, and then being able to add those really unique or relevant spans to the existing datasets or even creating net new datasets from them.

  464. 1:22:02

    Cool. That's, uh, activity three. Questions?

  465. 1:22:05

    Yes. Can you aggregate those scores?

  466. 1:22:09

    Aggregate-

  467. 1:22:09

    Like average or...

  468. 1:22:13

    If you go to the Settings. Aggregate scores here. So you can do that. So, uh, like... Yeah, you could create a true north, right? And just based on the aggregate score, decide if it's improving or regressing.

  469. 1:22:31

    So, so the SDK is meant to be pushed to production then to get those data logs?

  470. 1:22:36

    If you wanna do online scoring, you'll need to push scores into Braintrust, right? So that they can then be referenced, 'cause online scoring, if you're using the SaaS model, right, we're running that compute.

  471. 1:22:48

    We're, we're ex-- you know, we're grading the, the logs.

  472. 1:22:54

    Um, so yeah, I think the-- when you would wanna push to Braintrust is for online scoring. If you have teammates that are using the playground and they want access to the prompts or the tools that you've made in your SDK, that's when it makes sense.

  473. 1:23:10

    Uh, but you don't need to, right? Like, you could still do everything with code and get value from, from Braintrust.

  474. 1:23:18

    Um, maybe just m- to highlight really quick, um, we also have just sort of scores out of the box that you can, that you can configure. Uh, so if I go to

  475. 1:23:29

    the playground, like if you don't wanna push up certain things. Now, th- there's not as much flexibility in this as, as there are with your custom scores, but there is an auto evals package that is developed internally by us that you can use to, to add to these, uh, these tasks as well, and then you can use,

  476. 1:23:46

    uh, an online scoring also.

  477. 1:23:53

    Any other questions about logging or online scoring?

  478. 1:24:00

    Okay, cool. Now we're gonna move into the human section. Human in the loop.

  479. 1:24:07

    My favorite part. I'm needed. All right. Let's get, let's get into it. So, uh, yeah, it's, it's honestly really helpful for certain companies to implement human in the loop, right?

  480. 1:24:19

    Uh, not all AI mistakes are created equal. Depending on the industry that you're in, like healthcare, finance, or legal tech, a single failure can have, you know, huge, serious consequences.

  481. 1:24:32

    Uh, so that's where you may want to have and invest more resources in putting humans in front of those AI responses and making sure that they're up to par.

  482. 1:24:43

    This can be really helpful for catching hallucinations, especially in, uh, specific fields that require specific knowledge, establishing ground truth. So if you're gonna start filling out that expected column in the dataset of ideal outputs or, you know, ideal responses, it may be helpful to have a human help create that.

  483. 1:25:01

    Uh, and, um, you know, you also w- wanna capture user feedback and make sure that that's being incorporated in your feedback loop and, and, uh, being used in,

  484. 1:25:13

    in the prompt iteration. So this is great for quality, reliability, providing that ground truth, and aligning with what human, uh, humans need i- in specific areas where the, the LLM just doesn't have as much, uh, visibility.

  485. 1:25:33

    There are two types of human-in-the-loop that we're gonna cover. Human review, uh, so this is a, an annotator and a subject matter expert that's gonna come into Braintrust. They're gonna manually label a specific dataset or a specific log, right?

  486. 1:25:47

    They'll score it or they'll audit outputs. Uh, and this is really useful for, uh, ground truth to the dataset. Uh, they can audit certain edge cases, and they can also help, uh, an, uh, eval the LLM as a judge.

  487. 1:26:01

    So based on what that human annotator determined was a good response, you could train the LLM as a judge to also label it the same.

  488. 1:26:11

    And then for user feedback, that's, you know, closing the feedback loop. You're incorporating what your real users are thinking. If they keep thumbs downing a certain edge case, you can incorporate that feedback and add it into a specific dataset, you know.

  489. 1:26:25

    And that is then used to, to help improve, uh, that s- that problem that you're dealing with.

  490. 1:26:34

    So that's it. You know, now we're gonna jump back in and show you how you can enable user feedback and set up human review. And at the end, we'll, we'll do some questions.

  491. 1:26:44

    Uh, we also have a bonus lecture and activity about remote evals that if you stick around, we'll probably have time to get to.

  492. 1:26:55

    Cool. Uh, let's jump into activity four. Obviously, we're gonna look at, uh, user feedback and human review, the c- the sort of like two different, uh, routes here with human-in-the-loop.

  493. 1:27:05

    Maybe to give a little, uh, background here or just to give you, uh, show you where some of this is happening,

  494. 1:27:11

    uh, if you look at API feedback and then route.ts, the, the one thing I'll highlight is this log feedback. So this is really all the, like, what we need on the Braintrust side, again, using the, the Braintrust, uh, SDK.

  495. 1:27:25

    Uh, we know of the span that we wanna log this to from, uh, the current span that we are in. This is the user feedback score, right? So this is, like, essentially what we're calling it, user_feedback, and then the score.

  496. 1:27:38

    In our case, it's really just gonna be a thumbs up or thumbs down, so a zero or a one. We also optionally have the ability to add a comment.

  497. 1:27:45

    So if the user, uh, like your application has the ability for the user to define, uh, you know, free text, you can also put that in here as well, and then also, of course, using metadata.

  498. 1:27:56

    Metadata is-- becomes really powerful from the filtering aspect. Uh, so, like, we can use this within our custom views to filter on these different metadata, uh, keys.

  499. 1:28:09

    Uh, coming back to the app, and you may have seen this, uh, at the bottom when we ran this last time,

  500. 1:28:16

    um, but we have at the bottom the, the thumbs up and the thumbs down. Uh, this will allow, uh, the application here to, to log this particular user's feedback.

  501. 1:28:26

    So based on what they think, we can, you know, obviously do a thumbs up.

  502. 1:28:32

    We can submit with comment. And now if we go back into the Braintrust application, we have that net new log. We have, uh, over here on the right, we have our comment from the user, and then we should see a net new column here within our, uh, within our logs for that user feedback.

  503. 1:28:49

    But again, as, as Carlos pointed out, we can now use this in a lot of different ways to help augment the, the workflows that we've already created. This can help, um, you know, create custom views on top of this.

  504. 1:29:00

    This can now, uh, feed into a workflow for our, like, human annotators to understand what that user feedback is and whether or not we should incorporate it into other datasets.

  505. 1:29:09

    Again, being able to click here, add a span to a dataset, and so on.

  506. 1:29:15

    Uh, that's the, the sort of user feedback route. Um, the other one is the, the human review. So back in our configuration pane over here,

  507. 1:29:28

    uh, we have the ability to create different human review scores. So again, this will be a little bit more specific to, uh, your, your use case and what you want your, uh, humans to, to review or the scores that they should, they should create.

  508. 1:29:40

    Uh, it could be a, an option, it could be a slider, or it could be sort of free-form, uh, text.

  509. 1:29:48

    Um, so we could say something like, "What's a, a better score?" Uh, and then give it sort of an option, you know, uh, thumbs up

  510. 1:30:00

    and thumbs down. So sort of like, uh, the, the human now can sort of look at what the output was and give their own sort of, uh, thumbs up, thumb down of this.

  511. 1:30:10

    And now if you come back into the logs UI, uh, users have a different-- uh, users that are doing human review can have a different, uh, interface into those logs where maybe something like this is a little bit too much for what they're, what they're trying to do.

  512. 1:30:26

    Let's give them a little bit more pared down, uh, look at the, um, at the inputs and outputs of this particular span. And then they have the different scores that we've configured over here, right?

  513. 1:30:37

    And so maybe this is a thumbs up. They'll go through, and they can do that review. Again, this now, uh, can add to that, that sort of flywheel effect that we need to create to build this, this GenI- GenAI app.

  514. 1:30:51

    Um, of course, yeah, I think I mentioned this a, a few different times, but being able to now, uh, add some of these spans to the, the dataset, uh, we can do that, uh, here as well.

  515. 1:31:07

    So now we, we should have, you know, uh, if you look at the, the sample data, we had seven rows. Through this process, we've added two net new rows to this dataset.

  516. 1:31:16

    This now can help our offline evals and really create a, a rigorous sort of, uh, workflow to building a robust AI app.

  517. 1:31:28

    Cool. Uh, any, any questions?

  518. 1:31:30

    Yes. So those human evaluators, they have to have their own-

  519. 1:31:35

    Yes.

  520. 1:31:36

    Yeah, but how do you sh- how will they see those, uh, evals that they need to evaluate?

  521. 1:31:43

    Yeah. So you can add them to the specific project that you would want them to, to annotate. You could assign them. Uh, so we have the ability to add comments under specific-- under datasets, under experiments, or a specific trace.

  522. 1:31:59

    You would tag them in something that you would want them to see. We have sometimes, you know, you'll create an experiment, and you'll assign it to a specific annotator, and they'll go in and fill out everything, uh, in that experiment or in a dataset.

  523. 1:32:14

    Yeah.

  524. 1:32:14

    But the f- uh, sorry. The first one is developed by the client, though, right? 'Cause it's shown in that webpage where you have thumbs down and thumbs up.

  525. 1:32:23

    Yeah. So the actual op-- like, the text box that you want them to fill out or, you know, the, the categories that you want them to thumbs up or thumbs down, you would define initially, and then you would let them go in and, and they would say, "Okay, based on completeness, I would give this a thumbs down,"

  526. 1:32:41

    or, you know, uh, "What w- what is the ideal output? Let me use this text box and, and type it all out." Uh, so, so yeah, you would essentially create the form that they would then go and fill out.

  527. 1:32:56

    Any other questions? Yeah.

  528. 1:33:01

    Um, this isn't so much about the human evals, but I saw on your documentation that you support, like, writing tools and uploading those as well and kinda like injecting those in your, um, in like the evaluation prompt, right?

  529. 1:33:16

    Um, is there a way to like version the tools as well? Because like tools will also like continue to over time, and so, like, how can we keep track of that as well in the system?

  530. 1:33:25

    Yeah, that's a good question. Yeah, I believe it supports versioning. Um,

  531. 1:33:33

    I guess this project doesn't have any tools.

  532. 1:33:35

    No.

  533. 1:33:41

    I'm gonna give you a soft like, yes, it should be versioned. Same with the prompts and, and scores. Like, they all-- At the very bottom, if you create a tool and add a new version of it, at the very bottom...

  534. 1:33:52

    W-we can show you with the prompts, potentially if...

  535. 1:33:58

    Yeah, so at the very bottom, we see this one version that was initially created, but if we were to go in and make a change,

  536. 1:34:09

    save version, now we'd have that new version. So I, I believe the same thing would happen with the tools.

  537. 1:34:15

    So like you would like add a tool to this, and that would just be like a new version of the eval?

  538. 1:34:20

    Yeah, yeah, exactly. Y-yeah, if you were to add a tool or change the prompt at all, it would be a new version of the prompt. You'd save it. It would now appear at the bottom.

  539. 1:34:29

    And your question was specifically like in tools, right? Like, if you change a tool, would it be versioned?

  540. 1:34:33

    Yeah, I guess-

  541. 1:34:34

    Yeah

  542. 1:34:34

    ... because you have like different... Oh, thanks. [laughs] Um, I guess you have like different prompt versions, and then you have different tool versions, and then you're going to like evaluate those, like mix and match those together, right?

  543. 1:34:45

    To like-

  544. 1:34:46

    Mm-hmm

  545. 1:34:46

    ... see which performs the best.

  546. 1:34:49

    Yeah. So there would be a scenario where the tool doesn't change, just the version that you're using and the prompt changes, and how would you then version the prompt?

  547. 1:34:58

    Right.

  548. 1:34:59

    Yeah. That's an interesting scenario. I, I think we version it correctly today, but if there's any issues, you know, the, the SE team is here to help, so feel free to, to reach out.

  549. 1:35:10

    Great. Thank you.

  550. 1:35:11

    Yeah. Any other questions? Human in the loop?

  551. 1:35:21

    Okay, cool. All right, so now we're into the bonus session, I believe. Oh, not this one, I guess. Let's see.

  552. 1:35:35

    No bonus for you guys. [laughs] No, uh, this one has it.

  553. 1:35:44

    And it's... Oh my God. [laughs] Okay, we can just do it this way.

  554. 1:35:59

    All right, so the new bonus session. Look at this. Extending playgrounds with remote evals.

  555. 1:36:07

    So as you've seen, right, the playground can handle one LLM call being made. It can handle multiple LLM calls chained together, right? The agent feature. But if you wanted to have intermediate code steps, uh, in between the prompts, not supported, right?

  556. 1:36:25

    You'd have to go via the SDK. Or if you wanted it to talk to, uh, these external systems that are part of your VPC, that also wouldn't work. Or essentially, if your task gets too complex, it is not-- it doesn't fit into the mold of the playground at the moment.

  557. 1:36:43

    Uh, so that's where Remote Evals comes in. It allows you to expose that local environment that you've set up in your laptop that has all the bells and whistles, has everything that you want set up, uh, you can now make that available in, in the playground.

  558. 1:37:01

    Uh, so if you have custom internal tooling, intermediate code steps, you know, a dynamic R&D environment that's constantly changing and you don't wanna keep pushing tools or pushing scores, everything into Braintrust, then you can just set up a remote eval and it will become available to anybody in the playground.

  559. 1:37:17

    Uh, so this is a great way of bridging the gap between very complex technical teams building crazy tasks in the SDK and non-technical people, maybe PMs, SMEs, that still wanna be hands-on, still want to, to iterate on the prompt, change certain parameters, and see how that impacts the scores.

  560. 1:37:38

    Uh, so the remote evals can really help there.

  561. 1:37:44

    So how this works, right? Braintrust will send an eval request to your remote server. This can be localhost, or you can set it up somewhere else. Uh, your server will then run the task and the score logic all locally.

  562. 1:37:57

    So you-- All the intermediate steps, crazy business that you have going on, it'll all work 'cause it's occurring on that server. Same with the scores. You could have crazy complex scores, that would all work as well.

  563. 1:38:09

    Um, and then it will return the outputs, the scores, metadata to Braintrust, and it'll populate the playground.

  564. 1:38:19

    To start this, it's very simple. You do the same Braintrust eval, and then point to a specific eval.ts file. Uh, the difference is that you use the dash, dash, uh, dev flag.

  565. 1:38:30

    Uh, so this will start that remote eval server. It will default to localhost, but you, you can change it. Um, and it will allow it to bind to, to all the-- It will allow it to listen to all the network interfaces, so you could access it, uh, via...

  566. 1:38:48

    Yeah, access remote servers out in the world.

  567. 1:38:52

    Cool. So now we can get into the activity. We have something to, you know... In-- If you follow along with the activity document, you could set this up yourself.

  568. 1:39:00

    Uh, but if not, we can just show you quickly.

  569. 1:39:04

    All right, let's first look at the, the code. It's gonna look very similar, uh, to what I showed the, the first time. Uh, slightly different prompt. Just trying to layer in a little bit more complexity here within this task.

  570. 1:39:16

    We now have all these different parameters that we want the user to be able to define and don't yet necessarily have the ability to do that within, uh, or in a great way within, within Braintrust.

  571. 1:39:27

    Again, this can, this can be very, very complex. Uh, whatever sort of like, you know, the, the intermediate steps that you wanna create, uh, certainly allowed here. Um, I'm gonna take down this server and I'm gonna run, uh, this command, pnpm remote eval, which just under the hood is running that Braintrust eval, and it's giving the, the

  572. 1:39:49

    location of that file that we're running the remote eval on.

  573. 1:39:55

    Back in the Braintrust Playground, this now is exposed as something that we can, um, we can load in here.

  574. 1:40:03

    So let me remove this. I'll remove this here as well. And now if you look down here, we have this, this other option for remote eval, and this should link out to my changelog generator eval.

  575. 1:40:17

    Here are the different parameter, excuse me, parameters that I'm able to now configure, uh, that I've sort of exposed in that prompt. But again, now think of like the, the more complex type of workflows that, that you may have, uh, within your code base, and you don't, uh, have the ability necessarily to push those into Braintrust.

  576. 1:40:34

    We can still leverage this platform with that existing code base via remote evals, and this sort of like, uh, interaction works very similar to, to what I showed earlier with, uh, the evals locally via the SDK as well as the, the playground.

  577. 1:40:51

    Um, we should... Oh, well, probably have to... Oh, we need to configure that. My scores. Uh, I'll switch this to the grid layout so we can start to see what the, uh, the output is.

  578. 1:41:03

    Uh, but this is how we can really, uh, meet a maybe more technical team with a more complex task, and they're not necessarily wanting to push all of these things into the Braintrust platform, but still leverage it i-in some way.

  579. 1:41:16

    This is how we can sort of, uh, bridge that gap.

  580. 1:41:22

    And grounding this in the Notion use case, they use a lot of services, right? They have a lot of tools that will retrieve different user information, and they run all their evals via the SDK.

  581. 1:41:37

    And if they notice that there's an issue, right, it's underperforming in a certain area, a certain edge case isn't great, maybe they isolate a certain row that they wanna keep rerunning.

  582. 1:41:48

    Currently, they have to go to the SDK and run everything over and over again. Uh, so if they were to use remote evals, they could bring that custom task into the playground, and they could isolate just one row and run it on one specific thing that they're working on.

  583. 1:42:05

    They'd save money, and they'd move faster. Uh, so that's something that we're gonna discuss shortly with them and a lot of our other customers. This is a brand-new feature, and it's something that we've been hearing, um, a lot of excitement around.

  584. 1:42:19

    Uh, so I hope you, you see the vision of where this could be effective. Um,

  585. 1:42:24

    and yeah, feel free to, you know, if you have any questions about remote evals or about anything else in the presentation, we've pretty much reached the end. And thank you for sticking around.

  586. 1:42:33

    I know it's, uh, been a long day- [chuckles] ... to say the least. [laughing] [clapping]

  587. 1:42:42

    For the remote, do you need to put your IP in public?

  588. 1:42:46

    I'm right here. I'm right here. For the remote, for the re- Hi.

  589. 1:42:52

    Yeah, it's okay. For the remote, uh, call, do you need your IP address publicly available so Braintrust can hit it or?

  590. 1:43:01

    Uh, yes. Or you can use localhost. Um, so I mean, I guess, yeah.

  591. 1:43:09

    I, I can't imagine. Like, if, if it's, uh, private, then Braintrust wouldn't be able to, to connect to it.

  592. 1:43:15

    Okay. All right.

  593. 1:43:26

    Any other questions? Great. Well, thanks everyone.

  594. 1:43:33

    Thanks all.

  595. 1:43:33

    Hope you have a, a good night, and we'll see you, uh, tomorrow, day two. [upbeat music]