Agentic Search vs Vector Search for Coding Agents: We Ran the Eval — Braintrust

Read the talk

Agentic Search vs Vector Search for Coding Agents: We Ran the Eval

Selected presentation frame from Agentic Search vs Vector Search for Coding Agents: We Ran the Eval — Braintrust at 906 secondsOpen full source frame
Results slide comparing vector and agentic search accuracy and cost.

Jess Wang walks through an end-to-end coding-agent evaluation: turn real bug fixes into tasks, hold the agent harness steady, score with the repository’s tests, and inspect traces to understand why similar accuracy came with very different costs.

From a talk by Jess Wang

At a glance

Ideas worth remembering

  • An eval combines a dataset, a task, a scorer, and an experiment configuration to answer a specific question about an AI system.

  • Merged bug-fix PRs can supply tasks by restoring the parent commit and generating a bug description from the fix diff.

  • A vector-only condition needs enforced tool restrictions; prompt instructions alone do not reliably keep the agent from using repository exploration.

  • In this small experiment, equal reported accuracy concealed about four times the cost for vector search. Traces linked the extra calls to repeated retrieval that lacked surrounding code relationships.

  • Repeated trials, stronger retrieval, and a broader dataset are the next steps toward a more trustworthy comparison.

Replace “looks good” with a question you can test

A few successful prompts can make an AI feature look ready to ship. They cannot tell you how often it fails, which users it serves poorly, or what an improvement costs elsewhere. Jess Wang, a DevRel engineer at Braintrust, opens with teams making release decisions on exactly this kind of impression: engineering says it is ready, or a product manager tries a couple of prompts and likes the answers.

An eval turns that impression into a testable statement. Wang’s illustrative alternatives are concrete: 200 test cases with a 94% pass rate, or a change that improves accuracy by 5% while worsening tone by 5%. These are examples of the decisions an eval can support, rather than results from the coding experiment that follows. The useful starting point is a question about the system.

Selected presentation frame from Agentic Search vs Vector Search for Coding Agents: We Ran the Eval — Braintrust at 93 secondsOpen full source frame
Slide with example evaluation questions, including model selection and cost efficiency.

Those questions can concern model choice, performance across languages or programming languages, cost, company voice, or regressions. A system that works well for English inputs may struggle with Japanese; success on Python does not establish success on TypeScript. Wang also recalls OpenAI’s April 2025 update that became overly agreeable and less truthful. That example gives regression testing a specific target: an apparent improvement in helpfulness can damage another behavior users depend on.

0:120:42
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:12 · section reference included

The four pieces of an eval

An eval needs an input collection, a defined behavior, a quality standard, and a configuration you can compare with another configuration:

  • Dataset. Inputs should include strong reference cases, edge cases, and known failure modes.
  • Task. Define how the system turns an input into an output, usually including the system prompt and chosen model.
  • Scorer. Specify what makes the output good or bad. Deterministic checks, an LLM judge, and human review are different ways to apply that standard.
  • Experiment. A particular dataset, task, and scoring system form one experiment. Changing one of them creates another configuration to evaluate.

Comparison should reach below the aggregate score. Individual rows show which cases improved and which regressed. Braintrust’s interface also supports natural-language queries across experiments and traces—for example, asking for recurring problems or the best-performing experiment. That can help navigate long, multi-turn conversations where hallucination or drift is difficult to spot manually. The talk describes UI, CLI, and MCP access to this analysis.

Selected presentation frame from Agentic Search vs Vector Search for Coding Agents: We Ran the Eval — Braintrust at 281 secondsOpen full source frame
Braintrust experiment comparison interface with a table of cases.
2:423:12
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

2:42 · section reference included

Turn production behavior into the next experiment

Observability supplies the inputs for the next round of evaluation. In Wang’s example, a newly shipped documentation chatbot starts receiving real customer questions. Wrapping application code with the Braintrust SDK brings production logs into the platform, where the team can monitor behavior and build dashboards.

How does a customer conversation become an application change? Sample perhaps 10–20 log rows into a dataset, define the task and scorer, and compare candidate changes. A promising experiment informs a code or system change, which is deployed back into the chatbot. The cycle below makes the feedback path visible: deployment changes the behavior that produces the next set of logs.

Selected presentation frame from Agentic Search vs Vector Search for Coding Agents: We Ran the Eval — Braintrust at 374 secondsOpen full source frame
Diagram connecting an app, logs, a dataset, and evals.

This work crosses job titles. Engineers bring in data and change code; product managers formulate hypotheses about success and help adjust prompts; subject matter experts label examples and define quality in fields such as insurance, medicine, and law; analysts examine the results. “Evals are a team sport” is a practical observation about who can supply the inputs and judgments the loop requires.

How it fits togetherProduction experience feeds evaluation

Customers ask questions about documentation.

The application generates real examples; experiments guide changes; deployment creates new behavior to observe.

5:436:13
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

5:43 · section reference included

Semantic proximity versus following the code

The concrete experiment begins with a request from Wang’s CEO after seeing a Cursor post about search improving coding-agent performance. Online discussion provided a hypothesis worth testing: how would vector search and agentic search compare on real coding tasks?

Vector search converts code or text chunks into embeddings that represent semantic meaning. Wang illustrates the idea with animals and fruit: wolf, dog, and cat sit near one another, while banana and apple form another nearby group. A vector database, such as Qdrant or Pinecone, stores these representations. A query asking where database logic lives retrieves chunks with similar semantic content.

Selected presentation frame from Agentic Search vs Vector Search for Coding Agents: We Ran the Eval — Braintrust at 593 secondsOpen full source frame
Vector-space diagram with labeled points grouped in a three-axis plot.

Agentic search gives an LLM tools such as grep, find, ls, cat, and shell commands. The agent can find a function name, open the file, read the function, follow a call into another file, and continue investigating. Each observation helps determine the next action. This resembles a developer navigating a repository: the route through the code develops as the agent learns what connects to what.

8:429:12
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

8:42 · section reference included

Replay real bugs through the same harness

Microsoft’s TypeScript Go repository supplied the dataset. The team selected merged pull requests with “fix” in the title and checked out each parent commit, returning the repository to a state where the bug still existed. Claude then generated a bug-task description from the difference between buggy and fixed code. The resulting task asked the agent to locate the faulty logic using either retrieval or agentic search. The dataset contained approximately 20 PRs.

Both conditions used Claude Code as the harness. Agentic search required little extra implementation because it was the default behavior. The vector-search condition required active restrictions: otherwise the agent could fall back to ordinary repository exploration and blur the comparison.

Two controls enforced that restriction:

  • Prompt instruction. Explicitly tell the agent not to use agentic search.
  • Tool restriction. Use the disallowed-tools flag to block tools it would normally use for that exploration.

The second control matters because an instruction alone does not prevent a tool call. Holding the harness steady while restricting the available search behavior made the experiment’s intended difference concrete.

Selected presentation frame from Agentic Search vs Vector Search for Coding Agents: We Ran the Eval — Braintrust at 781 secondsOpen full source frame
Slide titled “Implementing Vector & Agentic Search” with numbered steps and code.
11:1211:42
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

11:12 · section reference included

Make subprocess behavior visible before interpreting results

The Claude Code runs executed as subprocesses, and their traces initially became orphaned. The parent trace showed a “Run Claude Agent” step, but it did not expose what happened inside that process. That left a crucial gap: a final score could show success or failure without explaining the search that produced it.

Passing parent span IDs through environment variables connected the subprocess activity to the parent trace. After that change, the trace showed individual turns, LLM calls, and terminal commands inside the run. The observable change was from a single opaque subprocess step to an inspectable sequence of actions—the information needed to debug the search behavior later.

Selected presentation frame from Agentic Search vs Vector Search for Coding Agents: We Ran the Eval — Braintrust at 874 secondsOpen full source frame
Trace interface showing nested run activity and command entries.

Scoring used the repository’s own test suite. A run that passed received 100%; a failure received 0%. This binary rule gave both search conditions the same executable success criterion. It also separated two questions: whether the run passed, and what work the agent performed along the way.

13:1213:42
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

13:12 · section reference included

Same accuracy, four times the cost—and a missing connection

The two conditions reached the same reported accuracy, while vector search cost about four times as much. This is a result for the tested implementation and dataset: Wang later describes the vector implementation as relatively weak, the sample as approximately 20 rows, and repeated trials and additional repositories as important improvements. The comparison does not establish a general ranking of the two search methods.

The traces gave the cost difference a plausible mechanism. Retrieved chunks often omitted imports, tool calls, or surrounding calling code. A chunk could be relevant to the bug while still leaving the agent unable to understand how that code participated in the larger behavior.

One dataset row makes the distinction concrete. The vector agent performed 26 searches and still could not piece together how three functions interacted across the file. Repeated retrieval brought it near the code without resolving the relationship among those functions. Agentic search could grep a function name, read the full function, identify a flaw in its logic, and follow the calling logic into another file. Reading the surrounding code supplied the next investigative step.

Where did the extra work accumulate? The comparison below follows the search behavior described in the traces. The vector path repeatedly returned another chunk while the missing relationship remained unresolved. The agentic path used a discovered call as a route into further context. Wang’s phrase “connective tissue” captures the useful difference: proximity found relevant pieces; navigation helped connect their behavior.

That unresolved context drove more searching, which drove more LLM calls. Those calls accumulated cost even though aggregate accuracy ended up the same. The traces therefore changed the interpretation of the score: equal success rates concealed different amounts of work, and the expensive condition spent much of that work trying to reconstruct context.

Compare the ideasRepeated retrieval versus following a call

Semantic search finds nearby code.

The observed vector-search example accumulated 26 searches without resolving three functions’ interaction. Repository navigation exposed surrounding logic and a route into another file.

14:4215:12
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

14:42 · section reference included

What a stronger next experiment would change

The ending gives three practical improvements to the experiment:

  • Repeat each task. LLM behavior varies between runs; multiple trials make the resulting score more trustworthy.
  • Improve vector retrieval. A stronger implementation would test whether the observed context problem persists beyond this particular setup.
  • Broaden the dataset. More cases and different repositories would extend the comparison beyond approximately 20 tasks in one codebase.
Selected presentation frame from Agentic Search vs Vector Search for Coding Agents: We Ran the Eval — Braintrust at 1031 secondsOpen full source frame
Slide listing multiple trials, improved vector search, and an expanded dataset.

The useful achievement is an end-to-end evaluation that connects a real question to real bugs, controlled tool access, executable scoring, and inspectable behavior. Its most actionable finding comes from combining the score with the trace: when an agent keeps searching, examine whether each result gives it enough context to choose a productive next step.

16:4217:12
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

16:42 · section reference included

Resources

Read the complete timestamped transcript
  1. 0:12

    Everyone, I'm Jess. I'm from Braintrust. I'm a DevRel engineer at Braintrust, and in this talk today, I'm gonna be teaching you guys what evals are, why they're important, and then I'm going to be talking through a relatively complex real-life eval that I did to try to solidify some of the concepts I'm about to teach you. Uh, before we get started, has-- Can you just raise your hand if you've heard of evals before or, like, would confidently say that you know what they are? Okay, cool. Pretty much everyone here. Makes sense.

  2. 0:42

    Um, cool. So I'm just gonna start off with the basis of what evals are, why they're important. Um, I've been on, um, calls with customers where they've said that they've shipped features because either the engineering team told them that it was ready, or maybe a PM said that they tried out a couple of prompts, and they said it looked good. Uh, and this is what you don't wanna hear 'cause essentially, these teams are making ship decisions based off of vibes, uh, which is not good. Uh, and essentially, what you wanna hear is more something

  3. 1:12

    like, "Oh, I ran, um, two hundred different test, test cases, and ninety-four percent of them passed, and that's why we're shipping this feature." Or be able to say nuanced things like, um, you know, "I shipped this feature that increased accuracy by five percent but actually decreased tone by five percent." So when people ask me what evals are, I think the best way to answer it is that it's just a way to answer questions about your AI system. So for example, uh, which LLM is the best choice for our needs? Especially

  4. 1:42

    when there's new models dropping all the time, it's good to be able to have data to back up your decisions. Uh, how does AI perform across diverse real-world scenarios? So for example, your system could do really well with English outputs but not Japanese-- or English inputs but not Japanese, or types-- uh, do really well with Python but not TypeScript. Um, a big one is cost efficiency, so how can we achieve high performance without excessive costs? This is definitely becoming more topical now. Uh, brand consistency,

  5. 2:12

    so does AI reliably reflect our company's voice and standards? Uh, feedback, are we learning from users and improving iteratively? And then one of the biggest ones is how do we know when something breaks or gets worse? So a real-world scenario that I like to use a lot is, uh, about a year ago, so April of twenty twenty-five, um, OpenAI actually shipped a model change and model update, um, that essentially made their model more-- was supposed to make it more helpful but actually made it sycophantic and

  6. 2:42

    therefore, uh, too agreeable and therefore less truthful. So this is a really good example of where evals can help you catch that, um, especially because, uh, AI systems are so complex and undeterministic that it's really important to have these eval systems in place at your company or whatever you're trying to build. So we're gonna get down into more of the specifics of what evals are comprised of. So evals comprise of four main things. Uh, the first thing is a dataset.

  7. 3:12

    So this is the data that you're passing into your AI system, and it usually includes things like, uh, your golden standard, uh, cases, your edge cases, your failure modes. Um, then once you have your dataset, you want to write a task. So a task defines how you want your AI system to behave when it takes in the input and produces an output. So concretely, this usually includes a system prompt and then the model that you're picking to execute this task.

  8. 3:43

    Then the next thing you need to do is you n-need to write a scoring system that will tell you whether the AI's output is good or bad. Um, and so usually it's, uh, maybe the PM or engineer that is writing the criteria that comprises that goodness and badness. Uh, and you might have heard of terms like, uh, deterministic scoring, LLM as a judge, human-in-the-loop. Those are different types of scoring, um, frameworks. Uh, but yeah, that's-- So you write your scoring system. And then the last thing you do, which is the most fun part of writing evals, is

  9. 4:13

    you experiment. So every configuration of a specific dataset, a specific task, and a specific scoring system equals to one experiment. And as you are testing out different hypothesis for how to make your system better, you're tweaking either your dataset, your task, or your scoring system, and these all lead to different experiments. So once you have your different experiments, then you start to compare them. So you might run experiment one and experiment two, and you can see whether there's

  10. 4:43

    been improvements or regressions. And you can also, uh, click into specific rows to see exactly which test cases have changed. Uh, this screenshot is from the Braintrust UI. Another thing that you can do is you can use, uh, natural language to query across your traces and experiments. Um, so for example, you can ask it things like, "Summarize my experiments for me. Um, highlight the problems in these experiments." Um, I find this especially useful when you're running experiment-- or you're running, like, um,

  11. 5:13

    tra-- you see traces with, like, long multi-turn conversations. It's good for, um, helping you catch things like hallucination or drift that's kind of hard to catch manually. Um, you can also-- If you're running different experiments, you can ask it, like, which performance or which experiments performed the best of, out of all of these experiments. Um, so this is another way you can, like, find patterns in your traces. Um, we also-- This is using, like, the UI in Braintrust, but we also have a CLI and an MCP

  12. 5:43

    server if you prefer to do it that way as well. Um, and then the last pillar of Braintrust is just, like, regular observability. So you can basically use the Braintrust SDK to wrap your code, um, and then, uh, have your production logs show up in Braintrust and essentially do traditional observability, where you can monitor your production logs and create dashboards and things like that. Um, so kind of to tie this together, I'm gonna walk you through kind of a flow of how you might use observability and

  13. 6:13

    evals to improve your AI system. So we're gonna start on the top left corner where it says app. So imagine you have some sort of AI application. Uh, in this scenario, I'll say you have a, a new docs page that you just shipped with like an AI chatbot built in so that users can ask questions to your AI chat-chatbot about your documentation. So let's say you shipped that yesterday, and then you have customers start to use your AI chatbot. So when they start to hit your AI chatbot, you're gonna start to

  14. 6:43

    get real production logs coming into your system. Once you get those production logs, you can, uh, sample a subset of them, maybe like ten to twenty rows or something, and create a kind of dataset. Uh, once you get that dataset, then we can run evals on that, right? So you can... You have your dataset, you can write your task, you can write your scoring system, and you can run, uh, evals on them. And you can start to tweak your dataset, you can start to tweak your prompt or s- tweak your score and run multiple different evals on that. Once you have your

  15. 7:13

    evals, like run a couple different evals, um, and start to see which eval does the best, you might learn something about your system, and then you're gonna go ahead and make your code change or make a change to whatever AI system that you're using, uh, and then deploy that back into your application. And then the whole flow repeats again because, uh, those changes will show up in your production logs and the whole system happens again. So that's kind of how the flow works as a whole. Um, the last slide I'll show before I get into a

  16. 7:42

    concrete example is... I love this slide because, uh, conceptually, I, I wanna hammer into people that evals are a team sport. Uh, there's so much that goes into creating evals, uh, from getting real world data into your platform, labeling it, developing hypothesis, things like that. So the AI engineer kind of works on obviously making the code changes, but also getting the data into the platform. Uh, the product manager is very important. I think they're, they usually spearhead evals at a company, but they're the ones that develop the

  17. 8:12

    hypothesis of what success looks like, um, and help, uh, tweak, tweak the prompts and things like that. Uh, subject matter experts are really important, especially in like more niche, uh, areas like, um, insurance, medicine, law, and things like that. Um, they can help, uh, label the data or provide a ground truth for what, what excellence looks like. Uh, and then data analysts obviously to analyze the data. So I always think that this slide is, is really useful to look at. Okay. So let's

  18. 8:42

    go into a like interesting real world example of how you might use an eval. Um, in this one, I'm comparing agentic search to vector search. And the story behind this specific eval is, um, is that my CEO, Ankur, sometimes goes on Twitter and sends me articles of things he wants me to eval. Uh, so he sent me this, uh, tweet from Cursor, where Cursor talks about how they used semantic search, also known as agentic search, um, to significantly improve their coding

  19. 9:12

    agent ex-- uh, performance. And he was like, "Jess, can you eval this for me?" And I was like, "Okay. I'm busy, but okay." Um, and so I kinda went on Twitter and I saw that there was a lot of, uh, discourse or discussion around this topic, so I thought it might be interesting to eval. So, uh, we're gonna eval a vector search and agentic search, so I just wanna give a quick, uh, summary of what each one is just so everyone's on the same page. So vector search essentially is the process of taking, uh, data like, uh, code or

  20. 9:42

    text and turning them into embeddings, where embeddings basically represent the semantic meaning of that chunk of text or code. So, uh, if you look at the graph here, it's a very good visual example where, um, the embeddings, uh, wolf, dog, and cat are very close to each other 'cause semantically they all are animals. And then, uh, the embeddings banana and apple are also close to each other 'cause they're fruit, but then, uh, they're far apart from each other because fruit and animals are different things. Um, and so what happens

  21. 10:12

    is in, you know, um, in a real vector search implementation, you would take these embeddings, uh, and store them in a vector database, so something like Qdrant or Pinecone or something like that. And if you come up with a query like, uh, for example, where does the database logic in my code base live? Uh, what would happen is the vector database would return back the chunks of, um, embeddings that most closely semantically match the words in your query. So the, the chunks that are closest to things

  22. 10:42

    like database or logic. So that's how vector search works. Agentic search is slightly different. Agentic search works more like a human would. So essentially with agentic search, uh, you give an LLM tools like, uh, grep, find, ls, cat, um, bash commands, and it basically explores, explores your code base like a human would. So it would, uh, look up a function name, open a file, read the file, follow that function call to another file, and then open that file and read that file, and then continue

  23. 11:12

    until it's found what it's ne- what it needs to. So this is, uh, much more similar to how a human would kind of go through a code base and, and look up-- look for certain pieces of logic. Okay. So then how did we eval this? So the first part of creating eval is creating your dataset, right? So to do this, we basically used, uh, Microsoft's TypeScript Go repo, uh, which is open source. And what we did is we basically found all the PRs, um, the merged PRs with the title fix in

  24. 11:42

    them. And, uh, we basically checked out the parent commit. So essentially we checked out the commit to the point where it still has the bug in it. And then we used Claude to generate a, like a bug task or like a linear task based on the diff, so based on the diff between the buggy code and the fixed code. And then we fed that into an LLM, uh, and had that LLM either use RAG search or agentic search to find the buggy code, to locate where the logic, the buggy logic lives.

  25. 12:12

    So that's how we created our dataset. We had like maybe We, we found like 20 different PRs, um, and, and ran the eval on that. Uh, in terms of our task, uh, we actually had to implement vector search and agentic search obviously. Uh, implementing agentic search was really easy because Claude Code uses agentic search by default, so we just set Claude Code on these tasks. Um, implementing vector search was slightly more complicated because we still wanted to use the Claude Code

  26. 12:42

    harness, um, to execute the vector search just to keep the experiment consistent, but we had to basically prevent, uh, Claude Code from defaulting to agentic search. So we did two things. Uh, the first thing is that we explicitly said in the prompt instructions not to use agentic search. Um, but if you've ever done this before, you know that prompting is never enough. So what we also had to do is use the, uh, disallowed tools flag to prevent, uh, Claude Code from calling certain things, uh, that it would normally use in agentic search.

  27. 13:12

    So that's how we implemented, uh, vector search and agentic search, both using the Claude Code harness. Um, another kind of interesting thing that we ran into is, um, when we were trace, uh, running the traces, we ran the Claude Code, uh, runs as subprocesses, but it essentially mean, meant that the traces were orphaned. Um, and, uh, our solution was basically passing the parent span IDs as environment variables. So if that doesn't make sense, that's fine because I'll show

  28. 13:42

    you via pictures. So this picture that you see here is what the trace looks like when it's running as a orphaned subprocess. So you can see in the highlighted on like the left panel where it says, "Run Claude Agent," um, that's basically in the trace saying, "Okay, it's running the Claude subprocess," but you can't... You don't actually get any visibility into what happens when it's running that Claude process. With the fix that we did, uh, this is what you see now in the trace. So you can see that it's, uh,

  29. 14:12

    calling the Claude Code subprocess, but you can also see everything that's being run in that subprocess. So you can see all the turns, all the LLM calls, uh, every single terminal command that's being called, like grep, bash, ls, find, things like that. Uh, and it gives you visibility into what's happening in this actual run. Uh, and implementing this fix was really important because being able to have transparency into what's happening, um, is, is important for debugging, uh, later on. And scoring was really simple. Uh,

  30. 14:42

    for our scoring, we basically said, uh, because we're using Microsoft- Microsoft's TypeScript Go repo, we basically said if the test, uh, passed the, uh, t- uh, Microsoft TypeScript Go test suite, then it gets 100%, and if it fails, it gets a 0%. So it was very binary scoring here. Uh, okay, so in terms of our results, uh, the summary is basically vector search and agen- agentic search both got the same level of accuracy, but essentially vector search cost, uh, four

  31. 15:12

    times more than agentic search. So when we looked at the actual traces and, uh, kind of manually like went through to try to understand what happened, we had a couple of learnings. So the first thing is that vector search, uh, returns back chunks of code, but these chunks of code often miss things like, um, imports or tool calls or the calling code above it. Um, and so as a result, it's not enough. It doesn't give enough context for how the agent should actually solve the bug. Uh, so one example is in

  32. 15:42

    one of the rows, we saw that the vector agent made 26 searches and still couldn't piece together how three functions interacted across the file. In contrast, agentic search was actually able to like, as I said, work more like a human would. So grep the function name, read the full function, uh, spot the flaw in the logic, and then follow the calling, um, calling logic, uh, into another file. Um, and the way that I kind of like to think about it when I was just going through the traces is it felt like,

  33. 16:12

    um, vector search gave a lot of proximity to the correct code, but agentic search had more of the connective tissue to connect the actual logic of those chunks of code. Uh, and then tying that into the third learning I had is, um, the reason why vector search was so much more expensive than agentic search is because it made more, um, searches. 'Cause it was o- um, c- constantly searching, like returning a chunk of code and then searching again, returning a chunk of code, searching again. And as a result, um, all

  34. 16:42

    of those LLM calls, uh, piled up and made it a lot more expensive, uh, to execute. Um, and this is the last slide, but essentially, uh, you might be sitting here thinking, "Oh, there's so many holes in this eval," which is very fair. Like, there's lots of things I could've done. Uh, I, uh, you know, there's lots of best practices when it comes to running evals, like running multiple trials per task. Uh, especially because LLMs are so non-deterministic, it's good to run multiple trials to be able to trust the score that you're getting back a lot

  35. 17:12

    better. Um, my vector search implementation was relatively weak. Uh, I could've done a better implementation. Um, also expanding the data set. Um, I think we only ran this on like 20 different rows, so it would've been better to expand this, uh, to even different, um, different repos or things like that. Um, but the, the point of this talk is not to create the best, uh, off- like the best eval ever, but it's more to illustrate to you the concept of h- the kinda how to create a eval end to end.

  36. 17:42

    So hopefully I've done that, um, and that is the end of my talk. Um, so hopefully you learned a little bit of what evals are, why they're important, and how you might do one in real life. Um, and yeah, I think that's it. So thank you for listening.