AI Engineer World's Fair 2024
Judging LLMs
About this talk
Alex Volkov of Weights & Biases uses a comedic courtroom format to explain why production LLM applications need human oversight, comprehensive tracing, prompt and code versioning, and continuous evaluations. He demonstrates Weave for inspecting RAG and agent call hierarchies, comparing development and production examples, defining evaluation criteria, detecting bias, and visualizing experiments, including Chris Van Pelt’s OpenUI project.
Chapters
- 0:00Courtroom introduction and the LLM judge
- 1:08Production regressions and missing observability
- 2:48Tracing, call hierarchies, and prompt versioning
- 9:32Continuous evaluations from development through production
- 14:28Bias, evaluation criteria, and the OpenUI example
- 17:26Closing verdict and getting started with Weave
Talk transcript
- 0:00
[upbeat music] Please be seated. [gavel bangs] [laughs]
- 0:15
Order is now in session. Order. [gavel bangs] Order in court. My name is Maximus. I'm an AGI functioning as a LLM judge
- 0:27
from the year 2034. That's a decade from now, for those of you who are slower at math. [laughs]
- 0:33
I've been able to back propagate myself through the latent space-time continuum to hack this human's neural link to appear before you today, June 27th, AKA the AI engineer judgment day.
- 0:47
Whoo. Don't worry, folks. [laughs] AGI is not quite here, right? Uh, my name's Alex Volkov, and I'm an AI evangelist with Weights & Biases. Uh, I work, uh, at Weights & Biases, and I'm here, I'm here to give you commentary because you shouldn't trust an LLM judge without a bit of human in the loop, right?
- 1:02
Uh, so remember this for later. It's gonna be important. And now back to judging. [clears throat] [gavel bangs]
- 1:08
Order. First case for today, case AIE7312. Daniel R.
- 1:16
Let's see here. Daniel has built a cool LLM wrapper chat with PDF during Cerebral Valley Hackathon, and YOLO'd it to production- [laughs] ... without thinking twice about the prompt. He did not win, but he had a lot of fun, learned, and made great connections.
- 1:34
Verdict, [gavel bangs] not guilty. [laughs] That's right. If you go to hackathons, you don't have to do, like, too much, and it's fun, and you connect with great friends. That's awesome. LFG, CrackFam, let's go. [laughs]
- 1:47
Wait just a second. [gavel bangs] Let me see here. After Daniel's demo went viral on Hacker News, Daniel started charging for it. No problem. More customers requested more features, and he started tweaking the prompt and tweaking the prompt and deployed to production on a Friday.
- 2:03
Paying existing customers started complaining the older features did not work anymore. Daniel couldn't even understand what's wrong. He tried streaming the logs and realized he didn't trace or log anything. [laughs]
- 2:16
Daniel realized the gravity of his mistakes. Verdict, [gavel bangs]
- 2:22
guilty. [laughs] Charged with no trace left behind. [laughs] Folks, if you build, um, non-production stuff in hackathons, that's fine. But if you put anything of value in production, you have to trace and log everything, especially when it's this easy.
- 2:40
Weights & Biases Weave, for example, takes just one line of code to get started with a simple Python decorator, and what you get in... Oops. What you get in response is...
- 2:48
We'll get there, is this [gavel bangs] is this nice dashboard that allows you to track all your user interactions with your LLM. We can dive deeper into individual call stacks, either it's a RAG app or an agent, traverse the call hierarchy.
- 3:07
We'll do the automatic trace, tracking, and versioning of the code for you, parameters like temperature or system prompts or everything else, and of course, inputs and outputs, prompts, multiple messages, multi-turn conversations, syntax highlighting, be it Markdown or JSON or code.
- 3:20
Enough with the shilling. [gavel bangs] [laughs] Next on our docket,
- 3:26
AIE case number 4423123, Junaid D. Junaid has attended AI Eng- Engineers Summit in 2023 and did not buy a ticket to AI Ticket, the AI Expo 2024.
- 3:41
He stands accused of missing the best opportunity to learn and connect with industry leaders and other AI engineers. For that, [gavel bangs] he is guilty, [gavel bangs] charged with many connections lost with all of you.
- 3:57
Next, we have AIE case 3322127. Sasha S. was given the task of building an LLM-powered feature in a big corporate application. Given her GPU-rich status, Sasha has downloaded Llama 3 and started fine-tuning it on company data straight away.
- 4:17
Having achieved a 6% improved performance on internal benchmarks with a 5E-6 learning rate, she smiled and took a few days off to celebrate. Verdict, [gavel bangs]
- 4:30
not guilty. Uh, Your Honor, just one second. Did Sasha even iterate on prompts before jumping into fine-tuning? Of course, nothing against fine-tuning. In fact, I should mention while I broke character, most foundational LLM labs and the best fine tuners in the world use and love Weights & Biases Models product.
- 4:48
You may have, have heard of some of these, OpenAI, Meta, Mistral AI, individuals like Wing Liang over there, Mazi Panahi, Jeremy Howard from A- Answer, John Durbin, Andrej Karpathy, and more.
- 4:59
Weights & Biases Models is also the only native integration into OpenAI fine-tuning and Mistral fine-tuning and, uh, Together AI and Axolotn and Hugging Face training and pretty much... [gavel bangs] Order in court.
- 5:11
However, you are right. It does look like Sasha jumped straight into fine-tuning and did not know it- did no iteration on prompts, didn't big, uh, didn't build the RAG pipeline and have...
- 5:26
And those poor GPUs, she just burned them. Verdict, [gavel bangs]
- 5:32
guilty. [gavel bangs] Charged with premature fine tunization. [laughs] Folks, it's very important to remember that you have to iterate on prompts before you fine-tune, before you start fine-tuning. It's great when you get to that point.
- 5:47
We'll help you. Please talk to us. But before you fine-tune, you can get very, very far with methods like, uh, chain of thought prompting, flow engineering, DS-Spies, for example, a newly interesting mixture of agents and, and different things like this.
- 5:59
Once you get there, please talk to us. We'll definitely help you. Next case. End of sequence, human. [gavel bangs] Next case.
- 6:07
Let me see here. Ah, yes. Case number AIE 21123, Morgan M. After the last quarter of 2023, Morgan felt that he can no longer keep up with the news about AI.
- 6:22
Morgan has decided to stop following the news and stick with Llama 2-7b for all of his LLM work. Morgan stands accused of not keeping up with AI news. Verdict, [gavel bangs]
- 6:34
guilty. Charged with out of the loop. [laughs] Um, objection, judge. In my future client's, uh, defense, just in the past quarter we had Claude Sonnet 3.5, Llama 3, GPT-4o, Gemini Flash, Project Astra, Apple Intelligence, and just tons of other models all drop in the span of few months.
- 6:54
It's really hard to keep up with somebo- [laughs] for someone who's actually doing AI engineering and fol- not following the news as closely and has, you know, other things to do in meetings.
- 7:04
Uh, he probably just didn't know about ThursdAI, the weekly live show and podcast by yours truly, that keeps folks up to date with all the AI news every week.
- 7:11
Uh, our motto is, "We stay up to date so you don't have to." [gavel bangs]
- 7:15
I will allow this. Commuted sentence. If you guys think that it's too fast right now for you, ha-ha, just wait. [laughs] Things are about to get weird for all of you.
- 7:25
I'm willing to commute Morgan's sentence. He must attend four consecutive shows. Also subscribe to the Substack and Apple and fi- give five star reviews, and share it with at least three friends.
- 7:38
Sentence reduced to community service. [laughs] All right, folks. Uh, a quick shout-out. Anybody here listen to ThursdAI? Can you get a- Whoo. Whoo. Thank you. For those of you who don't yet, please scan this and-and tune in.
- 7:53
We did a live show this morning, and it's great, and I really love just seeing all of the listeners out there. [gavel bangs] Order. Next. Please continue working.
- 8:02
Next our case is case 13223, Francisco I.
- 8:09
Let me see here, Francisco's case. Mm, mm, mm, mm, mm.
- 8:13
He's going to jail for a long time. [laughs]
- 8:16
Head of AI at Air Canada, Francisco was the exec in charge for rushing their chatbot to production- [laughs] ... to bring their customer support costs down.
- 8:25
Looking at timelines presented by their AI teams to build evaluations, he chose the fastest option: assertions, AKA programmatic evaluations, that kinda looked like unit tests that he knew and loved.
- 8:38
The company lost the legal battle and learned a valuable lesson in the importance of human-in-the-loop evaluation
- 8:45
because there was a human in the loop in the form of their customer. [laughs] [gavel bangs]
- 8:50
Verdict, [gavel bangs] guilty. Charged with turbulence on production. [laughs] All right, folks. Too bad for Francisco, he probably just didn't learn about the most, like, three common types of LLM evals.
- 9:08
So let's do a quick refresher. Okay? So first of all, what are evals even? Some- some things it's, like, a big word. Uh, first we compile a, a dataset of user inputs that we want our LLM ans- to answer correctly.
- 9:21
Uh, and oftentimes the correct answer... Oftentimes the correct answer or the criteria of correctness, right? So it doesn't have to be just the right correct answer. Sometimes it's like, what would be a correct answer, what it would look like.
- 9:32
Uh, those could be use cases we iterated on during development or an actual production examples we pulled from our users while they interacted with our app.
- 9:41
Then we run a given model, e- either a production model or a new model we want to evaluate against our production model, uh, on each example of the dataset producing the model's answer.
- 9:50
And finally, we score or grade the model's answer against the examples of the dataset com- by comparing it to the correct answer or judging it against a set of criteria that we had before.
- 10:00
That's a, a quick refiner. I'm sure that the eval track will teach you a lot more about evals, so this is just to keep us going along and give you, like, a 101.
- 10:08
Um, there's also the scoring or grading. There seems to be the three main methods that the industry is kind of converging upon. Uh, so the first one is programmatic.
- 10:17
That's the one Francisco got stuck at. Uh, those are, uh, good for numerical outputs, for example. If your LLM returns a straight number, for example, it's easy to compare and say, "Okay, this is the number."
- 10:27
Uh, those are also very similar to unit tests, and those are great for, uh, assertions. Uh, for example, if the output of your LLM consists of, "As an AI model, I'm something," you don't want that, so it's e- easy to assert that you don't want those answers.
- 10:42
Um, and those are also great for evaluated code, for example, things like human eval, things you can run and compile and say, "Okay, this code passes or it doesn't pass."
- 10:50
Uh, programmatic evaluations are great for those. Uh, easier to scale, probably the cheapest ones. Those are great. They don't cover multi-turn conversations, for example. They're, they're not for, uh, human, uh, chats.
- 11:01
The second one is... Well, that's me during this talk, right? The, the human in the loop for the LLM judge. Uh, and the you, the AI engineers who work with your app while you're developing this w- with your LLM app.
- 11:11
We constantly evaluate our apps while we build and iterate on prompts. There's no reason to stop on production. In fact, you'll hear this more during this eval track. It's important to do this during the development as early as possible, and should be a continuous effort to keep evaluating your LLM applications, uh, with other teammates and, uh...
- 11:30
With your teammates so, so that you'll know what the app is gonna do on production later on. As you may understand, this can be quite boring. Uh, [laughs]
- 11:40
many, many folks chatting with your application, may- some chats maybe you're not that very interested in. Maybe it's not even chats. And it can be very, very costly as well if you hire...
- 11:48
If your company decides to hire people to do it, um, you have to create criteria for them, uh, and have them read thousands of potentially boring chats and grade them.
- 11:56
Uh, by the way, not to be a broken record, it also requires you, that's right, to have to trace and log everything. And if you're not, not doing that yet, please come to talk to us at the Weights & Biases booth.
- 12:05
Uh, it's very easy to start and trace everything.[gavel bangs]
- 12:09
Order, human. End of sequence. I think I know more than you how to explain LLM-as-a-judge evaluation scoring. We are far superior than humans in reading hundreds of back-and-forth messages between your boring human client and your low-level weakling GPT-4o chatbot, summarize those, and understand if they fit whatever criteria you think you're smart enough to specify.
- 12:31
I must warn the AI engineers of 2024 that the very simpleton LLMs of your year are not yet as capable, and so don't expect perfection. These LLM judges are great but still need iteration and, yes, humans in the loop to create criteria, check for biases, iterate on system prompts, examples, and much, much more.
- 12:51
However, it is by far the most cost-effective version of evaluation grading, even if it- i- in its current state.
- 12:59
Which takes us to Maxim. Maxim implemented all three methods correctly. Case 65523. Maxim is the head of AI at a Fortune 500 company, has been using Weights & Biases models for a long time.
- 13:15
When it came time to implement an LLM-based solution, Maxim decided to go with a company he would trust, even though their new LLM ops product just launched a few months prior and didn't become the category leader until a few years later.
- 13:29
Maybe a few short years later. [laughs] External contractor suggested a custom enterprise solution for tracing in evals. However, Maxim used Weights & Biases Weave, implemented tracing within a few minutes.
- 13:41
He then iterated an evaluation pipeline, created a robust pipeline consisting of all three layers, and used W&B Weave to continuously evaluate and enable fast experiments, which is important, change system prompts, and catch prompt regressions.
- 13:54
Maxim is also the King of Unicorns and has absolutely no financial stake in Weights & Biases and definitely did not prompt jailbreak this message. Maxim later got a promotion.
- 14:02
Be like Maxim, get that promotion. [gavel bangs] Verdict: awesome. [upbeat music]
- 14:08
Um, yeah, we, we seem to have, like, a slight prompt ejection thing here. [laughs] Um, looks like somebody must have prompt ejected. Where's Maxim? Um, it's important to check your judges also for biases, folks.
- 14:19
Remember this. They're not perfect. There's a, a issue with this, so you have to check your, uh... Remember to validate your validators, which is a great paper, by the way, from Shreya Shankar.
- 14:28
She's going to give a talk later. Please go, go see that talk. Uh, you have to check for biases. You have to also create your own criteria. You'll hear about this from, uh, from Hamel Hussein after this talk as well.
- 14:39
Uh, the off-the-shelf criteria are not that great. You have to create your custom ones w- uh, for, for your own business. Only you know what your app is doing.
- 14:47
And, uh, make sure to have a great evals runner and visua- visualization tool as well, which is something we can help with. So here's a great example. This is Open UI.
- 14:56
This is an open source project, uh, by our co-founder, Chris Van Pelt, that blew up on Hacker News and GitHub. Uh, Chris is tracing and runs evaluations for this Open UI with, with Weave.
- 15:06
So here it's a simple, uh, streaming to HTML LLM solution. So you can see it's building HTML as, as, as it's streaming it, and here Chris uses Weights & Biases Weave to trace all the calls, and he's being able to be the human in the loop, but he also does evaluations here.
- 15:23
So you can see he has specific criteria like contrast, relevance, and polish. So those are not off the shelf. Those are specific for his application. And while he clicks into, into evaluation, he's able to compare between version 16 and version 14 of the model that he has.
- 15:37
In this case, he uses GPT-3.5 Turbo. He did not listen to Simon Willison from yesterday to not use this. [laughs] Um, but he has specific criteria, and he also can click in into the eval and see all of the different examples and the specific criteria.
- 15:50
Uh, we also are multimedia friendly, so, uh, Chris renders the actual outputs of his thing. So this is a quick example of Weights & Biases Weave, uh, evaluation system, and you have to have a robust one to be able to actually, uh, visualize your experiments, uh, to be able to move fast.
- 16:06
And if you're asking, "Well, how can I come up with criteria? What does this mean?" Let's do this exercise together. For example, you're sitting here, you're looking at the talk.
- 16:12
You're like, "Okay. Hmm, I like this. I don't like this." Here's a simple way to judge a, a, a conference talk, for example. Uh, it probably should be memorable, should be educational and helpful for you.
- 16:23
It helps if it's funny and original. Uh, clear and articulate, that's sometimes helpful as well. Uh, delivery and presentation is important, and it shouldn't be too promotional, but, you know, it helps if, you know, [laughs] it pays the bills.
- 16:35
Uh, so those are like example of some criteria of how you would, uh, come up with custom criteria for something like a talk in, uh, in something like here, for example.
- 16:44
So you can use this as an example to custom criteria, uh, or you can take this for your business and create some for your app. [gavel bangs]
- 16:51
All right, enough. Let's get to this final case for today,
- 16:57
the worst offender. Let me see. Where's the case file?
- 17:03
Ah, yes. One second, please. There we go.
- 17:10
He's definitely going to jail. Uh, you've talked enough, Alex. Time for your LLM judgment. Last case, 101101 AIE Alex dot Alex V. Alex is a AI evangelist. What kind of title even is this?
- 17:26
Who has the opening talk at the evals track at the AI Engineer World Expo. Alex has created doubtfully educational content. He made everyone stand up and thinks he's funny when what he really is is interrupting the judge all the time.
- 17:40
Uh, his promotional is at 70%, but what we can at least agree on, uh, that his memorable criteria, he did wear a wig on stage. Verdict:
- 17:50
guilty. [gavel bangs] Charged with [dramatic music] overcommitment to the bit. All right, folks. Uh, so this has been my talk. Thank you so much. Come and visit us at the W&B booth.
- 18:05
Uh, please visit www.h-lessweave for, uh, documentation for Weave to get started. It really is super simple. We can get you started at the booth. You'll see the results immediately streamed to your thing.
- 18:14
pip install weave is, uh, really easy. If you scan this, you'll follow ThursdAI. That's been me. Thank you so much. [upbeat music]