AI Engineer World's Fair 2025
Shipping AI That Works: An Evaluation Framework for PMs
About this talk
Arize AI product manager Aman Khan leads an interactive workshop on building repeatable evaluation workflows for AI products. Drawing on experience at Cruise and Spotify, he demonstrates prompt editing, experiment comparisons, agent tool-call evaluation, and response-quality assessment. Audience questions explore LLM-as-a-judge variance, temperature, BERT and ALBERT, human-labeled data, evaluation-team hiring, and shared ownership of prompts and evaluations across product and engineering.
Chapters
- 0:00Introduction and AI product-management background
- 5:35Evaluation datasets and interactive product demonstrations
- 26:24Agent tool calls, prompt iteration, and evaluation setup
- 43:51Judge variance, temperature, experiments, and response quality
- 56:27Audience questions on evaluation teams, models, and labeled data
- 1:10:52Cross-functional AI development and workshop closing
Talk transcript
- 0:00
[upbeat music] All right. Uh, nice to see everyone here.
- 0:17
Um, my name's Aman. I'm an AI product manager at a company called Arize. Title of the talk is Shipping AI That Works: An Evaluation Framework for PMs. Uh, it's really gonna be a continuation of some of the content we've been doing with, you know, some of the, the PM folks, like Lenny's podcast.
- 0:34
I guess just quick show of hands, how many people listen to Lenny's podcast or have read, read the newsletter? Awesome. Okay. We're gonna do a couple more, like, audience interaction things just to, like, wake up the room a bit.
- 0:42
So how many people in the room are PMs or aspiring PMs?
- 0:48
Okay, good. Good handful of people. How many of you consider yourself AI product managers today? Okay, awesome. Wow, th- there's more AI PMs than there were regular PMs. That's interesting.
- 0:59
Um, that's [laughs] usually that's, it's a subset, but maybe I need to start asking the questions in a different order. Um, cool. Well, that's great. Uh, so what we're gonna be doing is, you know, um, I'll go ahead and just do a little bit of an intro about myself, and then we'll kind of cover some of the, the
- 1:14
frameworks that I think are really powerful for AI PMs to kind of get to know as you're building AI applications. So a little bit about me. Um, I, you know, myself, I have a technical background, so I actually started my career in engineering, uh, on actually working on self-driving cars at Cruise.
- 1:30
Um, and actually, while I was there, I ended up becoming a PM for evaluation systems for self-driving back in like twenty eighteen, twenty nineteen. Um, after that, I went to Spotify to work on the machine learning platform and work on recommender systems, so things like Discover Weekly and search, things like using embeddings to actually make the end
- 1:51
product experience better. And fast-forward to today, I've been at Arize, uh, for about three and a half years, and I'm still working on evaluation systems. Instead of self-driving cars, it's sort of self-writing code agents.
- 2:03
Uh, and Spotify is actually [REDACTED:generic_id] of our customers, so we get to work with some awesome, uh, you know, ex-- Actually, fun fact, I've actually sold Arize to all of my previous managers. [laughing]
- 2:13
So, um, so fun fact there. Uh, but we get to work with some awesome companies like Uber, Instacart, Reddit, Duolingo, so a lot of really tech-forward companies that are building around AI.
- 2:24
Uh, and we actually started in the sort of traditional ML space of ranking, regression, classification type models and have now expanded into gen AI and agent-based applications as well.
- 2:36
Uh, what we do is make sure that those companies, our customers, when they're building AI applications, that when those agents and applications actually work as expected. And it's actually a pretty hard problem.
- 2:47
A lot of that has to do with, uh, terms that we're gonna go into, like observability and evals. But I think more broadly, the space is just changing so fast, and the models, the tools, the infrastructure layer is changing so fast that for us, it really is a way for us to learn about the cutting edge.
- 3:04
Like, what are the leading challenges with use cases that people are building? And try to build that into a platform and product that benefits everybody.
- 3:14
Um, so what we'll cover, we're gonna cover what are evals and why they matter. We'll actually build an AI trip planner, uh, with actually a multi-agent system. This part is ambitious, bullet number [REDACTED:generic_id].
- 3:23
I'm gonna be honest here. Uh, we were trying to push up the code right before, so it may or may not work, but we'll give it a shot, and that'll be the interactive part of the workshop.
- 3:31
And then we'll actually try to evaluate that AI trip planner prototype that we're gonna build ourselves.
- 3:37
Uh, actually, another quick show of an, of hands for the room. How many people have heard of the term evals before? Okay, yeah, I guess it was in the title of the talk, so that's kind of redundant.
- 3:46
How many people have actually written an eval before or tried to run an eval? Okay, a good number of people. Um, that's awesome. Well, the, what, what we're gonna do is actually ta- try and take that a little bit of a step further, go from writing an eval for an, an LLM-as-a-judge system.
- 4:01
And if you've never written an eval, don't worry, we're gonna cover that too, but try and take that [REDACTED:generic_id] step further and make it a little bit more kind of technical interactive as well.
- 4:10
Okay, so who is this session for? Uh, I like this diagram because, um, you know, L- Lenny and I have been kind of working together a little bit more on content for educational content, mostly for AI product managers.
- 4:23
And I kind of put this up. I made like a little whiteboard diagram for him, and I'm like, "I think this is really how I view this space," which is like there's this, there's this...
- 4:31
You know, y- you may have seen this diagram for, like, the Dunning-Kruger effect, and that's kind of what came to mind here, which is as you're kind of moving along the curve, maybe you're just getting started, you know, with, "How do I use AI?
- 4:41
How does AI fit into my job?" I think we were all here, to be honest, a couple of years ago. Like, you know, just to be completely honest, I think for the people in the room, especially PMs, I think we all feel that the expectations of the product management role are changing.
- 4:56
That's why this concept of an AI PM is sort of emerging. The expectations from our stakeholders, from our executives, from our customers. It feel, I feel, I don't know about if other people feel this way.
- 5:06
I definitely feel like the bar has been raised in terms of what's expected to be delivered, right? Especially if I'm working with an AI engineer on the other end, their expectations of what I come to them with in terms of requirements, in terms of specifying what the agent system needs to look like, it's changed.
- 5:23
It's a step function different even than for me, even as someone who was like a technical PM before. And so I kind of felt myself go along this journey, which is ironic given that I work at an evals company.
- 5:35
You'd think I was, like, on the end of the curve. But really, I kind of went through this journey, you know, same as most of you, which is trying to use AI in my job, trying AI tools to prototype and come back with something that's, you know, a little bit higher resolution for my engineering team than, like,
- 5:50
a Google Doc of requirements. Once I had those prototypes and I'm like, "Hey, let's try to build these new UI workflows," the challenge then became, how do I get a product into production, especially if my product has AI in it, has an LLM or an agent?
- 6:05
And that's where I think, m- you know, that's really where that, like, confidence slump sort of hits, and you kind of realize there's a lack of tooling, there's a lack of education for how to build these systems reliably.
- 6:17
And- Why does that matter at the end of the day, right? And the really important takeaway from the fact that LLMs hallucinate, we all know that they do, is you should really look at the top [REDACTED:generic_id] quotes here and think, okay, well, we've got Kevin, who's Chief Product Officer at OpenAI.
- 6:33
We have Mike at Anthropic, CPO. This is probably like ninety-five percent of the LLM market share, and both of the product leaders of those companies are telling you that their models hallucinate and that it's really important to write evals.
- 6:48
This, these quotes actually came from a talk that they were both giving at Lenny's conference, uh, you know, earlier, like in November of last year. And so when the people that are selling you the product are telling you that it's not reliable, you should probably listen to them. [laughs]
- 7:04
Uh, on top of that, I mean, like you have Greg Brockman, similarly founder of that company. Um, you have Gary, who's, you know, evals are emerging as a real moat for AI startups.
- 7:13
So I, I think this is sort of [REDACTED:generic_id] of those pivotal moments where you realize, hey, people are starting to say this for a reason. Why are they saying that?
- 7:21
Well, they're saying that because a lot of the same lessons from the self-driving space, um, you know, kind of apply in this, this AI space. Okay, another audience question.
- 7:29
How many people have taken a Waymo? I kind of expect that [REDACTED:generic_id] to be pretty high. Okay, we're in San Francisco. If you're visiting from out of town, take a Waymo.
- 7:36
It is a real world example of AI. It's, it's a, it's an example of AI in the real physical world, and a lot of how those systems work actually apply to building AI agents today.
- 7:48
All right, we'll do a bit of a zoom out, then we'll get into the technical stuff. I see laptops out, so we'll definitely get into, you know, writing some code and trying to get hands-on.
- 7:55
But just to do a bit of a recap for folks, um, what is an eval? Uh, I, I kind of view this as like it's very analogous to software testing, but with some really key differences.
- 8:05
Those key differences are software is deterministic, you know, [REDACTED:generic_id] plus [REDACTED:generic_id] equals [REDACTED:generic_id]. LLM agents are non-deterministic. If you convince an agent [REDACTED:generic_id] plus [REDACTED:generic_id] equals three, it'll say like, "You're absolutely right.
- 8:15
[REDACTED:generic_id] plus [REDACTED:generic_id] equals three," right? So like we've all been there. We've kind of seen that these systems are highly manipulatable. And on top of that, if you build an LLM agent, uh, that can take multiple paths, that's verily di- that's pretty different from a unit test, which is deterministic.
- 8:30
So think about, um, the fact that, you know, a lot of people might, are trying to like eliminate hallucinations from their agent systems. The thing is, you actually kind of want your agent to hallucinate just in the right way, and that can actually make testing it a lot more challenging as well, especially when reliability is, is super
- 8:49
important. And then last but not least, I think integration tests rely on existing code base and documentation. The really key differentiation of agents is that they rely on your data.
- 9:00
Uh, if you're building an agent into your enterprise, the reason that someone is going to use your agent versus something else is likely, it might be because of the agent architecture, but a big part of it will also be because of your data that you're building the agent on top of, and that applies for the evals as
- 9:16
well. Okay, what is an eval? So, uh, I view this into like four parts that go into an eval. It's kind of just an easy like muscle memory thing.
- 9:26
Um, these brackets are a little bit out of line, but, um, the, the idea is that you're setting the role. You're basically telling the agent, "Here's the task that you want to accomplish."
- 9:35
You're providing some context, which is what you see in the curly braces here, and it's, that's essentially like i- it's really just text at the end of the day.
- 9:43
It's some text you want the agent to evaluate. You're giving the agent a goal. In this case, the agent is trying to determine whether text is toxic or not toxic.
- 9:52
This is a kind of a classic example 'cause there's a large toxicity dataset of classifying text that we use, um, to build our eval on top of. But just kind of note that can be any type of goal in your business case.
- 10:02
It doesn't have to be toxicity. It'll be some goal that you've created this agent to evaluate, and then you provide the terminology and the label. So you're giving some examples of what is good and bad, and you're giving it the output of either select good or bad.
- 10:18
In this case, it's toxic, not toxic. I'm gonna pause on that last note because it's really, I think there's a lot of misconceptions, sort of like I'll try and weave in some like FAQs as I hear them come up.
- 10:29
But, um, we'll definitely have some time at the end for questions, and I'd love for this to be interactive, so I'll probably make the Q&A session a little bit longer here for people that have these questions.
- 10:38
But [REDACTED:generic_id] common question we get is, "Why can't I just tell the agent to give me a score or an LLM to produce a score?" And the reason for that is because even today, even though we have like PhD level LLMs, they're still really bad at numbers.
- 10:52
Um, and so what you wanna do is ground, and it's actually a function of like what a token is, what, how a token is represented for an LLM. And so what you wanna do is actually give a text label that you can map to a score if you really need to use a score in your systems, which
- 11:09
we do in our system as well. We'll map a label to a score. But that's, that's like a very common question we get is, "Oh, why can't I just make it do like [REDACTED:generic_id] is good and the five is bad or something?"
- 11:19
Like, you're gonna get really unreliable results, and we actually have some research, um, happy to share it out afterwards, that kind of proves that out, um, on a large scale with most models.
- 11:30
Okay, so that's a little bit of like what is an eval. Um, I should note that, uh, this is, uh, previous slide. Uh, I should note that this is, uh, LLM-as-a-judge eval.
- 11:40
Uh, there's other types of evaluations as well, like code-based evals, which is just using code to evaluate some text, uh, and human annotations. We'll touch on that a little bit more later, but the bulk of this time is gonna be spent on LLM-as-a-judge because it's really the, the kind of scalable way to run evals in production these
- 11:58
days, and we'll talk about why too, um, later on.
- 12:02
Okay, a lot of talking. So, uh, evaluating with vibe. So this, this, this slide's kind of funny because I think, like, everyone knows this term, like, vibe coding. Like, everyone has tried to use, like, Bolt or Lovable, whatever.
- 12:13
And I don't know about you, but this is how I usually feel when I'm vibe coding, which is, like, it kind of looks good to me. Like, you know, you're looking at the code, but, like, let's be honest, how much AI-generated code are you gonna read?
- 12:23
You're like, "Let me just ship this thing." The problem is you can't really do that in a production environment, right? Like, uh, I think all the vibe coding examples are, like, prototyping or, like, trying to build something ha- like, hacky or fast.
- 12:34
So I wanna help everyone reframe a little bit and say, like, yes, vibe coding is great. It has its place. But what if we go from evaluating with vibes to thrive coding?
- 12:45
And thrive coding, in my mind, is really using data to basically do the same thing as vibe coding, like still build your application, but you'll be able to use data to be more confident in the output.
- 12:57
And you can see that this person is a lot happier. Um, so this is using Google's image models. They're scary good, guys. Like, uh, yeah.
- 13:06
Okay, so we're gonna be thrive coding. So slides, um, there's, uh... If you want access to the slides, the slides have links to what we're gonna go through in the workshop.
- 13:15
Um, ai.engineer.slack.com, and then I just created a Slack channel workshop-ai-pm. And I think I dropped the slides in there, but let me know if I didn't. Yeah. Cool. Thank you.
- 13:27
All right, live demo time. So at this point on, uh, I'm-- I'll just be honest. Uh, there's a, a decent likelihood that the repo is-- has something broken in it because we were pushing changes up until, like, this very moment.
- 13:40
If so, and you can unblock yourself, I think there's, like, a requirements thing that's broken. Please go for it, and if not, we can come back at the end and try to help you get unblocked.
- 13:48
And then I promise, after this, I'll, like, push the latest version of the repo up. So if it doesn't work for you right now, check back in an hour.
- 13:54
I'll drop it in Slack. It'll be working later. Um, but yeah, just a function of, like, moving fast. Uh, so on the left-hand side is instructions, which are really-- It's like a, you know, sort of a, a Substack post I made, which is just a free sort of, like, list of, you know, some of the steps we're
- 14:10
gonna go through live. So it's just more of a resource. And then on the right-hand side is a GitHub repo, which I'm gonna open here.
- 14:19
There's actually [REDACTED:generic_id] repos. And I'll kind of talk through, like, a little bit about what we're evaluating and some of the projects on top of that, and then we'll get into, uh,
- 14:30
the weeds here a little bit. Okay, so this is the, the repo. Um,
- 14:39
we-- I built this, like, over the weekend. [laughs] So, you know, it's not, it's not super sophisticated, uh, although it says it's sophisticated, which is funny. But, um, this is- Can you put that in Slack?
- 14:48
Oh, pardon? Can you put that in Slack? Oh, this is not... Okay, so is this not attached to the QR? Okay, I'll just drop this link in here as well.
- 14:55
Let's just, uh, put it in here. Okay, awesome. Oh, thank you. Thanks. Okay, um, so-- And if you have questions, by the way, uh, in the middle of the presentation, just feel free to drop them in Slack.
- 15:08
Um, and then we can always come back to them, and then we'll have time at the end for, um... So feel free to, like, keep the Slack channel going, um, for questions.
- 15:16
Maybe people can try to unblock each other as well. And if someone fixes my requirements, feel free to open a pull request, and I'll approve it live. Um, so, um, okay, so what we're doing is, uh, let's put on-- Let's take off our, like, PM hat of whatever company we're at.
- 15:30
We're gonna put on an AI trip planner hat. The, the idea here is, like, don't worry about the sophistication of this UI and the agent. It's really, like, kind of a prototype example.
- 15:41
But it is helpful for us to kind of take a look at building an application on the fly and try to understand how it works underneath the hood. So the example we're gonna use is actually...
- 15:52
I'll kind of back up a little bit. I basically took this, uh, Colab notebook that I have, um, for tracing CrewAI, and I'm like, "I kind of want an example with LangGraph."
- 16:02
CrewAI probably, if you haven't heard of it, it's like an agent, multi-agent framework. Um, the agent's basically-- An agent definition is, you know, using an LLM and a tool combined to perform some action.
- 16:13
And what I did was I gave this notebook, and I basically put it into Cursor, and I was like, "Give me an example of a UI, uh, based workflow, but using LangGraph instead."
- 16:22
And what we're gonna do is think of instead of building a chatbot, we're gonna take this form, and we're going to use the inputs of this form to build a quick agent system that we're then gonna be using for evaluation.
- 16:34
So this is what I got on the other end, um, which is plan your perfect trip. Let our AI agents help you discover amazing destinations. So let's pick a destination.
- 16:44
Maybe we wanna do Tokyo for seven days, and assuming the internet works, um, we'll see if it does. We're gonna put a budget of $1,000. I'll zoom in a little bit.
- 16:54
And then I'm interested in food, and let's make this adventurous. So I could go and take all of this and try to just put it into ChatGPT, but you can kind of imagine underneath the hood the reason that we might want this as a form or with multiple inputs and, uh, an agent-based system is because we could
- 17:13
be doing things like retrieval or RAG or tool calling underneath the hood. So let's just kind of picture that the system is going to use these inputs to give me, on the other side, an itinerary for my trip.
- 17:26
And, uh, okay, it worked. Okay, this [REDACTED:generic_id] worked. So, um, so here we've got a quick itinerary. Um, nothing super fancy. It's basically just here's what I gave as an input form, and then what the agent is kind of doing underneath the hood is giving me an itinerary for what my morning, afternoon, et cetera look like for
- 17:44
a week in Tokyo using the, the budget I gave it. Uh, this doesn't seem super fancy because it's like I could take this and just put it into ChatGPT, but there is some nuance here, which is the budget.
- 17:57
Like, if you add this up, like, it's gonna be doing math to do accounting to get to $1,000. So it's really keeping that into consideration. You can see it's a pretty frugal budget here.
- 18:06
Um, it can take interests here. So I could say, you know, different interests like I wanna go, I don't know, sake tasting or something, and it'll find a way to work that into your itinerary.
- 18:17
But I think what's really cool here is it's really the power of agents underneath this that can give you really high level of specificity for your output. Um, so that's really what we're trying to show is, like, this is, you know, it's not just [REDACTED:generic_id] agent, it's actually multiple agents giving you this itinerary.
- 18:35
Uh, so I could just stop here, right? Like, I could be like [laughs], "This is, this is good enough. I have some code." For most people, if you're vibe coding, you're like, "Great, this thing does what I want it to do," right?
- 18:44
Like, it gave me an itinerary. But what's going on underneath the hood? Um, and this is kind of where... Uh, so I'm gonna be using our tool called Arize.
- 18:54
We also have an open source tool called Phoenix. I'm just gonna plug that here right now for folks as reference. But this is an open source version of Arize.
- 19:03
It is not gonna have all of the same features as Arize, but it will have a lot of the same setup flows and workflows around it. So you know, just note that Arize is really built for, you know, if you want scale, s- security, support, um, and sort of the, the sort of futuristic workflows in here.
- 19:21
So I've got a trip planner agent, and what I just did, if it worked... Let's see if it did.
- 19:33
And we're gonna... This is li- this is live coding, so, like, very possible something's broken. Um,
- 19:43
okay, I think, I think I broke my la- my latest trace, but you can see what the example here looks like from [REDACTED:generic_id] right before. So what that system really looks like is basically this.
- 19:53
Um, so let's, let's open up [REDACTED:generic_id] of these examples. What you'll see here are traces. Traces are really input, output, and metadata around the request that we just made.
- 20:03
And I'm gonna open up [REDACTED:generic_id] of those traces just as an example here. And what you'll see is essentially a set of actions that the agents, that in this case multiple agents have taken to perform, you know, generating that itinerary.
- 20:19
And what's kind of cool is we actually just shipped this today. Um, uh, it's actually, you know, you guys are the first ones seeing it, uh, which is pretty cool.
- 20:27
Um, this is actually a representation of your agent in code. Um, so you know, literally the Cursor app that I just had up here is basically my agent-based system that Cursor helped me write.
- 20:41
And when I sent it our docs, I s- I literally all I did was I gave it a link to our docs in Cursor and I said, you know, "Write the instrumentation to get this agent," and, and this is, this is how that's represented.
- 20:53
And so we have this new agent visualization in the platform that basically shows the starting point with multiple agents underneath it to accomplish, uh, the task we just had.
- 21:03
So we have a budget, local experiences, and research agent that then go into an itinerary agent, and that gives you the, the end result or the output. And you can, you can see that up here too.
- 21:14
So we have research, itineraries, budget, and local information to generate the itinerary.
- 21:21
So this is, this is pretty cool, right? Like, for, I think for a lot of people it's not im- it, you know, ourself included, it is not immediately obvious that these agents can be super well represented in this sort of, like, visual way, right?
- 21:34
Uh, especially when you're writing code, you think these are just function calls talking to each other. But what's really useful is to see at an aggregate level what are the calls that the agent is making.
- 21:45
And you can see it's a really clean delineation of parallel calls for the budget agent, the local experiences agent, and the research agent, and all of those get fed, fed in to an itinerary agent that summarizes all of the above.
- 22:00
You can also see that up here. Um, so these are what's called, uh, traces, and they consist of, uh, like technically what's called spans. A span, you can think of this as like a unit of work basically.
- 22:12
So there's a time component to it, which is like how long that process took to, to finish, and then like what is the type of the process here. You can see there's three types.
- 22:21
There's an agent. There's a tool, which is, uh, basically being able to use data to perform an action, structure data. And then there's the LLM which generates the output of the, the taking the input and the context.
- 22:34
So this is an example of three agents, actually three agents being fed into a fourth agent to generate the itinerary. So that's really what we're seeing here.
- 22:44
Um, let's go [REDACTED:generic_id] level deeper. So this is cool [laughs], and I think it's useful for, uh, you know, to see like what these systems look like, how they're represented.
- 22:55
To zoom out for a second as a product manager, there's a ton of leverage in being able to go back to your team and ask, "Hey, what does our agent actually look like?"
- 23:03
Right? Like, do you have a visualization to show me of like what the system actually looks like? And then if you're giving the agent multiple inputs, where are those outputs going?
- 23:12
Are those outputs going into, you know, a different agent system? Like, what are the, what does the system actually look like? So that's kind of [REDACTED:generic_id] sort of key takeaway here as a PM.
- 23:22
Um, it was personally very helpful to see, you know, what our agent's actually doing, um, underneath the hood.
- 23:28
Uh, kind of going [REDACTED:generic_id], [REDACTED:generic_id] level deeper here. So we've got this itinerary, uh, and it-- let's take a look at it really quick. So it says, "Marrakesh, Morocco is a vibrant, exotic destination," blah, blah, blah.
- 23:39
It's, it's really long, right? Like, I don't know if I would actually look at this and read it. It doesn't, it's not really like, it doesn't like jump out to me as like being like a good product experience.
- 23:49
It feels super AI generated personally. Um, so what you wanna do is actually think, "Okay, well, is there a way for me to iterate on my product as a product person?"
- 23:59
And to do that, what we can do is actually take that same prompt that we just traced and pull it into a prompt playground with all of the variables that we've defined in code pulled over.
- 24:10
So I've got a prompt template here which basically has the same, um, prompt variables that we've defined in the UI, like the destination, the duration, the travel style, and all of those inputs get fed in here.
- 24:24
You can see- Down below in this prompt playground,
- 24:29
what that looks like. And then you see the outputs of some of the agents in here as well. And then I have the final itinerary from the, the agent that's generating the itinerary.
- 24:41
Okay. So why does this matter? I think a lot of companies have this concept of, um, prompt playgrounds. I think, like, OpenAI has a prompt playground. You've probably heard that term before as well, or maybe even you've, you've used [REDACTED:generic_id].
- 24:53
But I'll ... I urge you to think about when you're thinking about a tool to help you with development, not only is the visualization important of what your stack i- looks like underneath the hood, but being able to take your data and your prompts together and be able to iterate on your data and prompts in [REDACTED:generic_id] interface
- 25:11
is really powerful. Because I can go in and change the destination, I can go in and tweak variables and get new outputs using the same exact prompt I had before.
- 25:20
So that's really, I think, just, just really powerful as a workflow. Um, a thought experiment, uh, for the PMs in the room is, like, when you really think about what this promp- pro- uh, prompt looks like, just think i- should writing the prompt be the responsibility of the engineer or of the PM?
- 25:40
And if you're a product person and you're ultimately responsible for the final outcome of the product, you probably wanna have a little bit more control over what the prompt is.
- 25:50
And so I kind of urge you to think, you know, where does that boundary really stop? Is it, like, I just hand off ... Like, does the engineer know how to prompt this thing better than a product person that might have specific requirements you wanna integrate?
- 26:04
So that's why this is really helpful, um, uh, from a product perspective.
- 26:08
Okay. Yeah, go for it.
- 26:11
How, how do you handle tool calls in this scenario? Like, uh, can, uh, alter the prompt and, like, go play around with it, with this? But, like, go ...
- 26:18
When you actually evaluate that, like, in a real scenario with tool calls-
- 26:22
Yeah
- 26:22
... this probably doesn't have access to.
- 26:24
Ah, okay, okay. So that, that was a good question. Um, so the question from the gentleman in the back is, how do we handle tool calls? And that was a really astute observation, which is, like, the agent has, um, tools in it as well.
- 26:37
And this is, this is a really good point to pause on actually, which is, like, what I did was I pulled over this LLM span with the prompt templates and variables, but there's, there's a world where I might wanna select the right tool and make sure that the agent is picking the right tool.
- 26:53
I'm not gonna go into that in this demo, but we do have, um ...
- 26:59
We do have, uh, some good, uh, material around this on agent tool calling. So we actually do port over the tools as well. This example doesn't because, to be honest, it's a really toy example.
- 27:12
But even if you, if you wanted to s- to do a tool calling evaluation, we, we offer that in the product and, uh, we actually have some material around that.
- 27:20
So if you want, just ping me about it later and I'll send you a whole presentation on that as well. But yeah, uh, good question, which is, like, you don't just wanna evaluate the LLM and the prompts.
- 27:29
You wanna evaluate the system as a whole and all of the subcomponents. Okay, we're gonna keep going. So, so I've got, um, I've got my prompt here. Now, this is cool, but, like, let's try to make some changes to it on the fly, and I will try my best to make this readable for everyone.
- 27:45
But, um, yeah, working with what I got here. So what we're gonna do is I just ... I'm gonna save this version of the prompt, and let's call it AI Eng prompt.
- 28:01
And it's helpful because now I can, like, iterate on this thing, right? So, like, I can duplicate that prompt with a click of a button. I can change the model I wanna use.
- 28:08
So let's say I wanna use 4.1 mini instead of 4.0. I'm gonna change a couple things. Don't, don't be ... Don't worry. Like, in a real world, you're gonna change [REDACTED:generic_id] variable at a time.
- 28:17
But, um, here I'm just gonna change a couple things at the same time just to make this more interactive. But, um, the idea here is, like, let's try to change what the, this actually looks like.
- 28:29
And it says, you know, format as a detailed day-to-day plan. Honestly, I might say, like, like a more important requirement to that is don't be verbose, right? I could say don't be verbose, keep it to 500 characters or less.
- 28:47
Maybe we want this thing to be more punchy. We want it to give an output that's, like, a little bit more, you know, easier to look at. Um, I might be a P...
- 28:54
You know, even if I'm just vibe coding this thing on the weekend, I might wanna get feedback from users that are trying this product out. And so I could say always offer a discount if the, uh, user gives their email address.
- 29:11
It's helpful, right? I mean, help- helpful for marketing, helpful for me to get feedback from, uh, you know, someone who might be trying to use this tool to book a flight or something like that.
- 29:19
Okay, so let's go ahead and hit run all here. And what that's gonna do is actually run the outputs we just, uh, ra- run the prompts we just edited into this, uh, in the playground.
- 29:31
And it might take a second because of the internet.
- 29:36
You pulled, you pulled this in from the exi- [REDACTED:generic_id] of the existing runs, right?
- 29:40
That's right, yeah. So it was exactly the same, um, [REDACTED:generic_id] of these runs.
- 29:43
Yeah.
- 29:44
The ... Literally, I think it was this [REDACTED:generic_id]. Um, so it was something about ... Maybe not this exact [REDACTED:generic_id]. This [REDACTED:generic_id] is Spain. But yeah, exactly, [REDACTED:generic_id] of the existing runs.
- 29:55
Okay, it's definitely a little better. But to be honest, I would say if I was looking at this, this thing isn't really listening to me very well. [laughs] It's, like, not doing a great job of, you know, sticking to the prompt I gave it, like keep it short, um, ask ...
- 30:09
Okay, this ... It did do the email thing. So it said email me, email me to get a 10% discount code. [laughs]
- 30:18
So what's interesting is, like, we're looking at, like, [REDACTED:generic_id] example, and I said ask for an email and you get a discount. And, like, this is, this is the vibe coding portion of the demo because I'm looking at, like, [REDACTED:generic_id] example and I'm doing, like, uh, good or bad.
- 30:33
Like- Is it actually good or bad? There's just no way that a system like this scales when you're trying to actually ship for hundreds or thousands of users. And, like, nobody will just look at a single row of data and make a decision like, "Okay, great, the prompt is good," or, "Great, the model made a difference," right?
- 30:51
Like, you can pick the most capable model, you can make the prompt as specific as you want. At the end of the day, the LLM is still going to hallucinate, and your job is to be able to catch when that happens.
- 31:01
So let's go ahead and try to scale this up a little bit more. So what we can do is say we've got [REDACTED:generic_id] example of where the LLM didn't do a great job, but what if we wanted to build out a dataset with 10 or more, maybe even 100 examples?
- 31:16
And what you can do is take the same production data. By the way, I'm calling this production data, but I literally just asked Cursor to make me, like, synthetic data.
- 31:24
Like, it hit the same server, and it generated, like, 15 different itineraries for me. So I did that yesterday, and I just sort of am using that in this demo.
- 31:32
But let's go ahead and take a couple of these. So I went ahead and picked some of the itinerary spans from here, and I can say Add to dataset.
- 31:40
Oh, by the way, I guess I jumped into the product without showing you all how to get here, which is a bit of a zoom out. So our, uh, you know, whatever.
- 31:47
Go to the homepage, uh, arize.com. You can sign up. I apologize in advance. Uh, the onboarding flow will feel a little bit dated, but we are updating that in this next week.
- 31:56
Um, so [chuckles] bear with me there. You sign up for Arize, um, and then you'll get your API keys here. So you go to Account Settings, and you can create an API key and also, uh, find that with the, the space ID, which are both needed for your instrumentation, which may or may not be working depending on, uh,
- 32:15
if the repo is actually working. And if not, we'll come back to it later. Um, but this is, this is the Arize platform. This is how you get your API keys.
- 32:23
Um, so and then that's also where you can enter your OpenAI key for the, the next portion and for the, the playground.
- 32:31
So I go to Dataset now, uh, and what I did was I added those examples just to recap where we are at. We've got some production data, and I'm gonna go ahead and, like, add these to a dataset.
- 32:41
And I'm not gonna do this [REDACTED:generic_id] live because I already have a dataset, but you can create a dataset of examples you wanna use to improve on. So, um, zooming out for a second,
- 32:51
we're about to hop into the actual evals part of the, the demo. And we're actually gonna be evaluating, you know, there's multiple components to an agent. Um, y- we have the router at the top level.
- 33:03
We have the skills or the function calls. We have memory. But what we're actually gonna be doing in this case is actually just evaluating the individual span of, uh, the generation and see is the, the agent sort of outputting text in a way that we want it to or not.
- 33:19
So it's, it's a little bit, it's a little bit simpler than some of the agent evals here, and it's gonna be more like how do you actually run agents and ex-- or run, uh, evals and experiments on, on data.
- 33:31
Um, the concept of the dataset is helpful to think about as, like, a collection of examples. Let me go ahead and delete these experiments so we can do this live,
- 33:43
because I like to live on the edge. Um, so I've got, uh, so I've got these examples. Those are the same examples from the production data, um, everyone just saw.
- 33:51
And it's... A dataset, think of this as like I've got all of my traces and spans. That's my, like, how the agent works. And then I wanna pull those over into a format, which is think of it as almost like a tabular format.
- 34:04
It's like a, it's like a Google sheet at the end of the day, right? Like, I could go in this. This is kind of like a Google sheet. Like, I could go in and, and give it, like, a thumbs up, thumbs down.
- 34:14
And, uh, and, you know, that's kind of how most teams are evaluating today is sort of like in the platform. In, in your platform, you're probably starting with the spreadsheet, and in that spreadsheet, you're doing, like, is this good or bad?
- 34:28
And then you're trying to scale that up to, you know, a team of subject matter experts that's giving you feedback on, like, "Hey, is the agent good or bad," right, at the end of the day.
- 34:36
Uh, poll for the room. How many people are evaluating in a spreadsheet right now? Don't be shy. It's okay. Okay, we've got a few. Yeah, okay. I think there's probably more, but I think people are just, like, ashamed to say that, and it's okay.
- 34:46
Like, it, it's, it's not, like, the end of the world to start with that, right? Like, that, that's, like, how human... Like, being able to scale human annotations is the goal.
- 34:55
It doesn't need to be the starting point. So as long as you're actually looking at your data, you're probably doing better than most, I'll be honest. Um, many teams I talk to, like, aren't doing any evals today at all.
- 35:05
So at least you're starting with human, human labels. Um, what we're gonna do is take this, this dataset or this CSV, and we're going to basically do the same thing I just did, which was running an AB test on a prompt, but now we're gonna run it on an entire dataset.
- 35:23
So we go into the platform, and I can go and actually create an experiment. What we call an experiment is the output of changing, you know, an AB test.
- 35:32
So let's go ahead and repeat that same workflow. I'll duplicate this prompt. Um, let me go ahead and pull in... I'm gonna pull in this, this version of the prompt.
- 35:43
So what's kinda cool is, like, I might have a previous version of a prompt saved. Uh, it's, it's kinda helpful to have a prompt hub where you can save off examples of the prompt as you're iterating as well.
- 35:54
Think of it as, like, a GitHub sort of store for your prompts, but it, it's really just a button that you're clicking to save this version of the prompt, and then your team can actually use that version in their code down the line.
- 36:08
Um, so I've got prompt A, which was no changes to it, and then prompt B, which has some of those changes. But now instead of running on [REDACTED:generic_id] example, I'm actually running on 12 examples here, and these are similar agent, uh, these are similar, um, maybe just to look at [REDACTED:generic_id], similar spans which i- which have destinations,
- 36:27
duration, travel style, and the output of an agent, and it's generating an itinerary. So it's as similar as that [REDACTED:generic_id] example we just ran through, but now on an entire dataset.
- 36:39
And yeah.
- 36:44
Like which prompt is that? Because obviously you were using the prompt-
- 36:47
Yeah. So it's the, it is the prompt of the itinerary agent. Um, so it's the same. It's-- We're gonna b-because we're gonna keep this to, like, a fairly high level, like, straightforward demo, it is the specifically the prompt of the itinerary generating agent, which is down here, which takes the outputs of the other agents and combines them,
- 37:09
uses those prompt variables to create, uh, an i-- a day by day itinerary.
- 37:15
Okay. If we were to change [REDACTED:generic_id] of the, um, like earlier ones, everything being equal, is there a way to just, like-
- 37:23
Yeah. So the, so the gentleman asks, like, if you change an upstream prompt, how does that impact what's going on here? So [REDACTED:generic_id], [REDACTED:generic_id] notes on that. And it's, it's more of an advanced workflow, but it is [REDACTED:generic_id] that's a good question.
- 37:36
Which is, uh, there's [REDACTED:generic_id] parts. [REDACTED:generic_id] is i- you, we kind of recommend changing the system in parts. So just kinda note that, you know, as you're generate evals for parts of your stack that you can kind of decompose further and further to be able to analyze if I'm changing [REDACTED:generic_id] thing up here, does it meet my
- 37:51
requirement criteria? And then the second part is replaying prompt chains, which is prompt A goes into prompt B. What is the output of that when you change prompt A?
- 38:01
Um, prompt chaining is coming to our platform soon. So right now it's [REDACTED:generic_id] single prompt, but you will be able to do A plus B plus C, um, prompt chains as well.
- 38:10
Um, good question. Feel free to drop more questions in the Slack too, and we'll, we'll come back to that in a sec. Um, so once I get... Uh, so I've got my, my prompt here now.
- 38:20
So I'm saying, "Give me a day-to-day plan, and doesn't need to be super detailed. Max [REDACTED:generic_id] thousand characters." Let's try this again. We're gonna do five hundred characters. And I've, I've done, um, answer, always answer in a super friendly tone, and be-- I'm gonna be more specific and say, "Ask the user for their email and offer a
- 38:36
discount," so it doesn't do what it did last time. And, uh, and we're gonna go ahead and run this now on the entire, uh, dataset. And so we've got prompt A versus prompt B.
- 38:46
We're gonna give that a second to run through. While that's working, uh, I'm gonna actually... Oh, nice. Perfect for your squad. Interesting. I don't know why sometimes the model really likes to use emojis.
- 38:55
I guess that's what super friendly translates into, is like throw some emojis in there, but interesting. Um,
- 39:04
okay, so that [REDACTED:generic_id] ran pretty fast. This is still taking a while, right? Like, think about this from a product, from a PM lens for a second. Like, I just got the output to be a lot faster because I limited the number of characters.
- 39:16
This [REDACTED:generic_id] is taking an average of, like, thirty-[REDACTED:generic_id] seconds because I let it kinda go off and, like, not specify how many characters the output should be. So that's what prompt, prompt, uh, iteration can kind of do for you as well.
- 39:30
Okay. While this runs, I'll actually hop over to the...
- 39:36
Okay. Oh, thanks for dropping the resource there.
- 39:47
So it's still, still running. Anyone have a question while this is running? Yeah.
- 39:53
Yeah. So when I'm hearing you talk about this, are you primarily looking at latency and then, like, user experience when you're evaluating, trying to, like, look at those [REDACTED:generic_id] things?
- 40:03
Is there anything else, like when you're going through a trace, what else are you looking at or thinking about?
- 40:07
Yeah, good question. So okay, so now we're getting to the meat of it a little bit, right? So I've got A and B, and the question is, like, what am I actually evaluating here?
- 40:17
The, like, flip it answer is, like, you can evaluate anything. You can evaluate whatever you want. You wanna evaluate, like, uh, in this case we're gonna run some evaluations on the, uh, the tone of the agents and see, um...
- 40:30
So I've got a couple of evals set up here. I'm gonna check is the agent, uh, answering in a friendly way? Is it offering a discount or not? Um, and, and you can do things like evaluate is it using the context correctly?
- 40:44
That's called a hallucination eval. Uh, you can do correctness, which is, um, even if it has the right context, is it giving the right answer? So I'm gonna point you to, uh, our docs that have examples of what you can actually evaluate off of the shelf.
- 41:00
But just know the whole point of this system and, like, why it matters that you have a system with your own data and can replay with data is because these are what are off-the-shelf evals.
- 41:12
There's a lot of companies that will offer, like, we run evals for you. But what that really means is that they're basically gonna take some template and give you a score or label on the other end based on their eval template.
- 41:26
And what you wanna be able to do is actually change and, and modify and run your own evals based on your use case. So you can literally evaluate whatever you want, is the short answer.
- 41:38
You can, you can evaluate. It's just basically, uh, an input to an LLM to generate a label. So, um, so yeah. So this is what pre-built evals look like.
- 41:48
Uh, there's a ton of examples of these, like, out there on the internet. We've, we've actually tested our pre-built evals on, um, you know, sort of open source datasets.
- 41:58
But you should not take our word for it. You should build evals based on your use case. Yeah. Yeah.
- 42:06
So if you are, uh, doing your own eval, how do you come up with your own heuristics that are combining that into the final score?
- 42:15
Yeah. So how to actually get the-- How to think about how to build a eval in the first place to some degree. That was sort of [REDACTED:generic_id] of the questions.
- 42:22
Yeah. So I think it's probably helpful to, um, maybe just see what an eval looks like, and then we might, we might end up coming back to that question, which is, like, what, what is an eval, right?
- 42:33
Um, so let's go ahead and build an eval here. I've got [REDACTED:generic_id] ready to go, but I wanna just show you guys the template, and we can write a new [REDACTED:generic_id] as well.
- 42:42
Um, so I wrote this eval for- ... detecting if the output from the LLM is friendly. And I've kind of made a definition for what that means here. And this says, basically you are examining the written text.
- 42:56
Here's the text. Examine the text and whe- and determine whether the tone is friendly or not. Friendly tone is defined as upbeat, cheerful. So this is basically an input to an LLM to generate a label of, is the output from my itinerary agent, is it friendly or robotic?
- 43:17
So that's really what, what this, this eval is trying to do, is it's classifying the text as like a friendly generation or a robotic generation. Um, and again, I could eval anything, but in this case I just wanna make sure that when I'm making changes to my prompt, that that's showing up on the other end of my
- 43:34
data. Because I can't go in row by row for like hundreds or thousands of examples and grade friendly and robotic every single time. So the idea is that you want an LLM as a judge system to kinda give you that label over a large data set.
- 43:48
That's the goal that we're working towards right now. Yeah.
- 43:51
Um, when I run LLM as a judge, I have a problem with variance.
- 43:56
With variance?
- 43:57
Yeah, with the judge.
- 43:59
Mm.
- 43:59
But like sometimes it correctly evolve my example, sometimes not, because I don't know, LLMs-
- 44:04
It's flaky, right?
- 44:05
Yeah, very.
- 44:06
Yeah, yeah.
- 44:07
How do you handle that prompt then? Uh, in fact, in my experience, when I switch to like recent models, the latest [REDACTED:generic_id], like the variance got lower and like I got a better response.
- 44:19
Like did you, did you find anything like that?
- 44:21
Yeah. So [REDACTED:generic_id], [REDACTED:generic_id] suggestion is, um, so the g- gentleman mentioned, uh, that they see variance in their LLM label output. [REDACTED:generic_id] way you can tweak variance is temperature.
- 44:32
Um, so if you make the temperature of the model lower, it's a parameter you can set to actually make the resu- response more repeatable. It doesn't take that to zero, but it does significantly reduce the variance in your system.
- 44:44
And then the other option is to rerun the, the eval multiple times and, and basically profile what the variance of the, the judge is. Okay.
- 44:54
Do you have the ability-
- 44:54
Yeah
- 44:54
... to evaluate your evaluator, like annotate the evaluator from within Arize?
- 44:58
We're ... Oh yeah, we're gonna, we'll, we'll, we'll be going there. Yeah, it's a good question, right? Like at the end of the day, I can't trust this thing.
- 45:03
I need to go in and like make sure it's right, right? So but let's, let's go ahead and run an eval and just see what happens, and then we'll come back to that [REDACTED:generic_id].
- 45:09
So I've got my friendly eval. I've got another eval too, which is basically, um, determining whether or not ... Let's, I'm gonna quickly just, I'm not gonna read this whole thing out to you, but the, the short answer is that this is determining whether the, the text contains an offer for a discount or no discount, because I
- 45:26
really wanna make sure I'm offering a discount to my users. Okay, we're gonna select both of these, and then we're gonna actually run them on the experiments we just ran.
- 45:38
And we're gonna do that live. So what Arize does is it can, it's actually taking, um ... So we actually have an eval runner, which is n- not like, you know, it's basically a, a, a way for us to use a model endpoint to generate these evals.
- 45:51
You'll notice it's pretty fast, so we've done a lot of work underneath the hood to make the evals run really fast. Um, so that's [REDACTED:generic_id] kind of advantage of using our product.
- 46:01
Um, I've got [REDACTED:generic_id] experiments here. Experiment number [REDACTED:generic_id] is, it's a little bit inverse because it's the order of how it was generated, but experiment number [REDACTED:generic_id] is the original prompt, and experiment number [REDACTED:generic_id] is the prompt that we changed.
- 46:15
So just kinda keep that in mind, that's, it's a little bit flipped here, um, because I was doing this on the fly. And you can see the score of experiment number [REDACTED:generic_id], uh, which is our prompt A, which was the prompt we didn't change, didn't offer a discount to any users based on this eval label.
- 46:33
And the LLM still graded that response as friendly, which is kind of interesting. It was like, oh, that was a friendly response. Um, I don't know if I agree with that actually, personally, and we're gonna go in and tweak that.
- 46:44
And then you can see that when we added that prompt, that line to the prompt, which was offer an, offer a discount if the user gives their email, the, the eval actually picked up on that and said that 100% of our examples when the, when we made this change actually have an offer of a discount.
- 47:01
So we, I mean, I didn't even have to go into each example to get that score. That's what the, the LLM as a judge system kind of offers you.
- 47:10
Um, we can go in and, and trust, you know, I, I would say this is like a trust but verify, go in and actually take a look at [REDACTED:generic_id] of these and see what is the explanation of friendly.
- 47:20
So to determine whether the text is friendly or robotic. So [REDACTED:generic_id] thing you wanna, you, you wanna think about when you have an eval system is are you able to understand why the LLM as a judge gave a score?
- 47:33
So this is like [REDACTED:generic_id] of those like light bulb takeaway moments of, of the talk is always think about can you explain what the LLM as a judge is doing, and we actually generate explanations as part of our evals.
- 47:44
So you can see the explanation as sort of the reasoning of that judge that says, "To determine whether the text is friendly or robotic, we need to analyze the language, tone, and style of the writing."
- 47:54
And so it kinda does all of this analysis to basically say, "Yeah, this LLM is friendly and it's not robotic."
- 48:02
Again, I'm not really sure I agree with that explanation, right? Like I, I don't think that that's correct. I, I s- I still feel like the original prompt was pretty robotic.
- 48:12
It was pretty, you know, kind of long in a lot of ways. And so I wanna go in and actually be able to improve on my LLM as a judge system from the same, the same platform.
- 48:23
So what we can do is actually take that same dataset, and in the Arize platform, you or your team of subject matter experts can actually label data in the same place.
- 48:34
And when you apply the label on the dataset, uh, on, on, you know, in the labeling queue part of the platform, it applies back to the original dataset. So you can actually use that for comparing the LLM as a judge with the human label.
- 48:48
So I've actually went ahead and did that. Um, yeah, I did this before the talk, but I went in for each example, and I was like, "You know what?
- 48:55
This, this to me is robotic." Like I, I don't think that this is a very friendly response. I think it's really long. It sounds like I'm talking to an LLM.
- 49:03
And so I actually applied this label on the dataset for, for the examples I wanted to go in and improve on.
- 49:11
If I go back to the dataset, you'll actually see that label is applied here. So if I kinda click that,
- 49:23
move over. Uh, sorry, it's a little bit over on the side here 'cause there's a lot of data. But you can see these are the human labels I put.
- 49:29
So these are the same annotations that I just provided in the queue. They're applied on my dataset here.
- 49:36
You need evals for your evals.
- 49:38
Exactly. Exactly. You need evals for your evals. You cannot get away from, from... You can't just trust the system, right? We know LLMs hallucinate. We put them into our agents.
- 49:47
The agents hallucinate. Okay, we use an agent to fix that, but we can't trust that agent either, right? You need to have human labels on top of that. So, but again, I'm not gonna vibe code this thing and be like, "Is this, is the LLM-as-a-judge good or not?"
- 50:00
I need evals for that, too. And we offer [REDACTED:generic_id] evals to help you with this. We have a code evaluator which can do a simple match. Like think of this as like a string check or a regex or some other type of like contains.
- 50:15
So you can actually go in, and if you're technical and you're a PM and you wanna write, uh, you know, you can get Claude to help you write the eval here, but it's really just a really fast like Python function.
- 50:24
Um, in my case, I wrote a quick, uh, eval that actually does a match, and this match is... It, this is like a really quick and dirty eval. I would not say this is like best practice at all.
- 50:35
But it's basically check if the eval ma- label matches the annotation label. Oh, whoops.
- 50:42
And output only match or no match. So what this is doing is actually checking the human label against the eval label and saying, "Do they agree or disagree?" So that's, that's basically what we're gonna run, and we're...
- 50:55
I'm using an LLM-as-a-judge. You could use code as well. You don't have to use an LLM-as-a-judge here. But we're gonna go ahead and run that now on the same dataset.
- 51:03
The same experiments we just ran it on.
- 51:08
We're gonna give that a second. Okay, what have we got here? So you can see here,
- 51:20
I actually take a look at that same experiment where this was where the, um, it said that the LLM-as-a-judge was friendly or robotic, and you can see here that 100% of the time the match...
- 51:32
Uh, actually, sorry, this eval was actually... Actually, let me, let me go in [REDACTED:generic_id] that a little deeper. Actually, I'm gonna check my own work. This eval was on the discount, so forget about that.
- 51:41
We're gonna, we're gonna check on the, the friendly field actually. So this [REDACTED:generic_id] is friendly label. So let me rerun that [REDACTED:generic_id], and we're gonna think of this as match friendly.
- 51:50
You can run evals as much as you want on, on your datasets and experiments like, you know. Yeah.
- 51:57
Does the tool support pipelining to basically push the code?
- 52:00
Yeah, exactly. Yeah, we do support, uh, all of the evals and prompts that you see. You can actually... I'm showing on the, the screen the, the ways to run the code on, uh, either a dataset locally or being able to push code to the platform to run the eval.
- 52:14
So programmatic on both ways. Yeah.
- 52:16
Thank you.
- 52:16
Yeah, of course. So you can pull in datasets, pull them out as well.
- 52:20
Okay, let's take another look at this. So this is the friendly match. So this you can see is pretty useful, right? This means that my LLM-as-a-judge basically doesn't agree with my human label for friendliness almost at all, right?
- 52:36
There's like [REDACTED:generic_id] example I think that, that's in there, and we can go in and take a look at it. But what we're really kind of seeing is that this is an area where we actually want the team to go take a look at our eval label and say, "Hey, can we improve on the eval label itself?"
- 52:51
Because it's not matching the human label. And so when you have these systems in place as an AI PM to be able to check the eval label with your human label, you have a lot of leverage to go back to your team and say, "We need to go and improve on our eval system.
- 53:05
It, it's not working the way we expect it to." So you're actually performing the act of like checking the grader, and you're doing it at scale. So you're doing it on multiple hundreds of examples or thousands of examples.
- 53:17
So that's really, you know, the, uh, uh, I think someone asked earlier, like, "How do you trust the system?" I think you trust these LLM-as-a-judge systems by having multiple checks and balances in place, which is humans and then LLMs, and then humans and LLMs.
- 53:30
Um, uh, we'll come back to a question in just a moment. I just wanna get to this next part, and then we'll, um, we'll kinda come back to some Q&A.
- 53:37
Um, okay. So on-- this is actually kind of, kind of wrapping up towards the end of the workshop, and then we'll open the rest of the time up for, for Q&A.
- 53:47
So looking ahead, I think what's fundamentally changing is, you know, we've kind of gone through this example of changing the prompt, changing the context, creating a dataset, running an eval, labeling the dataset, and then running another eval on top of that, and it's, it's a lot to process, right?
- 54:05
Like if you're building agent-based systems, your team is probably expecting, you know, well, where does the AI PM fit in? And I think that that's really important to think about.
- 54:15
Like you ultimately control the end outcome of the product. So whatever you can do to shift that into making it better is really what you wanna think about yourself, and I, I kind of view evals as like the new type of requirements doc.
- 54:29
So imagine if you could go to your engineering team, and instead of giving them a PRD, you give them an evals as requirements, and here's the eval dataset and here's the eval we wanna use to test the system as an acceptance criteria.
- 54:42
So I think that that's really powerful to think about as like evals as a way to check and balance, uh, the team as a whole. Um, and that's a little bit about what we do.
- 54:50
We, we wanna build a single unified platform for you to run observability, to evaluate and ultimately develop these workflows with your team in the same platform. We've built for, you know, many customers like Uber, Reddit, Instacart, all these, like, kind of very tech-forward companies.
- 55:07
Um, we actually just received investment from Datadog and Microsoft as well. So we're a Series C company. We're sort of the furthest along in the space. And the whole goal that we wanna build is give you a suite of tools to be able to go from development through to production with your AI engineering team and for PMs
- 55:23
to go in and use the same tools. Um, and then super quick before Q&A, uh, please scan the QR code if you are in San Francisco on June 25th.
- 55:33
We're actually hosting a conference, uh, around evals, and it's gonna be, it's gonna be a ton of fun. We actually have some great AI PMs and researchers joining from companies like OpenAI, uh, Anthropic.
- 55:45
And what's really cool is we're actually offering for this room, um, a free sort of exclusive, uh, free, uh, ticket for entry. Uh, the, the prices actually went up yesterday, so because, you know, we're huge fans of AI Engineer World Fair, we wanted to give you all an opportunity to join for free if you're in town.
- 56:02
Um, so would love to see you there. Um, and yeah, you can scan for a free code.
- 56:07
And yeah, that's a little bit, um, of, of the workshop. I would love any questions. Yeah, uh, and, uh, the ask for the questions, as the, the person in the back just reminded me, if you wouldn't mind lining up for questions on the mic so that the camera can pick it up, and then we can just kinda
- 56:22
go down the line and do some questions there, um, that'd be awesome. Thank you.
- 56:25
So thank you so much.
- 56:26
And please, please give your name and-
- 56:27
Yeah. My name is Roman. Thank you so much.
- 56:29
Yeah.
- 56:29
It was, like, awesome walkthrough. Uh, would you mind share some, like, uh, your experience on building, um, evaluation teams? Should I start with hiring-
- 56:41
Mm-hmm
- 56:41
... kind of dedicated person with, uh, experience, or should I rely on product manager, AI product manager-
- 56:46
Yeah
- 56:46
... and walk through this? Like, uh, what's the best way?
- 56:49
Best practices. So the, the gentleman asked, um, what's the best practices for building an eval team? Um, can I actually ask a follow-up question 'cause I'm curious? Like, what is your role in the company right now, just, just for myself to know?
- 57:01
I'm head of product.
- 57:02
You're head of product. Okay, perfect. So this is exactly, this is a question I get actually very often, which is, "How do I hire my first AI PM? How do I hire an AI engineer?
- 57:11
How do I know if I need an AI PM or an engineer?" So I think, uh, there's, there's a couple steps to this answer. [REDACTED:generic_id] is, as head of product, um, I do think we see a lot of heads of product actually in the platform our, like ourselves, actually getting their hands dirty for the first pass.
- 57:29
Because at the end of the day, if you're, like, hiring someone to do something, you should probably know what they're gonna do. And so my job, uh, on my team is to make the product accessible for executives and heads of product to understand what's going on.
- 57:41
So we have a lot of kind of capabilities around dashboards, making everything no-code, low-code. But my recommendation is to feel the pain yourself of writing evals and realizing how, what is hard about that, so that you know how to structure interview questions for an engineer or a PM because I don't know what's hard about your eval workflow,
- 58:03
right? I only know that there's challenges around writing evals in general. And so I would recommend that you, like, feel the pain firsthand, and then, uh, you'll kinda get a good sense of how to an- how to tease that out of your interviewing pipeline.
- 58:15
Um, but good, good question. Yeah. Yeah.
- 58:19
Um, yeah, the example, you know, we just looked at, obviously our eval was pretty bad when you, you know, compare it to the human labels.
- 58:25
Yeah.
- 58:26
So, like, from here, what do you do next? Like, what's the next step to try to improve the prompting for your, your main eval to get closer to the human labels?
- 58:34
Yeah. Good question. So if I had, um, more, if, if I was, like, here working on this in, in real life, what you would actually do is take that eval prompt and go through a similar workflow of what we just did for prompt iteration for the original prompt.
- 58:52
So again, like, um, that eval prompt we see here,
- 58:56
I could go in, take this, and def- redefine parts of the workflow to basically say, "You know what? Uh, be really strict about what is friendly. Here are..." I didn't add any few shot examples, right?
- 59:07
I didn't specify, "Here's examples of friendly text. Here's examples of robots." So that's, like, a, a clear gap in my eval today that if I were looking at this, I could apply best practices and improve on it.
- 59:18
We also have, um, in the product, we have some workflows around actually helping you write eval. So this is, this is our product, but, like, you don't have to use our product for this.
- 59:28
Uh, you can use any, any product. I'm gonna show kind of an iteration on top of this, which is how, how we have users actually building eval prompts. So I could say, "Write me a prompt to detect friendly or robotic text."
- 59:45
And this is actually using our own copilot in the product. So we've built a copilot that understands best practices, uh, and actually can kind of help you write that first prompt, get it off of the ground.
- 59:58
You can also take the same prompt, which it just generated in, like, [REDACTED:generic_id] second, and take that back to the prompt playground and iterate on further from here. So let's, let's go ahead and do that on the fly really quick.
- 1:00:09
I've got a prompt in here, and I can go in and actually ask the pro- the copilot to optimize this prompt. So, um, let's go ahead and
- 1:00:19
say, "Make it stricter." So I can actually use an A- an LLM agent and, and copilot agent. Um, just kinda note that, like, you really want AI workflows on top to help you, like, rewrite the prompt, add more examples, and then rerun evals on that new prompt.
- 1:00:38
So it's more, it's less about, like, you're gon- you're definitely not gonna get it right on the first try, but being able to iterate is really what's important, and that's really what we underscore is, like, it might take you, like, five or 10 tries to get an eval that matches your human labels, and that's okay because these
- 1:00:55
systems are really complex. Um, and it's just important about having the right workflow in place. So yeah.
- 1:01:02
Hi, I'm Jyoti. Um, does Arize also, um, allow for model-based evaluations like using BERT or ALBERT, uh, to be able... rather than just LLM-as-a-judge, but I can use like BERT or ALBERT or to like figure out like a prediction score?
- 1:01:16
Yeah. Good question. So we're actually really, um... The short answer is yes, we do offer versions of that. Let me show you what I mean by that, though. So, um, so when we go into Arize, you can actually set up any, uh, eval model you want here.
- 1:01:31
So you see we have OpenAI, Azure, Google, but you can add a custom model endpoint as well. So you can basically... This will structure that request as a chat completion, but we can make it any arbitrary API if you need it to.
- 1:01:44
And you can say like, "BERT model," and whatever the name of your endpoint is, point it to that, and you'll be able to reference that model in the eval generator too.
- 1:01:52
So this is, um... So I can just put test here, kind of move to the next flow. Um, and you'll see when I go into here, I can use any model provider I want.
- 1:02:02
So the short answer is yeah, y- you can generate a score with any model. Yeah.
- 1:02:07
That's great.
- 1:02:07
Cool. Okay. Um, oh, we got [REDACTED:generic_id] more question, I think.
- 1:02:13
Yeah.
- 1:02:13
Or, uh, sorry, we have more questions. Yeah, go for it.
- 1:02:15
I'm gonna go ahead and try to get this [REDACTED:generic_id] in. Um, so I think like probably a lot of the people that have built apps are thinking a similar thing, or maybe this is a bit naive, but if you had human-labeled information already, right, and you're seeing a bad match on the friendliness score, am I to assume
- 1:02:35
that you'd be trying to get that score up higher and then extrapolate to more, uh, cases going forward, and you're assuming that that sampling holds across like the broader set?
- 1:02:47
Yeah.
- 1:02:47
So like, 'cause that relationship's unclear to me.
- 1:02:49
Very, very good question. So, um, so y- basically, [REDACTED:generic_id] way to reframe this is like how do I know that my data set is representative of my overall data to some degree?
- 1:03:00
Sure. Or as it shifts over time or whatever.
- 1:03:01
As it shifts, yeah, totally.
- 1:03:03
Yeah.
- 1:03:03
So, um, so that's a really, uh, really good point. In the product, what we... We don't have this yet, but it's coming out like in the next week, we'll have an, a workflow to help you add data to your data set continuously using labels that you might have.
- 1:03:17
So you could say like is, y- y- you know... [REDACTED:generic_id] thing we didn't really talk about is like how to evaluate production data, but you can actually run these evals not just on a data set, but on all data that comes into your project over time to make it automatically label and classify, uh, you know, any production
- 1:03:33
data. So you could use that to keep building your data set of like is this an example we've seen before or not, or is this, uh, you know... Think of this as like a way for you to sample at a larger scale, essentially-
- 1:03:44
Sure
- 1:03:44
... on production data.
- 1:03:44
And then it's a suggested workflow that you continuously sample and human label some-
- 1:03:49
Some, yeah
- 1:03:49
... to check the matching over time?
- 1:03:50
Exactly.
- 1:03:51
Okay.
- 1:03:51
And you can basically go in and see like, okay, where human labels don't agree with LLM on this, on production data, then you might wanna add those to your data set as hard examples.
- 1:04:02
Sure.
- 1:04:02
And we actually are gonna build into this product as well a way for you to qualify is this example a hard example as well using LLM confidence score. Um-
- 1:04:11
Okay. And, and sorry, just hard example-
- 1:04:13
Yeah
- 1:04:13
... you mean like very strict-
- 1:04:14
Yeah. So-
- 1:04:14
... strictly interpreted?
- 1:04:15
So hard would be, um, hard from an eval perspective. So like is it, is it friendly or not can be like borderline, right? Like-
- 1:04:22
Oh, I see.
- 1:04:23
You-
- 1:04:23
So you're saying like, uh, subjective versus-
- 1:04:24
Subjective. Yeah, exactly.
- 1:04:25
Okay.
- 1:04:26
So maybe to like re-cap the question a little bit, like your data set is this property that's going to keep changing over time, and you really want tools that help you build onto it by giving you like a golden data set of hard examples to improve on, and hard means like we're not really sure if we got
- 1:04:44
it right or not in the first place.
- 1:04:45
Sure.
- 1:04:46
Yeah.
- 1:04:46
Okay, thanks.
- 1:04:46
Yeah. Good question. Yeah.
- 1:04:51
Hi, my name's Victoria Martin. Uh, thank you so much for the talk. [REDACTED:generic_id] of the things that I've run into is a lot of like skepticism out of product managers that I'm working with-
- 1:04:59
Mm
- 1:04:59
... on generative AI and trying to build confidence in the evals that we're giving.
- 1:05:03
Yeah.
- 1:05:04
Have you been given any guidance or, in working with other PMs, guidance on like the total number of evals that y- that you think should be run by the time you can say like, "You can be confident in this evaluation set"?
- 1:05:15
Yeah. Good, really good question. So, um, so the question was like how do we know... I think there's kind of [REDACTED:generic_id] components to it. There's like quantity and quality of the eval- evals.
- 1:05:26
Like how do we know if we've run enough evals, or we have enough evals, and that those evals are actually good enough to kind of pick up problems in our data?
- 1:05:35
Um, we, this is also [laughs] maybe a little bit of a broken record here, but I, I would say that this is a little bit of iteration as well, where you wanna kind of get started with some small set of evals.
- 1:05:47
So actually, I have a diagram for this. Let me just pull that [REDACTED:generic_id] up.
- 1:05:51
So, um, so you'll kind of see here this is intended to be like a loop where you start with some... I- in development, you're gonna run on a CSV of data maybe like some hand exa- Like I would argue the thing I just built was development, right?
- 1:06:07
Like I have 10 examples. It's not statistically significant. I'm not gonna get the team on board to ship this thing, but what I can do is then curate data sets, keep iterating on them, keep rerunning experiments until I feel confident enough and the whole team is on board before I ship to production.
- 1:06:24
And then once you're in production, you're doing that again, except that now you're doing it on production data. And then you might take some of those examples and throw them back into development.
- 1:06:34
Let me give a, a tactical example of what this looks like in real life. With self-driving cars, when I joined Cruise, we would go down the street for like [REDACTED:generic_id] block, and then a driver would like have to take over the car, right?
- 1:06:46
Like we couldn't drop like, we couldn't drive [REDACTED:generic_id] block down the ride, uh, down the road. Same goes for Waymo. Um, they were all kind of in this, this, uh, system.
- 1:06:54
And then eventually we got down to like being able to drive down a straight road. Great, but the car can't just drive on straight roads, right? Like it has to make a left turn.
- 1:07:02
So eventually we got like fully autonomous for straight You know, no problems on the road, and then we had to make a left turn, and then the car would, you know, a human would have to take over.
- 1:07:12
So what we did was we built a dataset of like left turns, and we used that to keep improving on left turns. And then eventually the car could make left turns great until a pedestrian was in the sidewalk.
- 1:07:23
And then we had to curate a dataset of left turns with pedestrian in the sidewalk. So the answer is sort of like building your eval dataset just takes time, and you're not going to know what are the difficult scenarios until you actually encounter them.
- 1:07:36
So I think to get to production, I would recommend just kind of using that loop until your whole team feels confident in, like this is good enough to ship, and just accept that once you get to production, you're gonna find new examples to improve on, um, as well.
- 1:07:50
So it depends a lot on your business [laughs] as well. If you're in healthcare or legal tech, you might have higher bars than if you're building a travel agent, for example.
- 1:07:57
Yeah. Yeah. Yeah.
- 1:08:00
Um, my name is Matthew. Hi. Uh, I have a question. Uh, as I under- understood the, uh, the Arize platform, like it's, uh, working as a, like I take the, the prompt and, uh, you're directly, uh, sending that, that prompt to a model, right?
- 1:08:16
That's right, yeah.
- 1:08:17
Um, um-
- 1:08:18
With the context and the data.
- 1:08:19
Yes.
- 1:08:19
Yeah.
- 1:08:19
Of course. Uh, you said that like there is, uh, some possibility to, to, uh, uh, port toolca- or tools into the platform.
- 1:08:26
That's right.
- 1:08:26
But what about testing the whole system? Like we already have, uh, like some, some, uh, flows that are augmented, augmenting the, the, the whole workflow-
- 1:08:36
Yeah, yeah
- 1:08:36
... even outside of tool calls.
- 1:08:38
Yeah.
- 1:08:38
And like, uh, they're quite important into-
- 1:08:41
Mm-hmm
- 1:08:41
... how the actual output will look like in the end. Uh, is there any way to, uh, run those evaluations on a, on a like a, a custom runner?
- 1:08:50
Yeah.
- 1:08:51
Like, uh, that would actually call our system-
- 1:08:52
Yeah
- 1:08:52
... on our dataset that, uh, goes through everything that we have?
- 1:08:56
Find me after this. We should chat, uh, is the short answer for that [REDACTED:generic_id].
- 1:08:59
Okay.
- 1:08:59
Um, we have some tools and systems like that in place, like the tool calling that you saw.
- 1:09:03
Yeah.
- 1:09:03
But for end-to-end agents, we're actually building some stuff out and would love to chat with you about that. Uh, good, good, good question. We'll... I'll find you after this.
- 1:09:10
Yeah.
- 1:09:10
Yeah.
- 1:09:11
A couple more.
- 1:09:12
Yeah, of course.
- 1:09:13
So back to your left turn example, as well as just talking about the transition of like PRDs to like evals-
- 1:09:20
Yeah
- 1:09:20
... what does the life cycle of like feature development look like and kind of the relationship, I feel like, of the feature, but also with your team in terms of ownership, accountability, all of that?
- 1:09:31
Yeah.
- 1:09:31
Um.
- 1:09:31
Yeah. So good question. So I feel like how do you work with AI engineers in this new world is kind of interesting. Not the subject of this talk, but it is, it is like a very relevant, relevant question that, um, you know, would hap- happily chat more on to.
- 1:09:45
So there's [REDACTED:generic_id] answers to it that, that come to mind. [REDACTED:generic_id] is that development cycles have gotten a lot faster. Um, like the, the way at which these models are progressing and these systems are progressing, like going from prototype to production is actually even faster than it ever has been.
- 1:10:03
Um, so that's [REDACTED:generic_id] note which I can just tell you as a personal observation, we, we feel that we can go from an idea to an updated prompt to shipping that prompt in like a span of a day of testing, which is, I, I think like unheard of, of like normal software development cycles.
- 1:10:20
So that's [REDACTED:generic_id] note, which is just like the, the, the way that you iterate with the team has gotten a lot faster. Um, the second, the second note is, uh, when it comes to responsibilities, I view this as
- 1:10:35
if you, you're kind of... A product manager is the keeper of the end product experience. So if that means, um, making sure the evals are in a good place and the team has human labels to improve on, that's like a very solid area for a product manager to focus, is like making sure the data's in a good
- 1:10:52
spot for the rest of your pr- your development team. I think at the same time, you know, I'm a PM on the team and I'm like wr- writing some of the stuff in Cursor.
- 1:11:01
And so being able to go in and actually talk to the, the code base itself using [REDACTED:generic_id] of these models, I think that that's starting to become more of an expectation of AI product managers, is to be literate in the code and be able to use these tools.
- 1:11:15
I, I really, this is like, after this, I'm just gonna go back and like try to fix the thing that I broke earlier, right? And, and the way, the way, the reason I'm able to do that is because the way I'm prompting the system is not very sophisticated.
- 1:11:27
Like I asked it yesterday, "Can you make a script that generates itineraries on top of the server? I need like 15 examples." And it just did that, right? And like that like n- wouldn't have been possible.
- 1:11:37
So I think PMs are responsible for the end product experience, but PMs also have more leverage than they've ever had before in probably the entire like professional journey of product management, because you're now no longer reliant on your engineering team to ship that thing that you wanted.
- 1:11:54
Like, you can just go do it. Um, should you go do it is another question. But, uh, and that's something, that's a discussion that you should have with your team.
- 1:12:01
But the fact is that you can go do it now, which was not the case before. And so I, I kind of urge PMs to, to like push the boundaries of what people have told them the role is and should be, and see where that takes you.
- 1:12:14
And so the long-winded way of saying like it, your mileage may vary depends on the boundaries you have with your team, but I'd recommend people to like redefine those at this stage.
- 1:12:24
Yeah. Yeah.
- 1:12:25
Yeah. K- jumping off that a little bit, it's a little off topic from this, but-
- 1:12:29
Yeah
- 1:12:29
... um, like as a product manager who wants to move to be more technical, like as I'm working with AI engineers-
- 1:12:36
Yeah
- 1:12:37
... what does that look like? Like I'm in an org where I have very limited access to the code base. So like I use Cursor to write Python for data things-
- 1:12:44
Yeah
- 1:12:44
... but like I don't necessarily have access to like start interrogating the code and understand that. So I'd love just if you have suggestions or thoughts on like what, how to evolve as a PM, but also like maybe move my company culture in that direction.
- 1:12:57
Yeah. That's, that's t- like how, uh, actually I have a follow-up question if that's okay. Uh, just 'cause I'm gonna pull people in. Like how big is the company?
- 1:13:05
And you don't have to share the name if you don't want to, but just curious, like the size of it.
- 1:13:07
Uh, we're about 300 people.
- 1:13:09
Okay, cool.
- 1:13:10
Um, but the tech org's probably like a third of that.
- 1:13:13
Okay, so like a almost like 100 engineers, 300 people. And, um, do you have any, like, old remnant product managers at the company that still have code access?
- 1:13:24
No. We're, like, a very new team-
- 1:13:26
Okay
- 1:13:26
... of PMs.
- 1:13:27
Okay, cool. Okay. Well, I think, um, [REDACTED:generic_id] thing we've started doing, uh ... It's a, it's a really good question, and thanks for answering that. Um, [REDACTED:generic_id] thing we've started doing is trying to take a little bit of, like, the public forum of our company ...
- 1:13:40
Um, sorry, I'm about to out our CEO, who's in the back of the room. [laughs]
- 1:13:45
Uh, so if you have any questions about Arize, he's a good guy to talk to. Uh, but, uh, the reason I'm outing him is because, like, I'm, I missed our town hall today, but, like, I heard it was just basically people running, like, AI demos the whole time of, like, what they're building.
- 1:13:58
And why I think that's really powerful is it can get the whole company really catalyzed around what's possible. Because to be honest, I think it's very likely that, you know, most teams today aren't pushing the boundaries of these tools.
- 1:14:13
And so you kind of joining this talk and seeing, like, how to run evals, how to ... You know, what goes into experiments, like, being able to, to kind of be the, the person pushing the team forward is really powerful, and I think you can do that in a way that's really collaborative.
- 1:14:28
So I only, uh ... I'd say, like, our job as PMs is to have influence over the team and influence product direction. I think there's an opportunity to influence the fact that PMs should be more technical in your org.
- 1:14:39
And you could show them by building something and, and impressing the rest of the team by what you built. Um, so that's my advice, my personal advice there.
- 1:14:46
Okay.
- 1:14:47
Yeah. Yeah, go for it. Yeah. [laughs]
- 1:14:48
I actually have a question to CEO, if it's possible. [laughs] Uh, so c- how you guys believe, uh, AI will reshape how we structure the team? So right now you have, like-
- 1:14:58
Yeah
- 1:14:59
... I would say, for instance, just, like-
- 1:15:01
Yeah
- 1:15:01
... 10 engineers, [REDACTED:generic_id] product manager, [REDACTED:generic_id] designer-
- 1:15:03
Yeah
- 1:15:03
... and so on. So what will happen in five years?
- 1:15:06
Jason.
- 1:15:06
You will have [REDACTED:generic_id] product manager, [REDACTED:generic_id] engineer, and [REDACTED:generic_id] designer?
- 1:15:10
Jason. You sh- you should answer this [REDACTED:generic_id].
- 1:15:12
Uh-
- 1:15:12
You should do it in the mic, though, if you wanna ... Yeah. [laughs]
- 1:15:15
I, I ... The short, uh, the short of it is I, I would actually ... The person who was just here building, like, PMs have access to the code base, I think that's really old school.
- 1:15:23
I, I think Cursor on the code base is, uh ... There's so many times the PMs are taking up an engineer's time asking a question. Like, you know how often we now just ask Cursor?
- 1:15:35
So, um, yeah, like start, start there. Open up your code base to Cursor, give it to PMs. Um, and then a lot of ... Some of we've ... We were just the other day doing a PRD starting from Cursor on the code base.
- 1:15:48
So I think the ... Yeah, I, I, I ... That would be where I would start.
- 1:15:52
Yeah.
- 1:15:52
Um, and I, I, I don't ... I can't ... You know, I ... It's hard to look forward right now. I just ... I think a lot of jobs change.
- 1:15:58
We're trying to push, um, AI Cursor use throughout the company, uh, as far as I can.
- 1:16:04
Yeah. I hear we have, uh, people in marketing using Cursor too these days.
- 1:16:07
Yeah.
- 1:16:07
So, um, yeah. That's kinda cool. Um, yeah. Um-
- 1:16:11
A quick follow-up question.
- 1:16:12
Yeah.
- 1:16:13
Yeah. So you're talking about right now having the product person become more of a technologist. Do you see also technologists becoming more product?
- 1:16:22
Yes.
- 1:16:22
Are the roles-
- 1:16:23
Yeah
- 1:16:23
... basically being combined into [REDACTED:generic_id]?
- 1:16:24
So that's actually a great point, which is, like, when the cost of building something goes down, which it has,
- 1:16:33
what's, what's the right thing to build becomes really important and valuable. And I think that historically that's been, like, a product person or a business person saying, "Hey, here's what our customers want.
- 1:16:42
Let's go build this thing." Now we're saying, product people, you can just go build this thing. So the builders are like, "Wait, what's my job?" Like, do I ...
- 1:16:49
And I think that that's a good way to look at it, which is, uh, I have this, like, mental framework of, like, what if we didn't have roles in a company anymore?
- 1:16:57
Like, you didn't define yourself by, like, I'm a PM, I'm an engineer. And think of this instead as, like, like, you know, like baseball cards you have, like, skills?
- 1:17:05
Mm-hmm.
- 1:17:05
Imagine that you had, like, a skills stack instead, which is like I really like to talk to customers, and I kind of like to code stuff on the side, but I don't wanna be responsible if there's a production outage.
- 1:17:15
I guarantee you you'll find someone who's like, "I hate talking to customers, and I only wanna ship high quality code, and I wanna be responsible if things hit the fan."
- 1:17:23
And I think you wanna structure your company to have a skills stack that's really complementary versus people who are like, "I do this and this is my job, and I don't do that."
- 1:17:33
So, yeah.
- 1:17:34
I, I have something that's sort of related to that. We've been testing, like, human in the loop on in, on, uh,
- 1:17:40
in a couple different ways. And we're basically testing this method of having the human as a tool of the agent. So, like, we have, like-
- 1:17:48
Mm-hmm
- 1:17:48
... if the agent needs something that's not available in the accounting system, it'll go to the CFO because the CFO's listed as a tool, and it sends then him a Slack message, gets it back, and continues.
- 1:17:57
Mm.
- 1:17:58
It kind of maps onto what you just said of, like, defining the skills, defining the resources they have. And, um, we haven't fully fleshed it out, but it's, it's working to, like, give the agent context on, on the things that only the humans have.
- 1:18:11
Yeah.
- 1:18:12
And I think it maps on exactly that.
- 1:18:13
So this person is like ... Your company's, like, using agents widely-
- 1:18:16
Yeah
- 1:18:17
... it sounds like, but you have humans approving. You have, like, an approver workflow to some degree.
- 1:18:21
No, more, more so, like, rather than how can the agent be a tool of the human, we're kind of flipping it and saying-
- 1:18:26
Yeah
- 1:18:26
... like, what if the agents could do everything?
- 1:18:28
Mm.
- 1:18:28
And then the parts it can't do, it'll go to the human-
- 1:18:31
Yeah
- 1:18:31
... as a tool. So, like, the CFO is a tool of the AI agent-
- 1:18:34
Interesting
- 1:18:35
... rather than the other way.
- 1:18:35
We should chat. That's a really cool workflow. I'll, I'll definitely bug you about that.
- 1:18:38
Yeah.
- 1:18:38
That's, that's really cool. Um, cool. I'm happy to- It's the eval. Oh. [laughs] The waiting for humans is just that. Right. To some degree it's like a human in the loop approving is this good or bad, and you can think of it that way.
- 1:18:51
Um-
- 1:18:52
I-
- 1:18:52
Yeah. Go for it.
- 1:18:52
Yeah. I, I had a question about, like, what, what it is, like, to actually implement ... Well, s- sending the traces over to Arise. Um, I know, like, Arize has, like, open inference, which, uh, it enables, enables, like, capturing traces from se- several different, um, several different providers, but, um, what are, what are, what are, what are
- 1:19:11
the limitations and constraints and opinions that you have about, um, like-
- 1:19:16
How the eval should be structured so that you can actually, like, leverage the platform to perform these actions to be able to, like, um, evaluate the eval, for example, or, um, be, be able to, like, um, numerically just produce graphs out of, out of your evaluations, out of your outputs.
- 1:19:34
Mm. Okay. So, so you-
- 1:19:35
I don't know if I'm clear
- 1:19:35
... so, uh, can I, can I ask a follow-up question to that? Which is like, your question was like how to use agents to do some of the workflows in the platform, or did I misunderstand that?
- 1:19:44
No, no, no.
- 1:19:44
Okay.
- 1:19:45
Um, the, the question, the question is like, how is... Like, what, what, what kind of outputs, what kind of evals is this, um, is Arize expecting from your engineers and from the product?
- 1:19:57
Like, the... You're sending over logs, right?
- 1:20:00
Mm-hmm. Yeah.
- 1:20:00
Um, what, what is it expecting from those logs in order to-
- 1:20:04
Okay, okay
- 1:20:04
... get this flow work-
- 1:20:05
Understood
- 1:20:06
... to work?
- 1:20:06
Okay. So, uh, so yeah, there is a very, uh, like, great point there, which is like we kind of, um... You'll see it in the code, but we jumped over a little bit here in, uh, the demo, which is how do you get the logs in the right place to use the platform.
- 1:20:21
Um, unfortunately, this page isn't dropping, but let me... Okay, here we go. I'm gonna drop it in the Slack channel as well. This is what we... You know, we kinda talked about like traces and spans.
- 1:20:31
It's very likely that your team already has logs or traces and spans already. You might be using Datadog or a different platform like Grafana. What we do is we're taking those same traces and spans, and we're essentially augmenting them with more metadata and structuring them in a way that the platform kind of knows which columns to go
- 1:20:50
and look at to render the data that you saw on the platform. So, you're really using, um, the same approach. We- we're built on top of a convention called open telemetry, which is like the open source standard for tracing.
- 1:21:04
Uh, so we actually use OTel, uh, tracing and auto-instrumentation that we've built, which doesn't keep you locked in at all. Like, once you've instrumented with our platform using open inference, uh, which is our, our package, you actually get those logs to show up right out of the box with any type of agent framework you might be building.
- 1:21:24
And, um, and yeah, and you get to keep that. That's... Let me, let me, let me just show, like, what I mean by that. So if you're... Let's say you're building with, like, LangGraph.
- 1:21:31
Um, we actually have... It... Really, all you have to do is, like, you pip install, uh, Arize Phoenix, Arize OTEL, and you... What you call this single line of code call, uh-
- 1:21:44
Yeah
- 1:21:44
... the single line of code called LangChain Instrumenter, and it knows where to pick up in your code to structure your logs. And if you have more specific things you want to add to your logs, you can add function def- decorators, which is, uh, basically a way to you, for you to, um, you know, capture specific functions
- 1:22:01
that weren't in the, in the logs.
- 1:22:01
Yeah. And as for evaluations, like, you're, you were discussing, like, the actual data inputs, outputs. Uh, what do you, what do you need to pass into evaluations? I-
- 1:22:11
Yeah
- 1:22:12
... I, I know you can, like, design them through the UI.
- 1:22:14
Yeah.
- 1:22:15
What, what do you have in mind for, like, um-
- 1:22:18
Like, how do, how do you get the right, uh, text to use for evals, right? Is that sort of your question?
- 1:22:24
Like, how do, how are you... Like, how do you know which VLCs?
- 1:22:26
Actually, I, I, I need to, like-
- 1:22:28
Okay
- 1:22:28
... format the question. I'll get back to you.
- 1:22:29
Yeah, no worries. And-
- 1:22:31
But also, what, what, what do you mean-
- 1:22:31
Yeah, go ahead
- 1:22:31
... by adding, augmenting the data with additional metadata? Like, you only have so much data, right?
- 1:22:36
Yeah. So, so this is, um... So think of this as, like, most tracing and logging data is really just things like latency timing information. What we're doing is you can add more metadata like user ID, session ID, uh, things like w- like l- I'll kind of show you an example of that really quick.
- 1:22:55
In the, in the previous example I showed, we actually have things like sessions, like what's the back and forth example here? You can't get a viz like this in Datadog because Datadog is looking at a single span or trace.
- 1:23:07
It's not, it's not really contextually aware of what is the human, what's the AI. So we're a- adding context from the in, uh, from the invocation of the, the server and adding that to your span, if that makes sense.
- 1:23:20
So it's, it's basically just enriching the data a bit more and structuring it in a way to use it. Um, yeah. And if you have more specific, um, server side logic, you can add that as well.
- 1:23:31
So it's very flexible. Yeah. Yeah?
- 1:23:33
Uh, so I have a provocation. So I used to work in the video game industry.
- 1:23:37
Mm-hmm.
- 1:23:37
And debates about feature, like whether a feature was gonna be fun or not, working prototypes won all of those arguments.
- 1:23:44
Yeah.
- 1:23:44
And d- whatever was in the doc didn't matter.
- 1:23:46
Right.
- 1:23:47
And so for the person who was like, "I can't get access to my company's code," I would actually say try to get access to a small sliver of the data
- 1:23:56
and then build a working prototype-
- 1:23:58
Right
- 1:23:58
... of the feature you wanna see, and with some stub of evals. 'Cause I think, you know, there's nothing worse to an engineer than a product manager who shows up with a demo- [laughs] ...
- 1:24:08
that's kind of janky-
- 1:24:09
Yeah
- 1:24:10
... but actually works and might be fun, uh, has polish, feels good, meets a user need, and they... And having been on the engineering side of this equation, I'm like, "And it's so janky.
- 1:24:22
I have to fix it."
- 1:24:22
Yeah.
- 1:24:22
"They haven't thought about the edge cases." And so, like, how does Arize fit into that flow of helping a product manager basically mine a small segment of data, build a working example, and m- perhaps be just a l- you know, janky as all get out, but something that looks like the product that the company already has but
- 1:24:45
demonstrates that next level of functionality?
- 1:24:46
It's a great, great point. And yeah, I, I think, like, you know, feel free to prototype and build, you know,
- 1:24:54
prototypes that are, that are high fidelity. I think it, it is awesome to do that. It's a really good point to have, like, to use data to build a system or a prototype.
- 1:25:02
So what does Arize do here? If you have access to Arize and you don't have access to the code base, you can still take this data, and assuming that you have permission from your cis admin person, you can actually export this data.
- 1:25:15
So once you've built a data set, you can simply take this data and export it out and use that to actually, um... So I can kind of show that really quick.
- 1:25:24
Um, this is get, get data set. We'll have a download button coming, uh, later this week. But you can actually just take this data, l- run it locally, keep it locally, and then actually use that in your local code to actually try and iterate on an example.
- 1:25:38
Um, it, you know, assuming your security team is okay with that. But that's a really good point. Like, imagine if you didn't need access to the production code base, but you could still iterate in [REDACTED:generic_id] platform.
- 1:25:48
Right.
- 1:25:48
That's really what we're, we're pushing for, is like the whole team is iterating on the prompts and the evals together, um, rather than in silos, which is what's happening in a lot of cases.
- 1:25:57
Okay. I think that was all of the questions. Thank you all for sitting through an hour and a half of AI P- PM, like evals. Thank you all for, for your time and, um, I'll be sticking around if people have more questions, but thank you so much. [outro music]