AI Engineer Europe 2026
Evals Are Broken, Use Them Anyway
About this talk
Cline's Ara Khan argues that coding-agent evaluations are indispensable but misleading when benchmark scores are treated as objective truth or replaced entirely by subjective impressions. He examines SWE-bench Verified's limitations, describes constructing evaluations that better reflect real programming workflows, and uses Terminal-Bench to show how standardized Linux environments, CPU and memory configuration, and model-specific agent harnesses affect results.
Chapters
- 0:00Why evaluations are broken but still necessary
- 2:54Subjective preferences, changing models, and SWE-bench limitations
- 6:10Cline's approach to realistic coding-agent evaluations
- 10:43Terminal-Bench infrastructure and standardized task environments
- 14:09Agent harnesses, model-specific behavior, and closing remarks
Talk transcript
- 0:00
[upbeat music] All right.
- 0:15
All right. Um, first of all, thank you so much for, for coming. I'm actually rather surprised. Um, a lot of times, like, you're, you're working on this stuff and you're, like, cooped up in a room, and you're like, "No one cares."
- 0:27
And then it's like so many people showed up, so I suppose someone cares. Um, so anyway, so the, the title of my talk today is Evals are Broken and You Should Use Them Anyway.
- 0:37
And a lot of this talk is just, like, a straight-up critique of, like, the way we do evals these days. And I kinda wanna help you out. I kinda wanna give you a way out of this.
- 0:46
It's like you have this, like, interesting technology and you can use it, but there's, like, so many ways to, like, mess it up. So I'm gonna help you out.
- 0:53
So my first claim is that people are wrong about evals. Actually, let me, uh, let me correct myself. Most people are wrong about evals. And I want you to be right about evals.
- 1:06
I want you to, I want you to use them. I want you to, like, I want you to be able to build with them, interpret them, use evals in your own agentic flows.
- 1:14
Um, leverage them in any way sense that you can. Um, so that's, that's, that's basically the point of the conversation. So to be right, to be right about something that has, like, a lot of nuances that can go in, like, many different directions, the fundamental question is, like, how are people wrong about that thing?
- 1:32
Right? So there's basically two camps of people who are wrong about things, right? So there's two camps of wrong on eval. The first camp is a camp of objective metrics.
- 1:42
So the objective metrics camp is this: like, there are people who would, like, look at a list, this dashboard, right? And they will interpret this as something akin to, like, like it means something, as in, like, GPT-5.4 is effectively the same as, like, Gemini 3.1 Pro Preview.
- 2:01
Believe me, they're not the same. Uh, there's a lot of these models which will, like, show up, like, s- with similar numbers and these numbers at a certain point it's like, this is...
- 2:09
the whole thing is a hoax. Like, you, you won't believe it at all. So there was this tweet that came out, like, just this morning, and it was a critique of Meta, where Meta came out and it was just, like, classic benchmark maxing.
- 2:20
Just like, "We're doing best on the benchmark. Everything's great." And I, I assure you, if you try a lot of these models, like, it just, just wo-won't hold the, hold the test of, like, actual real world evidence.
- 2:32
The other one, the other camp, the other way where people are wrong is that they go too far the other way. So this is basically, like, the test camp.
- 2:39
Um, so people in the test camp are kind of like this. They just, like, they're, they're an archetype. They're... They just... They think that it's like, it's... everything's about, like, vibes and everything's about, like, you know.
- 2:54
Like, it's like if you ask them, like, "Why do you like Opus?" And they'll say things like, "I like talking to her." Like, they'll, they'll, like, anthropomorphize it. And, um, that's also not right either, right?
- 3:03
So I think the truth is somewhere in the middle, that the evals are not the end all be all. They're also not completely useless. There are right ways to use them.
- 3:09
There are wrong ways to use them. Um, so in order to do that, I'll give you, like, three stages that will help you, like, uh, use them really well.
- 3:16
So the first stage is, like, you can, like, leverage evals from other people. The second stage is to, like, use evals to improve your own agents. And then the third stage is, like, build, actually build your own eval for specific use cases.
- 3:28
In the interest of time, I could personally talk about evals for hours, but, like, I can only talk about level one and two, and I think those would be most helpful for most people in the audience.
- 3:37
Um, so I'm gonna give you, like, a few heuristics, uh, to interpret eval. So the first heuristic is that whenever a model apps comes out with, like, a number, just don't believe them.
- 3:46
Just don't. Just, like, these are approximations. Coming back to the tweet, like, just, like, just don't believe the model app eval numbers. Like, they're somewhat of an approximation. Sometimes they're good, sometimes they're not.
- 3:56
Um, there was this tweet that's, like, a pretty cool one where, um, they couldn't say it, that, like, a lot of, like, AI researchers and engineers routinely dismiss evals.
- 4:06
They don't really, um, they don't really think of it as, like, something that's, like, the numbers to be taken that seriously. And it, uh, some- to some extent it's a matter of, like, actual trying and preferences, and I think, I think that is, like, somewhat more accurate.
- 4:19
Um, wait a second. Yeah. So the second heuristic is that, like, um, you wanna stay current, but you don't wanna be the earliest adopter. And why am I saying this?
- 4:29
So this is, like, a epoch index, um, is basically like the aggregate, uh, score of, like, different models on, like, evals. And if you notice, like, in the last, like, two years, every single couple months, like, the frontier lab, the frontier model is changing, and it's changing so fast.
- 4:48
Like, it's just, like, it's so hard to keep up with this stuff. And I've worked for... I, I've worked on this. I've been doing this for a living for years at this point, and I, I...
- 4:55
even, even me, I have, like, preferences changing so fast. So, like, I think that, like, when, when you're, when you're working through these things, like, the way I would recommend is that, like, let the thing come out first.
- 5:05
Let things set, set on fire for, like, a couple weeks and then if the, if the thing still stands the test of time, I think at that point you should, like, do your model switch and, like, uh, try something rather than, like, always trying to be on the cutting edge.
- 5:16
Like, the, the people who have to always try the new model and try the new things, like, don't be me, but, like, I do this for a living and y-you don't have to.
- 5:24
Um, the third, the third heuristic for evals is that, like, you need to look for very, very new and very precise evals. And the reason this is necessary is that, like, a lot of evals are, like, at this point that have become, like, standardized, they're actually kind of old.
- 5:38
Like, they're, they're not useful for you. So, like, this is a, this is a blog post from OpenAI where they straight up said SWE-bench Verified no longer measures frontier coding capabilities.
- 5:48
Um, and I think, like, to a lot of people in the AI research community that was, like, very obvious. It was very obvious that SWE-bench doesn't measure frontier co- coding capabilities because it had, like...
- 6:00
it would have tr- problems like solve the Fibonacci sequence. It would have problems like, you know, do a matrix multiplication or something, and it's just, like, it doesn't apply to, like, real world software engineering.
- 6:10
So you wanna have something that's like very new, but also like actually legit. Uh, and it takes some discernment to figure that out. So that's like the first part, but the second part is like, okay, now that we know that like, okay, that we have a few heuristics of like how to like use evals, like how do
- 6:24
we use evals to improve your agent, uh, upon them? And I think this is the part where I kind of lean into like the, the core philosophy of this conversation where like you wanna think of evals as like an engineering problem, but also as a philosophy problem, right?
- 6:38
So the engineering problem is obviously hard, but the philosophy problem is also very hard. Uh, the philosophy problem is that you, you wanna, you wanna ... You have a problem, and you can't exactly approximate the, the search space of like where, where the problem could go, where the problems could fail.
- 6:56
Um, it's, it's sort of somewhat easier-ish to do it for coding problems, but even then, coding problems have like an infinite search space. They can go in any direction.
- 7:05
So you wanna, you wanna build evals that like are somewhat more approximate representation of the actual thing that you're dealing with. And, uh, for us, like to give some context in Cline's journey.
- 7:16
So I work at Cline. Cline is an open source coding, uh, agent company. We have a, we, we have a very interesting, uh, product. I encourage you to try it out.
- 7:23
So in Cline's journey, one of the things that we dealt with is like in the last year, o- one of the things we found is that like there were like a few evals available.
- 7:30
At the time, it was like we were very rudimentary. Every, every other company was very rudimentary as well. And our thinking was like, okay, like if, if there's like, if there's like so few, uh, standardized evals available, and also they're not a factor, like they really are not measuring what it is that you're trying to do in
- 7:49
your day-to-day programming job, like what do you do? So our stance was, and this was the stance of the Codex team and a lot of other teams that we've talked to, uh, vibes.
- 7:58
Vibes was our stand, just like this, this eval, just completely ignore them. They're completely unnecessary. You're probably, probably wasting your time, and it's just like I don't know who, who will be appeased by them.
- 8:08
And then last year, we, we, um, we came to this idea that like, okay, like listen, um, I think, I think, I think we, we gotta up the ante, and we gotta like, we gotta have some measure.
- 8:18
We gotta try evals, and if no one else is doing it, we'll do it ourselves. We'll build actual evals from scratch, um, that would like actually test real-world programming problems, uh, of users.
- 8:29
So we got like, uh, we went through a lot of like our massive datasets of like people who had opted in to share their coding usage of Cline with us, and we offered them money, and we, we got a lot of this dataset of like, okay, this is what the problems that people are actually doing.
- 8:43
Then spent a lot of time parsing through that, figured out like an actual data stuff, like these are the problems that people are solving, and then just like completely cleaning it all up, like doing a lot of like really hard manual labor.
- 8:55
We're trying to make like very decent problems that can be solved with, say, uh, Cline or any other coding agent. Um, the hardest part for us when we were building, uh, evals is that like if you're building eval for anything that's like rudimentary, like if you're, if you're building eval for, say, an LLM model, you have a
- 9:13
very simple like one-shot use case of like how many toes does a cat have? And then the LLM can just be like, "I don't know, 11 or whatever." I don't, I don't know how many toes a cat has.
- 9:23
But, um, a single-turn eval is very easy to do because it has like a binary answer, and it has just like a very limited search page of what the answer could be.
- 9:32
But when you're working with an agent, that can't be the case. When you're working with an agent, you can give an agent a problem like, "Hey, uh, I have this new MCP server.
- 9:40
It's probably not working. Like, how do you, how do you, like, make it work for me?" And that's usually how a lot of you guys talk to, uh, Cloud Code or whatever agent you're using, and myself as well.
- 9:50
So in this, like it's very hard to gauge because like the agent like reads through files, searches through docs, uh, installs environment, sets things up, runs Python scripts, does all of that, and then in the end, like runs some tests and then maybe the whole thing works.
- 10:03
Like so we're, we're trying to grade the second thing. We're trying to ge- grade like all these things that will take a lot of time, and then figure out like, oh, did it actually work or did it not?
- 10:11
Did it work but like broke other things? Like you... So that's why it was like harder. So in the same time, uh, some very, very awesome, smart, bright people from Stanford University came up with Terminal Bench, which does the same thing, where they came up with like 89 coding problems, which are, which are just like very approximate,
- 10:30
decent representation of like real-world programming problems. So these could be things like, you know, um, race conditions, uh, database issues, um, like, um, uh, other stuff like it's like figure out this infra issue.
- 10:43
And I think that that was built, the Terminal Bench was built to like use with any coding agent CLI, so you can like, you can actually test and run things, um, like, uh, really fast with like, uh, CLI, and then it will take like a couple minutes to run.
- 10:57
So some of these tasks would take up to like 30 to 40 minutes. Uh, and that's how you know they're legit because like you, the agent does a lot of things and just like runs in circles and sometimes just goes crazy and, and yeah.
- 11:07
Um, so, uh, so we started using that. And the way to use that, like the way to use Terminal Bench is that like you think of an eval problem as like, uh, an evaluation suit which has a set of problems, and this one has 89 tasks.
- 11:20
Some others would have more. And what you wanna do is like you wanna give it an environment. You wanna give it an isolated environment where you just like you...
- 11:29
Let's say you have a task like, "Hey, figure out this race condition for me in this repo." And the race condition is that like this thing is not working.
- 11:37
So you wanna be able to give the, the eval run an isolated environment like a virtual machine. In that virtual machine, it has the whole setup. It has the repo.
- 11:46
It has everything. And then you install whatever agent you have, in our case, Cline, Cloud Code, Codex, whatever you wanna use, you can do that. Um, to do that, it's, it's not that it's like hard, but it is also not trivial.
- 11:59
And Harbor is another software that came from Law Institute, um, where they made this thing where, let's say we have 89 tasks. One way to do evals is that learn like each of these tasks in sequence, and then like do all the setup.
- 12:10
Another way to do is like have a very standardized configuration defined in infrastructure where each of these eighty-nine tasks have like the proper Linux machine, uh, the proper RAM/CPU usage.
- 12:22
Um, and then like being able to like isolate those environments and then run those eighty-nine tasks in parallel on infrastructure. So you could use a couple different things for the infrastructure here.
- 12:31
Uh, you could use Daytona, you could run it on your Docker Machine if you have like very powerful machines. I'm sure if you can handle those-- that much compute, sure, but I wouldn't.
- 12:40
Uh, we use Model. Um, Model, we're very thankful to Model. They've helped us a lot, um, so shout out to them. And uh, yeah, so in this case like Harbor basically lets you split up like the eighty-nine tasks, and then they all run in parallel.
- 12:52
So that way your limit, the limiting factor is basically the slowest task.
- 12:58
Um, yeah. So, um, yeah, so the slowest task is the limiting factor. So the process is this: you, you, you get a score, you first do a run, uh, on the eighty-nine task.
- 13:09
You get a score. You evaluate all the failures. So let's say you get like, say fifty failures, right, out of the eighty-nine tasks. What you wanna be able to do is you wanna portfolio allocate those failures.
- 13:21
So you wanna say, uh, you wanna run like a- another agent which goes through the traces of all the failures. So the trace would be like this massive file which has like every single LM call that the, that the e- that the agent did, and then be like, "Okay, this one, this specific problem failed because it didn't
- 13:37
run tests. This failed because the read file tool was broken." And once you portfolio allocate those failures, you figure out, "Okay, these are the small levers that I can pull.
- 13:46
If I pull those le- levers, like I can make like massive improvements to my AI agent." So what you're testing is like you're basically testing like, um, three things.
- 13:56
You are testing the model itself. Um, like if you have a very decent model, like somehow like you could have a horrible harness, you could have a horrible agent, but like the model just like overshoots so hard that just like you, you, you know, you, you get a great score.
- 14:09
You're testing the harness, you're testing your coding harness, so like you, you're testing, say, if you're using Claude Code, Codex. So sometimes you'll find ... I'm sure, I guarantee you, some of you have noticed that like, let's say Anthropic's models could potentially work with the cursor, could potentially work with Droid, could work with other coding agents, but
- 14:26
for some reason it just seems to work so much better- Yeah ... with Claude Code, right? And I think that, that that is like the testing the harness, that like is the harness actually really leveraging the best, uh, best of the model.
- 14:38
And the third problem, whether a problem is sane. If you're solving stupid problems, it doesn't matter if you score 100% all the time. So you really gotta make sure that like, um, um, the, the problems are sane, which, um, the LlaMA Institute has done a pretty great job of.
- 14:51
So for us, it was a case like this. Like, um, we, we basically like, um, had like this original score which was like much lower, like 43%. Uh, we made changes to like CPU, we made changes to memory in front of the containers.
- 15:04
Uh, we raised timeouts. We improved the thinking behavior. Sometimes we would ask the model to think more. Sometimes th- asking model to think more actually interferes with the quality of the response because it goes in like, it gets like a stroke, and it just like goes in like circles. [laughs]
- 15:16
And it's like it'll just like, "I am a model. I am a model." It would just like keep doing it for like, uh, like, uh, 2,000 tokens. So yeah, so like you, you gotta think through all of that.
- 15:26
And uh, yeah, so for us, like we have like a huge, um, like internal benchmark for all kinds of models, open source models. So like we just like, uh, keep like a list of like trying different versions and stuff.
- 15:38
Um, we encourage other people to try that as well if you, um... Pretty helpful. So whenever, whenever you get like zones improvements, you get like basically three zones of improvements when you get an original score.
- 15:48
The first one is the obvious flaws. Like sometimes your harness really has like very obvious flaws of like there's this bug that ca- straight up crashes the harness. Um, and those obvious bugs you gotta fix, right?
- 15:59
Sometimes you are not, like you're getting rate limited or whatever. Fix those, that's fine. I think the zone two is the most critical one where you actually do nuanced improvements.
- 16:07
And these nuanced improvements are things like there are certain prompt engineering techniques that apply to Anthropic model families that just straight up would not apply to Codex model family, that would be very different from Gemini model family.
- 16:20
And those are the nuance of like, "Why is it that this is a model that's so good, that so many people are saying it's so good, but for some reason it just isn't working for me?"
- 16:29
I think those, that's, that i- is the essence of like working with agents and hill climbing. That like you figure out those nuanced improvements of like tweaking your prompt, making it larger, making it smaller.
- 16:39
Uh, and then zone three is the danger zone where it's like you're straight up overfitting. So you're overfitting in the sense that like you're just straight up cheating to get the highest score, and then you can like make a tweet about it.
- 16:47
Don't, don't ... like a lot of people have done it. Don't do it. Like I wouldn't do it. I mean, never mind. Anyway, so yeah. So anyway, so this was like, this was like basically the rough outline.
- 16:58
So the final, uh, the final wording for me would be like basically i- regardless of the kind of problem that you have, I want you to like find a benchmark and like, like build the eval and just like hill climb on it.
- 17:09
So hill climbing means that like you get a score, and then you improve the score of your harness on the eval. And you have to do both. Like you can't just like have like a good number and be happy with it.
- 17:19
Like you, you, you gotta both have the vibe check of like does it actually feel good to use this product and this model, and at the same time you also have like a very, very, very decent, uh, score hopefully.
- 17:30
Um, if a new thing comes out you, you, you do your absolute best to like give it the right judgment. Uh, for us, like one of the things that we learned was that like we were very decent on Anthropic model families, not so much on, say, Gemini model family, not so much on, say, uh, Q-mini model family,
- 17:45
which again are very decent models. Um, so when we, when we started hill climbing, we learned that like, oh, if we support these models, we have this like entire swaths of people who love these models, and they can start using us.
- 17:57
And I think that in some reflection of that would also apply with you. Um, so my final, my final note to you guys is that like, you know, I've, I've done some hot takes or whatever, and if you work for, for some of the companies that I've said not so nice things about, I still love you and
- 18:10
everything. And uh, it was, it was, um ... I work at Cline, so if you find these problems fascinating, if you wanna learn more about these, like, uh, this my Twitter.
- 18:18
So like you can feel free to reach out to me, DM me about like if you wanna work on problems like these, like by all means, like I can put a word for you.
- 18:25
If you wanna learn more about evals, if you wanna learn like how do like ... I have a problem that's like completely orthogonal to everything you're defining for coding agents, like how do we work on that?
- 18:36
Uh, so feel free to reach out to me and I'll, I'll respond to you. And, um, once again, it's very, very kind of you to give me your time.
- 18:42
Thank you so much. [audience applauding] [upbeat music]