AI Engineer Code 2025
Why Agent Hype can fall short of reality – Joel Becker, METR
About this talk
METR researcher Joel Becker contrasts rising AI benchmark and human-calibrated task-horizon results with a randomized field study of 16 experienced open-source developers. He explains why benchmark scores and autonomous task completion do not necessarily predict productivity on messy, context-dependent software work, while emphasizing uncertainty and the limitations of small studies.
Chapters
- 0:00Joel Becker introduces METR and the benchmark-versus-field-evidence gap
- 1:53Why SWE-bench and GPQA scores are difficult to interpret
- 3:29Human-calibrated task suites, Claude 3.7 Sonnet, and time horizons
- 10:27Real-world deployment study with experienced open-source developers
- 17:29Potential productivity explanations and small-study uncertainty
- 20:58METR hiring and closing remarks
Talk transcript
- 0:00
[upbeat electronic music] Hey, guys.
- 0:21
Thank you so much for having me. My name is Joel Becker. I work as a researcher or member of technical staff at METR, which stands for Model Evaluation and Threat Research, as we'll see in a second.
- 0:31
I'm going to be talking about AI capabilities. How do we know how performant AIs are today? How, how performant they might be in the near future from these two different sources of evidence that seem to give somewhat conflicting answers?
- 0:43
You know, I, I could have done this whole talk without reference to METR papers in particular, but we'll look at two papers I've been, um, involved with as, as examples of benchmark-style evidence and then more economic-style evidence.
- 0:54
On the benchmark side, measuring AI ability to complete long tasks, this is the paper, um, that comes with the, the chart that many of you would have seen on, on Twitter and so on, that METR is well known for.
- 1:05
And then the second, this, um, RCT measuring how allowing AI affects developer productivity. And then we'll be talking about how to reconcile, uh, the, the gap that's implied between these two different kinds of measurements.
- 1:19
As I mentioned, METR stands for Model Evaluation and Threat Research. We are a independent research nonprofit that seeks to inform the, the public, policymakers, labs about the degree to which AIs might pose catastrophic risks to society.
- 1:33
The model evaluation part, uh, means that we seek to understand AI capabilities and propensities, and the threat research part means we try to connect those capabilities and propensities to potential catastrophic risks.
- 1:47
Okay. The first paper we're going to talk about, associated with this chart that, that many of you, I think, might have seen.
- 1:53
Um, take- taking a step back first before we dive into the paper. You know, how, how usually do we think about measuring AI capabilities using benchmarks on a SWE-bench or a GPQA, so on and so forth?
- 2:04
There's some notion of 0% performance, um, or, or random performance. So for GPQA, that's, that's twenty-five percent, which corresponds to this floor, the, the worst you can possibly do.
- 2:14
Perhaps there's a, um, human baseline that's below a hundred percent. For GPQA, I think this is something like seventy-five percent that represents maybe expert human performance. And then of course, you can go all the way up to a hundred percent potentially on, on these kinds of benchmarks.
- 2:28
But, but what does it mean? You know, if I'm getting fifty percent on GPQA, if I'm like half the way from the, um, from the floor to the, to the expert baseline, what, you know, what does that really mean about how performant the AIs are?
- 2:40
If I meet the human baseline, does that mean that the AIs are now as performant or even more performant than, than expert humans in a, in a relevant sense that I, that I care about?
- 2:49
It's hard to interpret. You know, a- another thing that you see from this graph is that, um, benchmarks seem to have less and less time between coming online, sort of gi- giving any signal at all, and being fully saturated.
- 3:03
It's harder and harder to create benchmarks that have, uh, plenty of signal that, you know, might, might be informative to us about how capable models are for, for an extended period of time.
- 3:13
So we're, we're going to go about this a different way. First, we're going to gather human baseline data for diverse tasks spanning a range of difficulties. You should think of these humans as, you know, experienced experts, but on their first day or, or, or first week on the job.
- 3:29
These are not people with context on the tasks in particular. It's not exactly the kind of thing that's come up in their work before. But if it's a software engineering task, you know, they're a relevantly skilled general software engineer.
- 3:40
Same for the machine learning tasks and the cybersecurity tasks here that we'll talk about. The, the type of tasks come from these three, um, buckets or task distributions. HCAST, which is a collection of, um, software-based tasks seemingly requiring autonomy, you know, i- interacting with tools,
- 3:59
um, uh, interacting with the environment, thinking, thinking through the problem, not, not just this kind of Q&A style, um, style dataset. Um, the SWAA suite, which are these atomic problems.
- 4:09
These are problems that, you know, maybe GPT-2 can do, maybe, maybe it can't. Problems like, um, here are four files. One of them is called passwords.txt, which file contains the passwords?
- 4:20
And then on the other end of difficulty, we have ARIBench, which are challenging, novel, open-ended, um, machine learning research engineering challenges, which are, are very difficult even for top human experts.
- 4:32
In addition to gathering the, the human baseline data, we'll also, under as close to identical conditions as possible, measure AI performance for the AIs that we're, that we're interested in on the same set of tasks.
- 4:44
And then we're going to convert the time it takes for humans to complete these tasks into an estimate of AI autonomous capabilities, as I'll, I'll show you in a second.
- 4:55
Here's an illustrative diagram, in, in this case for Claude 3.7 Sonnet, which was the, the frontier model at the time that this paper came out. You can see that, you know, for the, for the very short tasks, something like four minutes or below, S- Sonnet is getting the answers correct, you know, essentially a hundred percent of the
- 5:10
time or, or maybe even here literally a hundred percent of the time. For the very hardest tasks, it's struggling. And then, and then there's some range where we're kind of in the middle, you know, we're somewhere between ten and, ten and ninety percent.
- 5:21
I'll say that this empirical pattern where models are less performant at tasks that take humans longer is, you know, it's not a fact of nature, but it's, it's something that we see pretty, pretty commonly, pretty, pretty robustly across models, at least on this task distribution, and I'd conjecture for, for other task distributions as well.
- 5:37
So we try and fit this dark purple line to, to something like this data on, on how long it t- took humans to complete the relevant tasks that the models are, uh, um, are attempting.
- 5:47
And then we call the point on the X-axis, this horizontal axis, this human time to complete axis, at which we predict the models will succeed fifty percent of the time, the time horizon of those models.
- 6:00
There, there's much debate in the fifty percent number. I can, I can talk later about the reasons why we chose that.
- 6:05
And then, and then we'll do the same exercise for the other models. So here I have, uh, Claude 3 Opus has a time horizon of something like four minutes.
- 6:12
That's where we're predicting that it has a success probability on this task distribution of 50%. For o1-preview, I'm seeing something like 15 minutes, so on and so forth. And then, of course, all these models, you know, they, they come out over, um, calendar time.
- 6:26
So if we plot the time horizon, the X coordinate on, uh, on, on this set of plots against, um, against calendar time, we find something like this. It looks, you know, kind of like, um, kind of like a, an exponential trend that's, that's going up at some constant rate.
- 6:40
In fact, it doesn't just look like an exponential trend. If we had a perfectly straight line here, it would indicate, um, a, a perfectly exponential trend. Um, we, we see something really remarkably steady.
- 6:50
Actually, much more steady than we were anticipating when we, uh, went about doing this research project.
- 6:58
And that's continued to be the case. So many of you will have seen updates that we've made of, of this graph on, on, on Twitter. This is going all the way up to GPT-5.1-Codex-Max.
- 7:07
So extremely recent. Um, the, the predictions from this, you know, shockingly straight line have, have held up very well, I think.
- 7:16
Taking a quick step back, what are benchmarks telling us? Or, or here, kind of benchmark-like evidence. Well, one thing is that AIs can succeed at what for humans would be exceedingly difficult tasks.
- 7:27
The tasks in RE-Bench are, you know, really far beyond my capabilities, uh, personally. And, and, you know, the AI's having a good crack at them some, some decent percentage of the time.
- 7:37
And the second, you know, kind of obvious, is that progress is rapid.
- 7:42
On the other hand, um, you know, how much, how much stock should we put in the, um, the evidence suggested by benchmarks? Um, what, what limitations might they have?
- 7:52
Lots, but here are, here are three that I'll note. One is, as I mentioned, these are humans who are, you know, expert in some relevant sense, but they're low context.
- 8:01
It's something like their, their first week on the job. They haven't seen tasks exactly like this previously. They just have some relevant experience. Presumably people who were more sort of, you know, not, not just having the relevant experience, but also highly familiar with, um, uh, with the, with the set of tasks, would perform the tasks even sooner.
- 8:18
And then we think relative to those people, the AIs were more performant.
- 8:23
The second is that benchmarks can be low ceiling. Even, you know, GPQA, I'll use that example again, um, we're, we're beginning to get to the point where, where that benchmark is, um, is totally saturated, not providing, um, additional information for marginal models.
- 8:40
Whereas Time Horizon is providing this nice way to sort of chain benchmarks together in, in, in some sense over time.
- 8:47
Um, but, you know, nonetheless, it's, it's still very hard to, um, uh, to create these ever harder tasks when the, um, when the time horizon of models is doubling every something like six to seven months.
- 8:58
So even Time Horizon might be, might be saturated in not too long, or the benchmarks underlying Time Horizon.
- 9:04
And the next one is, you know, not, not a concern that's limited to the, to the METR task, to the task behind Time Horizon. It's also true for SWE-bench.
- 9:11
It's also true for, for many of your, um, favorite agentic benchmarks. The, the problems aren't very messy in some sense. They don't require a ton of coordination with humans.
- 9:20
They're often in relatively small, contained environments where, where not much can go wrong. You know, not these sort of massive open-source code bases or, or, um, other ways in which the, the problems can involve more interaction with the real world or, or, or be messy in, in, in some sense.
- 9:36
Um, so we did this, we did this project. And then, um, early this year, we were, you know, we, we were trying to think about, um, uh, how can we attack some of these limitations?
- 9:46
What, what's a different source of evidence that, um, might have its own, own, own pros and cons, but, you know, importantly be more externally valid in, in the scientific jargon?
- 9:56
Perhaps field experiments are the answer. Some more economic style evidence. So here we might be interested in very high context developers who are expert on the kinds of tasks they're already doing.
- 10:07
Speed up or some notion of productivity boost. You know, it seems to have more signal through even some, um, superhuman according to benchmarks range. You know, perhaps GPQA is fully saturated, and you're getting a 1.5X, 2X speed up, something like that.
- 10:20
But you can still achieve a 3X, 4X, 5X speed up. Even, even after that, we, we maintain more signal.
- 10:27
And the last is that, you know, the, the tasks are messier. They are tasks that are coming up in people's real work. They're not, um, synthetic. They're not small and contained.
- 10:36
Um, this is a real deployment scenario. Here's what we're going to do. For this paper, we're going to gather 16 experienced developers on large, mature open-source projects that we'll go through in a second.
- 10:49
Each of these developers will on average complete about 16 tasks from their real work. These are, these are issues on the, on the relevant GitHub repositories, the kinds of thing that they might otherwise have completed with the, with the caveat that the very longest issues we're, we're not going to include.
- 11:04
The tasks will be randomly assigned to AI disallowed or AI allowed. AI disallowed, you know, means... It means what you think it means. It means software development in 2019.
- 11:14
It means no AI-powered tab autocomplete. It means no Cursor agentic coding tools. It means no LLMs via the web UI.
- 11:23
Or they can be randomly assigned to AI allowed, in which case everything's on the table. You know, a- any of the AI tools I just mentioned, or not using the AI tools.
- 11:31
If you're in the AI allowed condition, you're not compelled to use AI. You just have the option. And we buy these developers Cursor Pro. So, um, for the, for the most part, that's the tool that they're using with typically 3.6 or 3.7 Sonnet at the time, uh, which was the frontier model when we conducted this work.
- 11:48
And then we're going to record the time it takes for the developers to complete each task and see the degree to which they might save time when AI is allowed versus when it's not.
- 11:58
These are some of the repositories. Many of you will be familiar with them. We've got the Haskell compiler represented. We have scikit-learn. We have Hugging Face Transformers. These are on average a million lines of code plus.
- 12:09
They've been around for ten plus years. The developers who are going to be working on these repositories as part of this study are on average the third top contributor out of hundreds or, or even in some cases, thousands of contributors to these repositories.
- 12:22
They personally have been contributing to the repository for something like five years on average. These are top experts.
- 12:29
Some of you might have seen this graph too, and, and so the punchline's been spoiled. For, for the rest of you, um, we asked, uh, economics experts, machine learning experts, you know, these are people at major AI companies and labs, um, uh, top academics, um, some graduate students, so on and so forth, you know, how, how much
- 12:45
they expect developers to save time when they're using AI. They say something like forty percent or a little bit less. We asked the developers themselves, the study participants, how much they expect to be sped up ahead of time, and they say something like twenty-four, twenty-five percent.
- 12:59
Then we asked the developers after the study has been completed how much they think they were sped up in the past by AI being allowed on the issues they completed as part of this study, and they say that it will have sped them up by something like twenty percent.
- 13:13
And the punchline is that we find that developers are slowed down by nineteen percent. They take nineteen percent more time when AI is allowed relative to when AI is not allowed.
- 13:24
You know, when I first saw the data coming in, saw sort of early versions of this plot, um, I thought presumably the same thing that many of you might be thinking right now, that we've messed something up, um, that, that s- you know, something's gone wrong.
- 13:36
There's some, there's some issue in, in how we've set up the experiments. How could it possibly be the case? You know, at least these, um, uh, these developers have access to the zero points because they can not use AI at, at any time.
- 13:49
Um, so we pored over, you know, many, many, many, many, many hours of screen recordings from these developers working on issues as part of the study. We looked to dive into, um, a bunch of hypotheses that might explain what's going on and try to categorize, you know, the things that, that we think are going on versus not.
- 14:09
Um, many of this is, is listed in the paper. I'll, I'll just quickly go through some of the things that we think are contributing.
- 14:14
First, overoptimism about AI usefulness. That, that seems like an obvious one. You know, the developers e- even after the study's completed, they think that, um, uh, that AI's going to be helpful to their work.
- 14:25
It's, it, it makes sense they might overuse AI, um, on that basis. Um, two more, implicit repository context and high developer familiarity. You know, these developers are, are coming to these problems already knowing the solution to the problem.
- 14:38
They don't, they don't s- Um, they're, they're so expert in this work, you know, I, I, I imagine them as, as not trying to spend a bunch of time thinking through the solution that the, the AI can, can work through.
- 14:49
Instead, they're just limited by how fast they can type, um, which, which means that, you know, using AI, instructing AIs to do it, um, comes with some significant time cost versus how they might otherwise have spent their time.
- 15:00
I think many of us had the sense that AIs might be less performant on, on large and complex repositories, which is a different from this, difference from this benchmark style evidence or, or from, or from some previous work.
- 15:11
And then low AI reliability. You know, um, maybe the AIs are very performant on these kinds of tasks, but, you know, they're only performant, um, fifty percent of the time or eighty percent of the time, twenty percent of the time.
- 15:23
And so at the very least, you need to check their work afterwards, and perhaps even you need to spend time correcting their work afterwards, which is, which is something we see quite a lot on these issues.
- 15:34
One thing from the factors with an unclear effect that I'll, that I'll mention briefly, I happen to talk to people about later, is below average use of AI tools, which came up in the public discussion.
- 15:43
This, this is in the u- unclear column because it's sort of evidence, evidence for and against. Um, that, that's true for, for many of the things here. We don't have anything so conclusive to say.
- 15:52
We're still working on, on this line of work.
- 15:56
Here are some, here are some carry outs, all important. Um, first, you know, obviously, we do not provide evidence for all software developers or tasks. These are extremely experienced developers working on extremely complex, long-lived, open source repositories.
- 16:10
I, in my own work, you know, am not, um, a, as expert in the relevant sense as, as these people are. I'm working on much smaller repositories. Um, I, I feel more comfortable saying that even at this time, I was sped up by AI tools, even if, even if the developers weren't.
- 16:25
This setting is weird. It's weird for the same reasons that it's, that it's interesting, this, this unusual developer population.
- 16:31
Second, the experiment is concentrated in March twenty twenty-five. As I mentioned, uh, we know that AI progress is rapid. Um, perhaps this, this result will have already changed by the, by the time I'm giving you this talk.
- 16:45
So there's a kind of puzzle suggested, right? The, the benchmark style evidence is giving, um, a very impressive sense of what benchmark... o- o- of what AI capabilities look like today, whereas the more economic style, you know, I include labor market impacts, um, uh, uh, working here too, in addition to our, in addition to our field experiment,
- 17:03
look somewhat more bearish or, or unimpressive. You know, why, why is the former not, not translating to the latter? At least naively, there seems to be a clash. H- How might we go about resolving this puzzle?
- 17:15
So one possibility is that, in fact, we, we messed something up. This is, this is still live and on the table. Uh, you know, maybe the developers really are, um, uh, not very capable at using AI, and if we continue to run this experiment, as, as in fact we are, they'll, you know, learn more familiarity with the
- 17:29
tools and, and so get productivity benefits that they, they weren't getting at the time. I'm a little skeptical of that story, but, but, but that's one possibility.
- 17:38
A- Another that economists like to bring up is that we're not incentivizing these developers to finish quickly. We're paying them per the hour, um, which we do for external validity reasons.
- 17:47
Um, you know, looking through their videos, I, I really, uh, do not think that they're developing differently in, in accordance with these incentives. But, but that certainly is one possibility that's on the table.
- 17:58
You know, another, um, more statistical in nature possibility is, you know, th- this is a small study. You shouldn't, you shouldn't o- over update so much from small studies.
- 18:06
We are, we are doing, um, uh, bigger things that I'm excited to release at some point.
- 18:11
Okay, but let's, let's assume we haven't messed something up, and this is, uh, th- this, this is a result,
- 18:16
um, uh, that, that, that we think, that we think does hold up. How could we resolve the puzzle? So one possibility, you know, as I, as I alluded to briefly, is that reliability needs to be very high to save time, that you need to be getting, um, the, the answers to these problems the developers are putting in
- 18:33
correct, you know, something like 95, 99% of the time in order for developers to tab, tab, tab through and, you know, not, not, um, not spend lots of time verifying the AI's work, which, which of course, um, is, is pretty costly from a time perspective.
- 18:47
Another possibility is SWE-bench-like or algorithmic costless scoring at the margin versus mergeability-like scoring. SWE-bench scores are not trying to account for, you know, whether the code is build-on-able by, by other people in future or whether it's matching quality considerations that aren't, um, considered by the unit tests.
- 19:08
You know, perhaps AIs really are performant according to SWE-bench-like scoring but not performant according to this kind of more holistic, um, uh, holistic scoring that we might care about.
- 19:17
Low versus high-context baseliners. I, I, I mentioned, I mentioned previously the- these are just much more skilled humans. You know, relative to those humans, perhaps the AIs are less capable.
- 19:26
Task distribution, maybe these are just different kinds of tasks, you know, in particular less, less messy than the, than the benchmark-style task. Maybe that's explaining what's going on here.
- 19:35
Suboptimal capability elicitation. A huge amount of work has gone in at Meta to making the agents as performant as possible given the underlying models on, on all kinds of tasks and, um, you know, that involves churning through a, a, a load of AI tokens.
- 19:49
Perhaps that's, that's less the case for Cursor in particular at, at, at the time when we completed the study.
- 19:55
And then independence across tasks. Maybe it's the case that, um, you know, if humans can complete task A and task B, AIs can only complete task A but not task B, and of course can do task A faster, then it still makes sense to do...
- 20:09
for humans to do task A and task B, not delegate task A because, you know, th- they need to know the outputs. They need to know how, how task A was completed in order to reliably complete task B.
- 20:19
I think that's, that's part of what's going on. You need to maintain context as you're working through these subtasks.
- 20:25
Um, lastly, I will say that we are hiring, not just for this kind of work that you've, um, that you've seen being extended, you know, ever longer tasks, ever more ambitious, um, RCTs, um, even more sources of evidence from which we can triangulate the truth about AI capabilities, but also for, for much more besides.
- 20:42
You can, you can find this at meta.org/careers. In particular, I'm excited about research engineers, research scientists who might be, um, hiding in the current audience. We're excited not just, you know, for, for research types with academic experience, but very much for scrappy startup people as well.
- 20:58
And we're also hiring for a director of operations.
- 21:02
And with that, thank you very much for listening. [audience applauds] [upbeat music]