AI Engineer Europe 2026
The Art & Science of Benchmarking Agents
About this talk
Snorkel AI co-founder Vincent Chen outlines how agent benchmarks can close the widening gap between AI capabilities and evaluation. Drawing on Snorkel's $3 million Open Benchmarks Grants, he distinguishes scientific requirements such as expert-validated tasks, adversarial quality control, broad task coverage, and meaningful headroom from practical concerns including researcher adoption, policy adherence, and long-horizon organizational context. Examples include GPQA, MMLU, Terminal-Bench, and ARC-AGI-3, whose launch exposed a sharp human-versus-frontier-model performance gap.
Chapters
- 0:00Introduction: Snorkel AI and the agent-evaluation gap
- 3:50Open Benchmarks Grants and consequential agent benchmarks
- 6:20The science: GPQA task quality and MMLU coverage
- 10:08ARC-AGI headroom and policy-aware evaluation
- 17:28Researcher adoption, long horizons, and future benchmarks
Talk transcript
- 0:00
[upbeat music] Hey, everybody.
- 0:16
How's it going? Lovely. Well, I'm very excited to be here hailing from, uh, San Francisco. Uh, it's a little bit of a trek over, but I'm super excited to, to, uh, chat with you all, um, today after the talk, uh, and beyond.
- 0:29
Uh, my name is Vincent. I'm a research fellow and, uh, co-founder at Snorkel AI. And, uh, you know, today I'm gonna be talking about, um, some meta-evaluations for building benchmarks, the art and science of what we found to be really useful when building effective benchmarks.
- 0:45
I have the great privilege at Snorkel of working with both our researchers, uh, great collaborators in academia, industry, uh, and the open source community to build great benchmarks, and I wanted to share some of the learnings that we've had over the, you know, last few years on, on what really makes for benchmarks that, uh, shape the field
- 1:03
and, and move it forward. So a little bit about us. Uh, we're a frontier AI data lab. Uh, we have, you know, labs both at, uh, a-academic settings. You know, our, our co-founders, uh, have, have labs at Stanford, UW, Wisconsin.
- 1:17
We have an internal team, uh, for deployed engineers, applied engineers. And, uh, I mention this because we get a lot of exposure to both, you know, the academic frontier, work with frontier labs, and also real enterprises and, and companies who are deploying in practice.
- 1:32
So our focus as a company is on building the best datasets and environments to define and advance future AI capabilities, and we are in a fortunate spot where we get to play at this unique intersection of both, you know, the frontier academically, but also, um, uh, uh, to contact reality with our deployments, um, in enterprises.
- 1:52
So today, I wanted to talk about an asymmetry that we see in our real world deployments. There's real excitement around agents today. You know, we see it. Ev-every person in this room, I'm sure, has, has played with these, um, these, these agents, and we see a real progress marked by, you know, hill climbing on model cards.
- 2:10
We see the vibes are improving, right, and, and truly shifting, especially in coding. But when you ask individuals and enterprises or these, you know, large scale, uh, organizations if they're fully ready to let these agents loose and, um, you know, deploy them in high stakes environments, you, you get a little bit of hesitation.
- 2:27
And that's not to say the capabilities aren't there, but our ability to actually measure these agents in practice, um, that is, uh, falling behind of where the capabilities actually are.
- 2:38
This is one of the challenges and research questions that I think are actually one of the most important in the field and one of the ones that we're very interested in, uh, here at Snorkel.
- 2:48
So closing that gap, that evaluation gap, uh, we believe requires a toolkit, right? As I mentioned earlier, um, we're strong believers in field deployments, right? So this is, uh, you know, actually deploying engineers, researchers on our teams to, um, again, contact reality and, and, and work with folks, uh, to deploy these models in real production settings where
- 3:07
the stakes are high, right? These are finance settings, insurance, um, you know, healthcare settings where it's not just about a number, it's about real outcomes. Uh, and we're also big fans of other eval tools, right?
- 3:17
This is red teaming, private human evals, um, crowdsource labeling, um, a lot of the, the themes that we saw, uh, talked about, um, today which, which has been, um, awesome to see.
- 3:27
Um, but one of the things that we, uh, feel most strongly about is that, um, benchmarks, open benchmarks in particular remain a really critical piece of the measurement toolkit.
- 3:38
The best open benchmarks aren't just about, you know, taking a snapshot of progress looking backwards. They're actually about defining progress and shaping the field and setting a goalpost about where capabilities need to go.
- 3:50
And, you know, even looking at the last few months, right, benchmarks like Terminal Bench, uh, Meters long horizon benchmark, uh, ARC-AGI, these are really exciting and critical g-guideposts for where the field is going and, and, and as a result, the path to safe trustworthy agents, um, will, will really depend on more of these benchmarks in practice.
- 4:12
So what are we doing at Snorkel? One of the things that again I'm, I'm very fortunate to, to be able to be a part of is, uh, the Open Benchmarks Grants.
- 4:18
We recently a few weeks ago, a month ago, uh, deployed three million dollars to commit to open benchmarks and this is a really fun job. I get to work with the best academic teams, you know, builders, uh, to really accelerate and fund, um, the next wave of benchmarks that's gonna really steer and, and, and guide where, uh,
- 4:36
the field is going. We've had a wild reception so far. I'm a little admittedly behind on some, some reviews, uh, but we've been really excited to see what the community has come up with so far.
- 4:46
And in this talk in particular, you know, we've reviewed I think over one hundred and twenty applications so far spanning academia and industry labs. We wanted to share a few perspectives, a few learnings, um, over the fast-- past few months about what we view as one table stakes for, for useful benchmarks, right?
- 5:02
How do you actually build good empirical measuring sticks that are actually useful, you know, to, to measuring progress? And two, what really separates, you know, those benchmarks that are shaping the frontier, right?
- 5:13
What, what is the art and, and the, the, the science if you will of building really effecti- uh, effective benchmarks at the end of the day? So as I'm doing this I'll have a fun opportunity, uh, maybe this is a little too [REDACTED:origin], but to pull a Timothée Chalamet and h-honor some of the greats, you know, some
- 5:29
of the great benchmarks over the last few years, uh, that have really shaped the field, um, in talking about some of these axes and I hope that these, you know, themes resonate with you and also inspire a bit of, um, you know, kinda new thinking about, hey, how, how can we actually deploy some of the learnings we're,
- 5:43
we're all kind of driving towards in our day-to-day work to, you know, shape the field and, and move it forward. So two themes here again on the science side, you know, how do we actually build effective measuring sticks?
- 5:53
We'll talk about task quality, distributional control, uh, robust evals in general. Um, and on the, on the art side, right, really the, the differentiators for great benchmarks, um, how do you build benchmarks with a thesis on where the field is going that inspire new roadmaps, um, and that critically are, are built for this audience, right?
- 6:11
A researcher audience, a builder audience so that adoption is, is something that is way smoother, um, and, and a first class citizen, uh, for a bunch of these benchmarks.
- 6:20
So let's start with the science. Uh, this is again what makes for really effective measuring sticks as we've seen them in, in practice, i-in, in the pl-- the, the deployments that we see, uh, in industry and academia and with frontier labs.
- 6:34
So the first theme I wanna talk about is individual task quality, right? This is the idea that individual tasks need to be exceptionally rigorously validated, right? They need to represent real-world complexity.
- 6:46
They need, uh, well-posed, well-structured instructions. Um, they need verifiable solutions, uh, that ideally have been, um, actually validated by real-world domain experts. Um, one of the, the benchmarks here, GPQA, is one of my favorites, um, not just because it's been a very lasting and enduring benchmark that captures, you know, uh, uh, graduate level and professional knowledge, uh,
- 7:08
even to this date, right? You kinda still, still see this on model cards. But one of my favorite contributions is actually tucked away in the appendix. GPQA, uh, introduced one new adversarial qual-- uh, quality control mechanism.
- 7:20
So the idea was that not only do these tasks need to be well posed, um, they need to be tractable for other experts to solve. So they had a ve-very rigorous multi-reviewer protocol where there was an original author, you know, there were, there were reviewers and adjudicators in the loop.
- 7:35
There was opportunity for revision, right? These were tasks that were really pushing the frontier of knowledge, and it was non-trivial for any single expert to say, "Yeah, this is actually a good task or not."
- 7:45
And so developing this sort of rigorous adversarial, uh, quality control mechanism was one of the, uh, contributions I was most excited about here. And if you read the appendix, you also see that they introduced new incentive mechanisms, right?
- 7:58
Payouts were actually based on whether there was certain agreement. And, you know, uh, uh, coming from academia, you know, there's, there's, uh, some inspiration here from the peer review process as, as flawed as that is.
- 8:08
But, you know, this type of innovation around how you actually get really rigorous, you know, multi-expert quality control, um, leads to the type of outcomes that we see around individual task quality that, um, we, we see as a, a, a key foundation for any benchmark that matters at the end of the day.
- 8:27
Two is, uh, distributional diversity, right? This is the idea that for any benchmark that really matters, you wanna define a clear taxonomy for the domain, for real-world tasks, and distribute those tasks, uh, intentionally.
- 8:40
So this might be, hey, I, I captured a, a trace or, or a kind of real-world stream of, um, traffic a-along, um, you know, how my agent is operating in the real world, and I wanna really represent that distribution.
- 8:52
It could also mean, hey, I'm specifically characterizing and taxonomizing the failure modes that are, you know, paradoxically rare but, you know, disproportionately im-important in, in production, right? If you take classic self-driving settings, right?
- 9:05
Yellow lights or, you know, pedestrians or, uh, motorcyclists, right? Might actually show up way less than other, other types of scenarios, but are disproportionately important, you know, to get right in these settings.
- 9:16
And so defining that taxonomy, being really intentional about distributing tasks across it is one of the hallma-hallmarks of, uh, great benchmarks in our view. Um, MMLU, um, few years old now, uh, constructed a, a quite ambitious taxonomy of, you know, fifty-seven academic and, and professional domains across STEM, humanities, et cetera.
- 9:35
It's remained one of the lasting benchmarks for understanding graduate and professional level knowledge. And, um, again, a lot of this was, uh, as, as we, we believe, um, a result of really thoughtful and intentional, uh, taxonomy design and, and, uh, uh, building towards that.
- 9:53
The third axis here is around, uh, difficulty of individual tasks and model headroom, right? It's really important that the benchmark is unsaturated, that it exposes real soft spots, uh, in capabilities and reliably separates where models sit, uh, at, at the frontier.
- 10:08
One of my favorite plots, uh, is, is the one on the top right. Uh, this was, you know, all, all credit to the ARC Prize Foundation team. Um, [REDACTED:generic_id], you know, for a very long time was, um, unsaturated, right?
- 10:20
For, for several, um, months and years. And when there was the big reasoning push, you know, maybe eighteen, twenty-four months ago, um, we saw a massive leap in capabilities that actually corresponded to a real leap in model capabilities, right?
- 10:34
This was a benchmark that was intentionally designed to represent a type of efficiency or capability that's-- that, that, that humans have but models didn't have. And, and they really kinda captured, well, hey, there's a lot of model headroom here.
- 10:46
Humans can do this. Um, where is that gap? And again, lo and behold, uh, it correlated quite well with the recent, you know, o1 style reasoning push that has really dominated the field in the, in the past eighteen, twenty-four months.
- 10:57
Just a few weeks ago, uh, the ARC team just launched AGI, uh, ARC-AGI3, and again, at launch, they had frontier models under one percent. Every single task was, uh, human solvable, uh, to some degree.
- 11:09
And so it remains one of, I think, the, the most, um, uh, meaningfully, uh, e-exciting, uh, benchmarks in the space where, you know, any new model, you know, people are kind of awaiting, hey, how does it do, do on ARC?
- 11:20
And, and I think they did this quite well, right? The, the kind of model headroom here is, is really, really exciting.
- 11:27
This last axis here I wanna talk about, um, on, on the empirical measurement side is all about, uh, robust eval methodologies. Now this goes really deep, so just kinda capturing some of the high level ideas.
- 11:38
Benchmarks need to ideally go beyond accuracy to capture real world dimensions that matter, right? This is everything from cost, latency, you know, the quality of the reasoning traces, uh, uh, some of the intermediate steps and, and, and tool use.
- 11:52
Whatever dimensions actually matter for the capability at hand, uh, capturing those as reward or, or supervision signals is really critical. And measuring what it, what it claims to, um, is, is actually a non-trivial, uh, feat, uh, in, in, uh, you know, building robust and, and reproducible benchmarks.
- 12:10
So Terminal-Bench is a benchmark that we're, we're a big fan of. You know, it's had multiple evolutions over the year but-- over the years, but it w- it was a benchmark that was built to evaluate both task completion Of these multi-turn agents, they built a, a clever, you know, kinda user simulator.
- 12:24
Um, but also, you know, not just accuracy and completion, but adherence to policy constraints. So a model, for example, on the right-hand side, this was the one of the, uh, examples from the paper.
- 12:33
A model that books the right flight but violates fare class rules still fails. It's still a kind of, um, no-go at the end of the day. So, um, this notion of, hey, being intentional about what axes we actually care about, uh, what do we actually want to measure, and measuring that rigorously, um, is one of the hallmarks
- 12:49
that, uh, matters when you're, when you're building, um, these, these frontier evals.
- 12:55
So I wanna shift a little bit now to the differentiators, or r- what actually leads to the benchmarks that push the frontier. That's not to say, you know, anything I mentioned on the last few slides have not pushed the frontier.
- 13:05
These are just special characteristics that I view as, you know, critical to, to the benchmarks that are real research contributions that are really shaping where's the field going, where are the, where are all the labs gonna hill climb next.
- 13:17
Um, and, uh, this is, this is the art, the special sauce that, um, helps push us forward.
- 13:23
So one of the key hallmarks here is, um, these benchmarks should have a thesis, right? They should have a research question about a subspace of capabilities, about where the field is going.
- 13:33
It should revisit previous capabilities. Um, and the most ambitious benchmarks are really a statement about where the world is going. Terminal Bench, um, is one of these bets, right?
- 13:44
It was a bet on the CLI, not just for coding agents, but for general purpose computer use. And in many ways, I think this has turned out to be a largely correct and, and consequential bet, right?
- 13:54
As, as we're seeing, uh, teams at, at kinda Claude and, and, and Codex build their general purpose, you know, enterprise capabilities on top of these, um, coding and CLI-based tools.
- 14:04
We're seeing this bet pan out, and again, Terminal Bench remains one of the most robust and kind of, uh, most important benchmarks that are measured on, on all the recent, uh, model cards.
- 14:13
So again, this was a bet early on to say, "Hey, we think the CLI is gonna be really important, um, as a core interface, a core abstraction and affordance for agents to interact with the real world in a general purpose way."
- 14:26
And, uh, by measuring those capabilities, you know, I, I'd argue that it actually, um, helped accelerate, uh, you know, how, how the field is operating in this way.
- 14:36
The second piece here I think worth, uh, mentioning is, uh, the ability to kinda roadmap for the field. Um, a, a great benchmark, you know, one that really shapes where, where all of us are going is producing new roadmaps, right?
- 14:48
It's inspiring new attacks against research problems. It's helping folks ideate and come up with new ways for thinking about, uh, benchmarks and methods in general, and I think SWE-bench is a really phenomenal example of this, right?
- 15:00
It was a simple idea, as often the best ones are, are quite simple, right? H- how do you, um, kinda leverage, um, you know, existing, uh, coding, coding type capabilities via, via PRs?
- 15:11
Uh, and it spawned a new family of benchmarks, right? All the way from SWE-bench Light, Verified, Pro, Multilingual, Multimodal, et cetera, and its evolution, I think, is, is still very relevant today.
- 15:21
It's evolved how we think about coding agents, and one of the things that's been awesome to see with, uh, the SWE-bench team is how many, um, new research directions and kinda inspired benchmarks have come after it, um, in this coding space.
- 15:33
And arguably I'd say, you know, there's a lot more room, um, to kinda innovate on top of this as well. What do the new ways of, um, coding look like?
- 15:40
How do these types of workflows apply to vibe coding and kinda this new layer of abstraction that, that software developers are, are applying? Um, I think it's been really exciting to see the, the foundation that the SWE-bench team sets and how that's gonna, um, shape, you know, how we think about coding agents moving forward.
- 15:59
So this theme here I think is severely underrated, and this is the notion of researcher UX, right? I think the most prescient benchmark builders are committed to the researcher and builder experience.
- 16:10
This is to say it's really simple to run models and agents against, uh, your benchmark. It's really simple to contribute new tasks, to extend, and also it's, it's really simple to leverage some of the, the signals that you're getting for the benchmark for RL or, or kinda tuning on post-hoc.
- 16:27
I think this is really underrated. Uh, it's, you know, a classic product principle to make what you're building and putting out there easy to use by the community or, or by your core users.
- 16:35
And in this case, benchmarks have core users, which are other builders or researchers. And so really putting in time and attention to building those interfaces has been important for the adoption of some of the most important benchmarks, right?
- 16:48
To call out, I think the, the Stanford team, um, at CRFM, uh, built Helm, you know, several years ago, which I, I'd argue kinda pioneered a standardized modular harness for evaluating reproducible, you know, different scenarios as well as kinda models against a standard test bed of models.
- 17:05
Um, Terminal Bench 2.0 just a few months ago again shipped with Harbor, which has been in many ways the de facto, uh, harness and, and kinda evaluation infrastructure for teams who are building agents, uh, more broadly.
- 17:17
And so, you know, thankfully, we have a, a bunch of open source, um, software out there today, you know, ba- based on this principle. But as you're building your benchmarks, right, kinda considering, "Hey, how easy is this to extend?
- 17:28
How easy is it for the community to adopt and eventually kinda hill climb against this?" I think is a severely underrated factor for, for what makes for really high adoption of, of these frontier benchmarks.
- 17:41
So this is the full framework. Again, can go into more detail and, and please find me afterwards. But again, uh, what makes for really empirically meaningful, uh, measuring sticks, right?
- 17:50
It's task quality and attention to distributional control and diversity. Uh, it's, it's difficulty and model headroom and, of course, a robust eval methodology that measures the concrete axes that actually matter in, matter in practice and is, is intentional about it.
- 18:05
And of course, on the art side, right, these, these great benchmarks really have a thesis on where the frontier's going. Uh, they set roadmaps for the field, and they really prioritize researcher UX.
- 18:17
Now, before I wrap up, um, I wanna propose, you know, a few dimensions that, uh, we're really excited about at Snorkel, uh, that we think are really gonna encapsulate, you know, the, the next wave of, of benchmarks.
- 18:28
Um, tried to leave some more degrees of freedom here for creativity, but these are areas where we think there's a lot of room to push complexity, to push, you know, realism in benchmarks, and so wanted to share a little bit of our internal roadmap and thinking around where the field is going and where we need more benchmarks
- 18:43
and more contributions. So this is our point of view. Um, we think that, uh, the, the axes for the next great benchmarks are threefold. I mean, I'll go into a little bit more detail about what I mean here, um, in just a second.
- 18:56
But one, it's environment complexity, right? How complex, how realistic, how dynamic is the operating environment that these benchmarks are working? Are they representative of real world settings that a professional, that, you know, a scientist, that someone using these tools could actually use?
- 19:10
Two is autonomy horizon. Do these benchmarks represent realistic and frontier, uh, uh, horizon lengths that these agents are operating against? Are they capturing different points on the autonomy slider that are, again, representative of how users are using them as copilots versus fully autonomous agents?
- 19:27
Is this an intentional design in the benchmark? And three, capturing the wide range of output complexity. I think this is very underexplored today, right? Lots of chat-based or document-based outputs, um, not as much around nuanced, uh, kinda differentiated reward signals, right?
- 19:44
Real artifacts that show up, um, you know, in day-to-day work that we, that we represent. And critically, um, new artifacts, right? New types of form factors that we haven't even imagined about, you know, how agents interact with humans, how agents interact with each other.
- 19:57
So a little bit about each one. Um, the first one here, I won't go into all of these in, in significant detail, right? Environment complexity is all about capturing the real-world complexity that, that is, uh, i- in our, in our day-to-day working environments, right?
- 20:11
And this, this gap is often where agents fail today. Consider coding agents, right? A real code base has org specific policies, you know, lots of Slack context, uh, screenshots, flaky tool chains, you know, CI that's, that's kinda distributed, um, human reviewers with knowledge in their heads about what they like and what they prefer, um, many contributors in
- 20:30
parallel. Benchmarks today capture a fraction of this complexity and, you know, not just in coding, but other domains. There's a lot of excitement and, and opportunity to up the level of complexity and continue to drive what these models, um, can do to, to represent real world uses.
- 20:47
Two, again, on autonomy horizon, right? Uh, this is all about how long an agent can operate before it reliabil- uh, before reliability breaks down, right? Let's take a customer experience agent.
- 20:58
Um, you know, in, in many cases, right, uh, these, these agents may lose track, you know, o- of, of context that was, you know, delivered a few weeks ago.
- 21:06
Different integrations or product specs might, um, you know, change, uh, the actual spec or requirements for, for a particular model. Reorgs can kinda shift, you know, priorities midstream. Um, real world settings actually represent a lot more complexity that again, is represented in these kinda long-term continual learning type, type settings that represent kinda changes in state and environment.
- 21:28
Um, so we again think that there's a lot of room to contribute and, and build out new benchmarks that represent very, very long horizon in an autonomous agents.
- 21:37
And lastly, this axis is all about producing more complex work, more representative work, and also nuanced signals that can be used for not just evaluation, but reward signals during training.
- 21:47
This also has to be complex, and this gap is, is growing as well, right? Let's take again a software example, right, or, or, uh, a complex report for making strategic recommendations.
- 21:56
It's non-trivial and subjective, you know, to define, hey, what is verifiable about a good recommendation, a good strategic proposal, a good roadmap in general. Um, the nuances of this need to be captured well.
- 22:07
They need to capture organizational context, really good human judgment, um, and tomorrow's benchmarks really, you know, uh, we're, we're excited about, um, signals that capture all of these settings.
- 22:18
Trustworthy outputs, right? The ability for agents to actually capture their own uncertainty and, uh, define, "Hey, I'm actually not sure about this. I actually need to stop or kinda ask for more in- information."
- 22:28
Again, different types of outputs that aren't just, you know, a, a kind of plain text answer is something we're really excited about.
- 22:35
So again, uh, hopefully, you know, this inspired a little bit of thinking around, "Hey, what am I working on? How can I turn this into a meaningful benchmark?" Um, we are still accepting, you know, benchmarks, uh, in, in the Open Benchmarks Grant.
- 22:49
So if you're excited, uh, please reach us at benchmarks.snail.ai. Feel free to reach me directly, and, uh, we're really excited to see, uh, where the field is going, and again, to use benchmarks to not just measure, you know, progress looking backwards, but really shape where, where things are going moving forward.
- 23:04
Thanks for your time, and excited to catch up soon. [clapping] [outro jingle]