Your LLM Judge Is a Confident Liar: Building Better Verifiers — Browserbase

Read the talk

Your LLM Judge Is a Confident Liar: Building Better Verifiers

A browser agent scored 74% with WebVoyager’s official judge and 38% with a verifier grounded in task-specific evidence. Miguel González Fernández and Corby Rosset explain how to assign credit, catch unsupported claims, and test whether the judge deserves your trust.

From a talk by Miguel González Fernández and Corby Rosset

At a glance

Ideas worth remembering

  • Generate task-specific criteria, retrieve relevant screenshot evidence for each one, and compare the agent’s claims with the visible environment.

  • Keep process credit separate from task completion: an accurate attempt blocked by unavailable inventory can deserve full process marks while still failing the purchase goal.

  • Penalize an upstream mistake where it occurs, preserving credit for correctly performed downstream work without declaring the overall answer correct.

  • Evaluate a verifier through held-out human judgments and the quality of models trained on its selected data.

  • The Opus 4.6 autoresearch run completed roughly thirty experiments in a day but reached about 70% of the human-built verifier’s agreement; a run primed with human findings exceeded the team’s previous result.

The web outgrew deterministic checkpoints

Early Stagehand evaluations used deterministic environments: controlled sites and checkpoints that could establish whether an agent had completed a task. Miguel González Fernández, Browserbase’s agent-platform tech lead, describes how that approach became harder to maintain as models and the agent harness improved. More capable agents needed longer tasks, and the evaluation infrastructure required constant patching, generated static sites, and additional checkpoints. Corby Rosset, a Microsoft Research researcher, joined the effort to build better verifiers.

A live website makes those checkpoints fragile. Several paths can lead to a correct result; a product can disappear; a blockage can prevent an otherwise sound action. The desired outcome may be clear while the environment no longer supplies the fixed conditions needed for a deterministic test. An LLM judge offers a way to interpret these varied trajectories without hand-building every possible path.

Selected presentation frame from Your LLM Judge Is a Confident Liar: Building Better Verifiers — Browserbase at 145 secondsOpen full source frame
The speakers describe why changing websites and multiple paths to success make deterministic checks difficult.

But human experts reviewing the judges bundled with Online-Mind2Web and WebVoyager found confidently wrong decisions. That creates a problem beyond misleading benchmark scores. If the same judge supplies reinforcement-learning rewards or guides harness optimization, its mistakes become targets for the optimization loop. A system can improve at earning approval while failing to improve at the user’s task.

0:120:54
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:12 · section reference included

The same agent went from 74% to 38%

Microsoft’s Fara-7B browser agent made the discrepancy concrete. On the same benchmark, the official WebVoyager judge, GPT-4o, assigned a 74% success rate. The team’s Universal Verifier assigned 38%. The agent had not changed; the decision about which attempts counted as successful had changed. The lower estimate came from a verifier whose agreement with human labels the team separately evaluated.

Selected presentation frame from Your LLM Judge Is a Confident Liar: Building Better Verifiers — Browserbase at 254 secondsOpen full source frame
The same agent receives different success scores from the WebVoyager judge and the Universal Verifier.

Several weaknesses explain why judging a trajectory is harder than reading an agent’s confident final answer:

  • Missing rubrics. Without explicit success criteria, the judge has no clear structure for assigning credit.
  • Poor evidence selection. A judge may omit relevant screenshots or ingest every screenshot until the context becomes crowded or overflows.
  • Incomplete trajectory review. Omitting the final answer or action history removes information needed to compare the agent’s claims with what it actually did.

The Universal Verifier addresses these together. A stronger judging model still needs a well-defined task and the right evidence.

3:474:17
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

3:47 · section reference included

Retrieve evidence for each criterion

Consider the example request to book the cheapest flight from Seattle to Boston. The verifier first generates a rubric containing the criteria that define success. It then examines the agent’s trajectory and ranks screenshots separately for each criterion. The top K screenshots become the evidence used to decide whether that criterion was satisfied and whether the agent’s account contradicts the visible environment.

This makes evidence selection part of verification. A screenshot’s relevance depends on the question being checked. Instead of asking one model to remember an entire visual history and deliver a single verdict, the system narrows the evidence around individual obligations, while also considering the final answer and action history.

Selected presentation frame from Your LLM Judge Is a Confident Liar: Building Better Verifiers — Browserbase at 327 secondsOpen full source frame
A diagram shows task criteria linked to selected screenshots and verifier outputs.

The output preserves two different judgments. A scored rubric supplies a process score, allowing partial credit for steps performed correctly. An outcome Boolean, accompanied by an explanation, states whether the result meets what a reasonable user would expect. How does the trajectory become those two judgments? The diagram shows the per-criterion evidence path and the separate outputs. Screenshots support specific criteria before those judgments become an assessment of the whole task.

How it fits togetherFrom task requirements to evidence-backed judgments

For example, book the cheapest flight from Seattle to Boston.

Each rubric criterion receives its own ranked screenshot evidence. The verifier returns partial-credit process judgments and a separate task-outcome decision.

4:475:17
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

4:47 · section reference included

Grade the request and isolate the mistake

Rubric generation can introduce its own errors. In the Jakarta example, the request is to find a cheap hotel for specified dates, use its address to locate the closest coffee shop, and return that shop’s name and address. A generated rubric added a requirement to report the total hotel-stay price. Price matters when choosing a cheap hotel, but reporting the total was not requested. Penalizing its omission artificially lowers the score.

Dependent steps require another distinction. One task asks for the net worth of the member of NSYNC or Backstreet Boys with the longest last name. The agent selects Timberlake, whose surname has ten letters, instead of Kirkpatrick. That selection is wrong. But if it then accurately finds Timberlake’s net worth, the lookup step has worked for the person it selected.

The observable change is in credit assignment: fail the name-selection criterion and preserve credit for an accurate downstream lookup. Otherwise, one upstream mistake produces several penalties, obscuring which capability needs improvement. This partial process credit does not make the requested answer correct; the outcome judgment remains a separate question.

Selected presentation frame from Your LLM Judge Is a Confident Liar: Building Better Verifiers — Browserbase at 509 secondsOpen full source frame
The example marks the mistaken name selection separately from the later net-worth lookup.

Evidence checking also catches small factual substitutions that a fluent answer can hide. In an image-captioning research task, the agent reported a 6.2% improvement in CIDEr score while the underlying paper’s abstract stated 2.8%. The error requires comparing the returned claim with the source, rather than merely recognizing a plausible metric and a plausible number. Such discrepancies can also escape human reviewers.

6:166:46
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

6:16 · section reference included

An out-of-stock toy separates effort from completion

The Amazon plushie example explains why process and outcome should remain separate. The agent searches for the requested toy correctly and discovers that it is out of stock. The environment has blocked the purchase. Under the team’s grading scheme, that attempt receives full process marks for its best effort, while the outcome remains unsuccessful because the toy was not bought.

Selected presentation frame from Your LLM Judge Is a Confident Liar: Building Better Verifiers — Browserbase at 582 secondsOpen full source frame
The plushie example distinguishes the agent’s attempt from the unmet purchase outcome.

Different continuations call for different judgments:

  • Unavailable product. Credit the correct work while recording that the purchase goal was unmet.
  • Successful alternative. In the scenario described, finding a similar alternative can receive both process and outcome success.
  • Controllable error or hallucination. Penalize a mistake attributable to the agent, including unsupported claims about what happened.

The verifier follows an explicit failure-mode schema for these cases. An external blockage and an agent mistake can both end without the requested purchase, yet they supply very different signals for improving the agent.

9:169:46
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

9:16 · section reference included

Verify the judge with humans—and with training

To evaluate the verifier, human experts first reviewed the full trajectory and judged the evidence and action results. Only afterward did they see the Universal Verifier’s judgment and indicate agreement or disagreement. This ordering preserved an initial human assessment while allowing the verifier’s explanation to expose something the reviewer had missed. The resulting labels were released through CUAVerifierBench.

On the evaluated labels, González Fernández reports that false positives fell from almost half to zero. That is a result on the tested set, rather than a guarantee that future false positives cannot occur. The team also measured agreement using Cohen’s kappa, reporting 0.58 and describing verifier–human agreement as comparable to agreement between the two human annotators per task. Human labels supplied a reference, but the verifier sometimes identified mistakes humans had overlooked.

Selected presentation frame from Your LLM Judge Is a Confident Liar: Building Better Verifiers — Browserbase at 727 secondsOpen full source frame
The slide compares verifier–human agreement with agreement between human annotators.

A second test asked whether the verifier could improve training data. The team held dataset size constant at 3,000 or 9,000 trajectories and changed which verifier filtered those trajectories for supervised fine-tuning, or SFT. At 3,000 examples, selection using the Universal Verifier’s process score produced a better model than selection using the baseline verifier; the advantage also held at the larger scale.

Holding the number of examples fixed makes selection quality the useful comparison. The causal path is straightforward: a better verifier chooses better demonstrations, and SFT learns from those demonstrations. This experiment tests whether verification helps the downstream agent, extending the evidence beyond agreement with human judgments.

10:2510:55
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

10:25 · section reference included

Autoresearch accelerates experiments; human findings still help

Building the Universal Verifier took González Fernández and Rosset about three weeks and roughly thirty experiments, with prompt and code changes evaluated against human labels. Could an automated research loop rebuild it after their code and prompts were removed? In the reported run using Opus 4.6, the loop ran the same number of experiments in about one day, but reached only about 70% of the human-built verifier’s agreement. That figure is relative to the verifier’s agreement level, rather than a claim of 70% task accuracy. Rosset had not yet tested the proposed follow-up with another model, so the speed and agreement comparison applies to this run.

Selected presentation frame from Your LLM Judge Is a Confident Liar: Building Better Verifiers — Browserbase at 909 secondsOpen full source frame
A chart compares the automated verifier run with the human-built verifier.

A second automated run received the findings from the three weeks of human experimentation. That primed run exceeded the level the team had previously reached. The comparison separates two contributions: automation compresses the time needed to run experiments, while human discoveries can give the search a better starting point. In this experiment, combining them worked better than asking automation to rediscover everything.

The released work includes a preprint, golden labels in the Microsoft Fara repository, and experiment code. González Fernández also describes a new benchmark used to find remaining capability gaps as established web benchmarks become saturated. His concern is that familiar test tasks may offer little fresh signal when they resemble the training distribution; the harder benchmark serves as a working instrument for finding where agents still fail.

13:2113:50
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

13:21 · section reference included

More telemetry does not remove the need for held-out labels

The first audience question extends verification beyond the browser. Microsoft was working on a version for enterprise-style desktop tasks, with a further benchmark planned. The proposed verifier would inspect terminal information, telemetry, and logs alongside visual evidence. Those additional channels expand what the system can check; the underlying approach still connects task criteria to evidence of what happened.

The final substantive question concerns overfitting: repeated prompt improvements can make a verifier agree with the labels used during development without making it a better judge on new trajectories. The team’s answer was to separate development labels from evaluation labels. Of approximately 150 trajectories, Rosset labeled about 50 and used them during the roughly thirty experiments. The other 100 were held out, with two paid human annotators labeling each trajectory.

The reported CUAVerifierBench results came from those held-out trajectories, on which the team did not iterate. Keeping at least two-thirds of the labels out of development gives the final measurement a different job from the optimization score. Rosset’s practical recommendation is to train human reviewers to perform the verification task themselves, then test whether the system agrees with them. Building a judge requires both a standard to learn from and examples that remain unseen until it is time to measure.

17:3518:05
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

17:35 · section reference included

Resources

Read the complete timestamped transcript
  1. 0:12

    Good afternoon, everyone. Um, thank you for coming. Hope you've had a great expo so far. We're in the last stretch, but I know it's been a wonderful, wonderful expo, at least for me. So hopefully you're learning a lot and we can teach you a little bit more. Um, so I'm Miguel. I'm the tech lead for, uh, our agent platform at Browserbase. For those of you who don't know, Browserbase is a infrastructure company for deploying agents on the web. Um, it basically unlocks all of the web, not just what is easily

  2. 0:42

    accessible, uh, by managing browsers in the cloud and all of the runtime that, uh, requires to do agentic web automation at scale. So, um, but-

  3. 0:54

    And I'm Corby. I'm a researcher at Microsoft, and, uh, I've been collaborating with, with Browserbase to help build really good verifiers.

  4. 1:02

    So today we're here to talk about some research that we published about a month ago. Um, and it is regarding evals, it is regarding RL environments and LLM judges as verifiers. So early on-- So I've been working on this open source framework called Stagehand for about a year and a half now, and from the beginning, as any good agentic product needs to have, we were building evals to make sure that the changes that we

  5. 1:32

    were introducing in the framework and that the model capabilities continue to improve and hill climb on those evals. And at the beginning, all those evals were deterministic environments. Um, it was relatively simple. Model capabilities were some-some somewhat constrained, and so we were able to build a lot of deterministic environments. As model capabilities continued to improve, as we continued to make improvements on the agent harness, that just didn't scale. It-- we found ourselves constantly patching or generating static sites, um,

  6. 2:03

    coming up with longer trajectories with a lot of checkpoints to make sure that the evals were deterministically verifiable. And so there's many different project-- problems with automating the web at scale and verifying that what you intended to do actually succeeded. And it's that the web is very open-ended. There's not just one path to correctness. Um, there is no ground truth. Sometimes the web changes, and the product that you were checking for before is no longer there, so that

  7. 2:33

    breaks the whole determinism of your environment. And there's blockages. There's a lot of the environment errors that you face that don't let you get sort of a verifiable deterministic signal. So we recurred to what most people do, which is trying to use an LLM judge to help improve, um, and scale this operation. But we quickly found out that a lot of the leading benchmarks, and some of you may know about OSWorld, uh, for computer

  8. 3:03

    use, but there's some analogous benchmarks for web use called Online-Mind2Web and WebVoyager. Those come with LLM judges as verifiers. And what we noticed when we were using our sort of human expert annotators to verify the verifier is that in many instances, they're very confidently wrong, and that leads to results that you can't really trust.

  9. 3:30

    So very quickly, we realized that even if you were using this as an RL reward or as a signal to improve your harness and auto research, you're not really training a better agent, you're just training a more confident liar.

  10. 3:47

    And so in a real world use case at Microsoft, we trained this Fara-7B model, which is a, a web browser agent. And the same model on the same benchmark, we judged it according to the official WebVoyager judge, which is GPT-4o, uh, and it said that it had seventy-four percent, uh, success rate. But there's a huge gap between the real truth, which is when we used our new universal verifier, uh, which has high agreement with human labels, that, that number quickly becomes thirty-- like, thirty-eight percent. So there's a

  11. 4:17

    huge gap between what the ex-existing verifiers say and what is actually the ground truth, and so we needed to find a way to, to close this gap. The existing verifiers have lots of weaknesses. Um, they're u-- they use smaller and, like, much dumber models, uh, as the LLM, as a judge. They use o4-mini or GPT-4o. They don't even use rubrics. So rubrics are the most critical thing that you need to have because you need to assign credit to where credit is due. They don't look at the, the relevant screenshots or they, or they try to look

  12. 4:47

    at all the screenshots and they quickly get lost and overflow the context window of the LLM as a judge. And sometimes they don't even look at the final answer or the action history of the model. So our verifier checks all of these boxes. Um, and this is kind of how it works, uh, at a very high level. You can, you can look in the paper for more details. But given a task like book a cheapest flight from Seattle to Boston, the first thing we do is we generate a really good rubric, and I'll give some examples of that. And a rubric has like maybe N different criteria of what success looks

  13. 5:17

    like. And then given that, that criteria in the rubric, we look at the agent's trajectory. Uh, all of the screenshots in this case are the, the evidence of the ground truth. And we rank what are the most relevant screenshots in the trajectory for each criterion. And then we use that group of, of top K evidence to determine whether that criterion was met or not and if there's any like differences or contradictions between what the agent said it did and what the state of the environment

  14. 5:46

    actually showed. And then given that, we output, uh, a scored rubric which we call like a process score because it gives partial credit in some cases, and we give an-- we output a, uh, an outcome Boolean, uh, value to turn-- to basically say whether the agent accomplished the task according to what a reasonable user was ex-- would expect. Um- And so this is an example of what the rubric, uh, the rubric output is. It's basically a list of criteria, um, as

  15. 6:16

    well as the outcome verifier is basically a true-false flag with an explanation as to why a reasonable user would expect th-um, this trajectory to have succeeded or not. There were four guiding principles when we created this universal verifier. One is around rubric creation, which is we wanted to grade only what was asked and not any extraneous criterion. Uh, we also did not want errors to cascade from one rubric, uh, criterion to another. We wanted them to be isolated, and I'll give an

  16. 6:46

    example of that. You also wanted to look at the ground truth screenshots, like this is a must. The agents will often overconfidently claim that they did something when in fact they did not do it, and so you have to look at the ground truth state. And then we also needed to separate, um, what the agent could control versus what the agent could not control. The agent operates in a web browser. The web browser is an environment that it doesn't always have ability to control, and I'll show some examples of that too. So this is an example of what a good and bad rubric looks like. The task

  17. 7:16

    here is to, uh, find a cheap, uh, hotel in Jakarta for these dates, and then use the hotel's address to search for the closest coffee shop, and then output the name and address of that coffee shop. Now, a bad rubric, which we did see happen, would output a criteria saying like, "Please tell me the total price for the stay at the hotel." This is actually an extraneous criterion that was not asked for in the task, but we saw a lot of rubrics naively generated do something like this, and they would artificially deflate the

  18. 7:46

    scores because the agent didn't do something it wasn't asked to do. Um, when it comes to not cascading errors, um, th-there are very subtle mistakes because a lot of tasks build on each other, uh, if they are like multi-step tasks. So in this example here, the task was to, um, determine the net worth of the individual with the longest last name from among the members of NSYNC and Backstreet Boys. Okay? And so the, the agent in this case

  19. 8:16

    mistakenly thought that Timberlake had the longest last name, which is only ten letters, whereas Kirkpatrick actually had the longest last name. So the, the model made one arithmetic mistake in one criteria of the rubric, and that should not cascade to the next criteria of the rubric, which is report the net worth of, of the person that you chose, right? So if you're able to report the net worth of Timberlake, uh, accurately, then you shouldn't get penalized in this criteria. You should only get penalized in that one. Hallucinations is the biggest one,

  20. 8:46

    uh, that we really wanted to focus on. So in this task, this is an example of a very subtle hallucination, uh, where you wanted to find some information about some image captioning model, and the agent said that this model had plus six point two percent in some CIDEr score. But in reality, the abstract of the underlying paper said it was only two point eight percent in CIDEr score. And so this was a, like a very subtle mistake that even humans will not catch unless an LLM will flag it. Um, so this

  21. 9:16

    kind of gives, uh, an overview of, of how process scores and outcome scores differ. Um, in the case we, we mentioned one of the principles was we wanted to separate controllable versus uncontrollable, um, failures. So if the task is to buy, say, this, uh, plushie toy from Amazon, and the agent accurately searched for the plushie toy but found that it was out of stock, it could not continue. And so here we say that the agent gets full marks for doing its best

  22. 9:46

    effort, uh, to achieve the goal, but the outcome was still not met because it couldn't buy the thing that it wanted to buy because it was out of stock. Um, so we basically enumerated a bunch of scenarios as to what could happen and how do we assign credit. If the agent was able to find a similar alternative plushie toy in a different way, um, it would still get success there, uh, for both, for both cases, and it would get penalized if it made a controllable mistake. Uh, it would get penalized if it made hallucinations, and so on and so forth. So we basically enumerated a, a

  23. 10:16

    schema of like all the possible failure modes and how we would assign credit to that, and the universal verifier adheres to those, to those schema. So-

  24. 10:25

    So basically, in, in order to evaluate the verifier and really tell whether or not we were able to hill climb the accuracy of that verifier, we started designing a lot of experiments with human experts and built a whole platform to collect data on what we call the ground labels that we've also released under CUAVerifierBench for others to train upon. But the process that we followed was the following. At first, we would show the full trajectory of

  25. 10:55

    evidence to the human verifiers, and the annotators would judge based on the evidence and the result of all the actions that the agent took. After that, they would be prompted to af-- seeing the judgment of the universal verifier to say whether they agree or disagree with the human-- uh, with the universal verifier. And we found this to be a very reliable signal to tell whether or not

  26. 11:25

    we were working in the right direction because... I'll jump to this one straight up, but the verifier started correcting humans at one point. They-- It would identify things that the humans were missing. And so ultimately, the result of the, of the universal verifier on these golden labels was that we were able to reduce false positives from almost half to zero. And Cohen's kappa is a metric to measure inter-annotator,

  27. 11:55

    inter-annotator agreement. Um, and what we can see here is that the hu- universal verifier agrees with humans as often as humans agree with one another. So we can see that point five eight Cohen's kappa score, uh, versus the two annotators per task, um, w-where we computed Cohen's kappa as well to see how much they agreed with one another

  28. 12:20

    Another way, another way to, uh, verify the-- another way to verify the quality of our verifier is to actually train on data filtered by it. So what we did was we did-- ran an experiment where we held constant the number of training examples. In this case, it was three K or nine K training trajectories. But we filtered those training trajectories based on whether our verifier said they were pass or not. And so when you're doing SFT, which is what we did here, you wanna train only on SFT examples where

  29. 12:51

    they were true, uh, like successful trajectories. And so if you train on three thousand trajectories filtered by the process score of our universal verifier, you will get a much higher quality model than if you were to train on a worse verifier, like our own, uh, a baseline verifier. And this held at larger scales as well. And so this basically-- this experiment told us that, uh, if you have a higher quality verifier, it means you can filter higher quality data, uh, and training on that data leads to a higher

  30. 13:21

    quality model. And so this was like the ultimate experimental proof, in addition to the, um, uh, to the, to the human agreements, uh, that this model, this verifier is working and is, and is reliable, um, in a real production, uh, scenario. Um, the last thing we did kind of as a side project is we wanted to determine whether AI in auto research can build the same verifier that we built. So basically what I did, and what Miguel and I did, is we

  31. 13:50

    sat down for three weeks, um, to build the universal verifier and to tweak the prompts and to write the code for it. And we basically ran about thirty different experiments over the course of three weeks to determine whether the verifier we were building agrees with human labels. Okay? And so that's what the Cohen Kappa score here on the y-axis is measuring. It, it measures whether an individual verifier system is agreeing with the human labels. And so the blue line here is the model Miguel and I-- the, the system Miguel and I

  32. 14:20

    were, were creating, this auto research, uh, this, uh, universal verifier. And then we wanted to see if, if we stripped everything that we did away, could AI, um, in an auto research loop, build the same verifier that we built with the same level of quality and fidelity? And that's what the red and green lines here are. The red line is, if we took away all of our code and prompts, could AI replicate that? The conclusion that we got from this is that while we took about three weeks to build this verifier, uh, an auto research loop

  33. 14:50

    could do it in about one day. It ran the same number of experiments in about one day. However, it only reached about seventy percent of the agreement that our verifier, uh, was able to reach. So there's still a gap there. And so the-- kind of the conclusion that I would draw here is that, uh, you can use auto research to build, uh, metrics and verifiers, um, at least to help you speed up, uh, experimentation. Um, but you probably still need some level of human intervention here. Um, then

  34. 15:20

    again, these, these results were from Opus four point six. I haven't tried it with Fable yet. Maybe it will do better. But it's a pretty powerful baseline, and it got pretty far, um, along the way. So not only can auto research help you build a model, but in this case, auto research can help you build a verifier as well, which we thought was a pretty cool, um, thing. And, and building verifiers is as important as building the models themselves, is one of the takeaways I wanna leave you with. So...

  35. 15:47

    Yeah. And just to double emphasize, um, the green line is the auto research verifier, um, primed with a lot of the findings that we had gotten from the three weeks of experimentation. And in that optimization, it was able to reach higher percentages than we were ever before. So combining that human intuition with auto research loops has proven to be the most powerful recipe for many research tasks. Um, so

  36. 16:17

    the-- all of the work is open source. The paper is published under this, uh, as a preprint, um, on arXiv. Um, and the Microsoft Fara repo contains the golden labels on CUAVerifierBench. It contains the code to run the experiments. But beyond that, um, this research also yielded an entire new benchmark. So we've talked briefly about Online-Mind2Web and WebVoyager. I

  37. 16:47

    spend a lot of time working with labs, helping make sure that we can provide a quantitative signal to improve their models. A lot of those open source benchmarks are very quickly getting saturated, and it is very difficult to provide any sort of signal because it feels like the, the test data is already in the distribution of the training data. And so this new benchmark has proven to have the biggest gap to complete success, and it's my daily driver

  38. 17:17

    to make sure that I can tell whether model capabilities are strong, whether they lack, and how to continue to close the gap. So...

  39. 17:27

    Did you have more slides?

  40. 17:28

    No.

  41. 17:29

    Okay.

  42. 17:30

    We can take a little bit of, I think, questions, but, uh-

  43. 17:33

    Yeah

  44. 17:33

    ... we're good on time.

  45. 17:35

    Thank you. Yes. So my question is on universal verifiers. Have you tried other things like just computing tasks and other kinds of services? Yeah. So the question is, the question is whether we've used the, uh, the universal verifier for other kinds of tasks beyond just web, web tasks. So one thing we're working on in Microsoft right now is to apply a version of this verifier for desktop tasks, like things that you would do on your laptop, like kind of enterprise workflows. And so we hope to publish another benchmark, uh, of not

  46. 18:05

    just web data, but also, like, enterprise-style desktop data. Uh, and we would have basically the sim-- the same kind of verifier for that too. It would look at a little bit more information, because on desktops you have the terminal, you have more telemetry, more logs than just the browser. So the verifier would take those things into account as well. Yeah. Any other questions? Yeah.

  47. 18:29

    So I personally never built a verifier, so I'm just curious. Um, but when you did the presentation, I was thinking about, like how do you think about, uh, ensuring that you're not overfitting, uh, you know, trajectories? Because if you do a supervised learning trajectory-

  48. 18:44

    Yeah

  49. 18:44

    ... you also have-- You did show some diverse use cases, so it tells me like you are definitely thinking about it. But what, what would you say in terms of like any guidance or prompting about not overfitting when you're building versus supervised learning?

  50. 19:00

    Yeah.

  51. 19:02

    How do you do it? Uh, how-

  52. 19:03

    Yeah, so the question is about how do you make sure you're not overfitting the verifier to the human labels, because that is a problem. Uh, Miguel and I spent a lot of time, uh, working on this question. And so, like one thing that we did was we, uh, kind of held out, uh, two sets of labels. One-- We, we basically had a set of like a hundred and fifty trajectories. Um, on about fifty of those trajectories, I labeled them myself, and I used them to hill climb the-- in those 30 experiments. But for the other hundred trajectories, they were

  53. 19:33

    basically held out to us. They were done by humans, uh, that we had paid, uh, with like two X overlap. So once our verifier was done being built, we gave it to those human annotators, and each human labeled each trajectory. Sorry, two humans labeled each trajectory, and then that's what we published in CUAVerifierBench. And so those are the numbers that we report here. Um, so we're pretty confident that we didn't overfit because we held out at least two-thirds of the labels.

  54. 20:03

    Um, like, we did not train, uh, like iterate on them. Yeah. So that's very important, is you don't wanna train on, like-

  55. 20:14

    Thinking about your own, like, like personal verifier, I'm gonna have to hire new people to be able to actually like

  56. 20:24

    Yeah. Yeah, yeah, the gold standard is, uh, is to hire humans, um, and train them to do the t- the verification task themself, and then see whether your system agrees with it. Yeah. Any other questions? Anyone wanna help us build more benchmarks?

  57. 20:43

    Yeah. Oh, yeah. Okay. Okay. Yeah, yeah, yeah.

  58. 20:49

    Awesome. Well, thanks a lot. Uh, hopefully you enjoy the rest of the expo. It was a pleasure. Um-

  59. 20:54

    Thank you so much.

  60. 20:55

    Thank you so much.

  61. 21:13

    Yeah, yeah.