Miguel González Fernández builds tools for browser agents and methods for judging whether those agents accomplish their tasks. His contributions include Stagehand, Browserbase’s open-source browser automation SDK, and the Universal Verifier research developed with Microsoft Research. He served as tech lead for Browserbase’s agent platform at the time of his 2026 presentation with Corby Rosset. His work connects the mechanics of operating software with a harder question: how to distinguish a convincing attempt from a completed task.
From open models to browser infrastructure
González Fernández’s early public projects explored how open models could control familiar environments. In late 2024, he contributed a Llama 3.2 Vision browser-agent quickstart to Meta’s Llama Cookbook. The notebook combined screenshots, natural-language instructions, browser navigation, and persistent sessions to demonstrate an agent interacting with websites. It was merged in December 2024.
His Llama and Home Assistant project explored using a small, locally running model as a private smart-home assistant. He also built a Home Assistant MCP server exposing controls for lights, climate systems, locks, alarms, and humidifiers. Both projects addressed a practical requirement for useful agents: connecting model instructions to actions in an environment people already use.
His Browserbase work extended that interest into shared browser infrastructure. González Fernández is a major contributor to Stagehand, which combines conventional browser controls with natural-language actions and structured extraction. Developers can ask it to find an input, perform an action, or extract a table into a specified schema. Its self-healing operations refresh their interpretation when a website changes. He also helps coordinate substantial community contributions to the project.
By October 2025, his work included evaluating computer-use models with Google DeepMind. He and Sean McGuire described a testing effort spanning more than 200 experiment runs and approximately 4,000 browser hours. They emphasized comparisons that other developers could inspect: shared constraints, public trajectories, human-verified results, and an open evaluation runner. The following month, González Fernández described evaluation of Microsoft’s Fara-7B, using consistent browser environments and human review of screenshots, action histories, and final task states.
Measuring what agents actually accomplish
Live websites complicate evaluation because the environment changes independently of the agent. González Fernández and McGuire give the example of a hotel-search task becoming impossible when its requested dates have passed. Quietly removing such tasks leaves teams measuring different things under the same benchmark name. Their evaluation approach publishes failures alongside successes and makes traces available for inspection. Cloned websites offer more control, but can omit the popups, loading delays, and other complications agents encounter in practice.
The Universal Verifier collaboration addresses a related problem: the judge can be wrong even when the task remains valid. A prescribed action sequence cannot capture every route to success, while an LLM judge can accept plausible behavior without checking that the requested outcome occurred. In the joint presentation with Rosset, the benchmark’s original judge reported 74% success where the improved verifier reported 38%. That gap matters for training as well as measurement: rewarding an incomplete task teaches an agent to repeat behavior that failed the user.
Task-specific criteria: Rubrics are generated before examining the agent’s behavior, reducing the temptation to invent criteria that fit what it happened to do. The verifier grades what the user requested rather than adding unrelated requirements.
Evidence for each judgment: The verifier ranks relevant screenshots for each criterion and checks them for evidence of completion. This helps catch cases where an agent’s account sounds successful but the visible state does not support it.
Independent errors: Criteria are evaluated independently so one mistake does not cascade into repeated penalties for subsequent steps.
Separate process and outcome scores: An agent can navigate correctly and still encounter an unavailable product or a login barrier. Its execution may deserve credit while the user’s request remains unfulfilled. Separating the scores distinguishes agent mistakes from environmental obstacles without declaring an unfinished task successful.
In the collaboration’s reported experiments, the verifier reduced false positives to near zero and reached agreement with human judgments comparable to agreement between human annotators. González Fernández’s account of the work connects that improvement to more reliable training rewards: better verification changes which behavior a learning system is encouraged to repeat.
Making model choice easier to revisit
González Fernández also addresses the infrastructure decisions that shape an agent’s cost and performance. In his Model Gateway introduction with Harsehaj Dhami, the authors explain why developers can remain tied to an early model choice: changing providers means managing additional accounts, keys, billing arrangements, and configuration. Browserbase’s gateway centralizes provider access so Stagehand developers can switch models as their speed, cost, and reliability requirements change.
Across these projects, González Fernández’s work has developed from demonstrations of models controlling browsers and household devices into infrastructure for operating, comparing, and judging browser agents. The defining concern is whether automation remains useful under real conditions—when websites change, models need replacing, and an apparent success still needs checking.
A browser agent scored 74% with WebVoyager’s official judge and 38% with a verifier grounded in task-specific evidence. Miguel González Fernández and Corby Rosset explain how to assign credit, catch unsupported claims, and test whether the judge deserves your trust.
Generate task-specific criteria, retrieve relevant screenshot evidence for each one, and compare the agent’s claims with the visible environment.
Keep process credit separate from task completion: an accurate attempt blocked by unavailable inventory can deserve full process marks while still failing the purchase goal.
The Opus 4.6 autoresearch run completed roughly thirty experiments in a day but reached about 70% of the human-built verifier’s agreement; a run primed with human findings exceeded the team’s previous result.