Corby Rosset’s research connects web search, language-model self-improvement, and computer-use agents. His work at Microsoft Research has included helping build Bing’s People Also Ask feature, developing Direct Nash Optimization with collaborators, and designing the expert-built Universal Verifier. Across these projects, he studies how systems can learn from questions, preferences, and evidence of whether they accomplished a user’s task.
From retrieving documents to understanding questions
Rosset’s earlier research addressed the machinery behind web search. In 2018, he and collaborators explored reinforcement learning for query evaluation: teaching a system to choose how it scans a search index, reducing the work required to find promising documents while preserving candidate quality. This placed machine learning inside the retrieval process, where efficiency matters before a user sees any results.
His work then moved toward helping people decide what to ask next. For conversational question suggestion, Rosset and his collaborators trained ranking and generation systems using patterns in search sessions. A useful suggestion needed to do more than resemble the original query: it should advance the user’s investigation. Their 2020 study reported different gains depending on placement in Bing’s results: clicks on suggested questions increased by 8.9% for suggestions at the top of results and 6.4% for suggestions at the bottom.
That interest continued in Researchy Questions, a dataset released by Rosset and collaborators containing roughly 100,000 demanding Bing queries. These questions require multiple perspectives and often several smaller investigations. The accompanying research found that breaking questions into subquestions helped GPT-4 answer them, exposing a gap between performance on familiar question-answering benchmarks and the harder information needs people bring to search.
Learning from feedback and changing preferences
By 2024, Rosset’s research also encompassed how smaller language models learn from feedback. He co-authored Orca-Math, which combined synthetic mathematical problems with an iterative cycle of attempted solutions, feedback, and preference learning. The seven-billion-parameter model achieved 86.81% accuracy on the GSM8K grade-school mathematics benchmark with a single generated answer, without external tools or a verifier at inference time.
Rosset was first author of Direct Nash Optimization, an iterative method that trains models through comparisons of their own responses. The method samples fresh answers, compares them through a preference evaluator, and trains toward the preferred behavior. Because the responses change as the model learns, the feedback can follow its current mistakes rather than relying entirely on a previously collected set of examples.
His explanation of Direct Nash Optimization emphasizes correcting mistakes that imitation alone may leave intact, while avoiding some of the complexity and instability of conventional reinforcement learning. The evaluator’s preferences therefore become a central part of what the model learns—a concern that also shapes his work on browser agents.
Checking what computer-use agents actually accomplish
Rosset contributed to Fara-7B, an open-weight computer-use model that reads screenshots and acts through predicted screen coordinates. Its training pipeline generates web tasks, attempts solutions, and filters the resulting interaction histories. In this setting, a browser agent’s account of its work can diverge from what happened: it may claim success despite an unfinished form, an incorrect selection, or a blocked website.
Evidence for each requirement:Task-specific rubrics define what success requires, and the system selects relevant screenshots for each criterion. This grounds judgment in visible results rather than treating the agent’s final account as proof.
Independent criteria: Evaluation should grade what the user asked for and isolate errors. One early obstacle should not automatically multiply into penalties for every later criterion; each judgment needs its own evidence.
Execution versus completion: Separate process and outcome scores distinguish how well an agent acted from whether it achieved the user’s goal. An agent that finds the correct product but encounters a login barrier can receive credit for completed steps while still failing the overall task. This also helps distinguish failures the agent controls from obstacles it cannot resolve.
The research tested the verifier against human labels, reporting agreement that approached the agreement between human annotators. That comparison matters because a more elaborate evaluator is useful only if its judgments better reflect actual task completion.
Rosset’s contribution to Fara1.5 extends this work into a larger training pipeline. It combines live websites with simulated environments, generates interaction histories, and evaluates correctness, efficiency, and adherence to actions requiring user approval. His progression from search suggestions to browser agents keeps the user’s intended result central: finding useful next questions, improving responses through preferences, and checking actions against observable outcomes.
A browser agent scored 74% with WebVoyager’s official judge and 38% with a verifier grounded in task-specific evidence. Miguel González Fernández and Corby Rosset explain how to assign credit, catch unsupported claims, and test whether the judge deserves your trust.
Generate task-specific criteria, retrieve relevant screenshot evidence for each one, and compare the agent’s claims with the visible environment.
Keep process credit separate from task completion: an accurate attempt blocked by unavailable inventory can deserve full process marks while still failing the purchase goal.
The Opus 4.6 autoresearch run completed roughly thirty experiments in a day but reached about 70% of the human-built verifier’s agreement; a run primed with human findings exceeded the team’s previous result.