AI Engineer Code 2025
Why 99% Accurate Browser Agents Still Fail — Derek Meegan, Browserbase
Read the talk
Why 99% Accurate Browser Agents Still Fail
Derek Meegan of Browserbase explains how small risks compound across browser workflows, why completed transactions matter more than individual runs, and how an insurance-portal architecture reduces the decisions a model must make.
From a talk by Derek Meegan
At a glance
Ideas worth remembering
Under independent step success, 99% accuracy compounded across 100 steps yields only about 36.6% end-to-end success. Reducing model-directed steps is a reliability lever.
Define completion through a concrete artifact and measure success per customer transaction, including the attempts permitted by its retry budget.
In the insurance example, separate authentication, a packaged download operation, OCR-based verification, and a workflow skill narrow the model’s responsibility to the remaining ambiguity.
Production costs include browser infrastructure, integrations, observability, developer repairs, and reevaluation as websites and models change.
From a supervised demo to an unattended browser
A browser-agent demo gets favorable conditions: one site, one run, and a person watching. Production asks for thousands of unattended runs while websites change and models sometimes choose poorly. Derek Meegan, a software engineer at Browserbase, starts with that gap. The engineering problem is making a task complete repeatedly when nobody is there to rescue the browser.
The basic action loop separates choosing an action from executing it. Input tokens describe the page, the goal, and the steps already taken. The model proposes an action from a distribution of possibilities. A structured tool call then executes that action deterministically. Reliable execution of a click does not establish that clicking was the right decision; the uncertainty enters before the tool runs.
A browser offers several ways to observe and affect its state: network requests, the document object model and HTML, screenshots, the accessibility tree, and storage, cookies, URLs, tabs, and console information. Beneath these representations, a runtime can execute JavaScript or interact through CDP. Which parts an agent receives determines what it can reason about and how it can act.
Three strategies connect those browser capabilities to the model:
- Text representation. Combine HTML and the accessibility tree into a textual description of the page. The accessibility tree supplies semantic information about the interface.
- Computer use. Represent the browser through successive screenshots, letting the model work from its visual appearance.
- Dynamic code execution. Remove tailor-made tools and let the agent write code against the browser in an execution environment.
Meegan describes production systems as commonly using text, screenshots, or a combination. Dynamic execution offers a more open-ended interface, moving more of the interaction logic into code the agent generates.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Every step costs something; the completed task earns the value
Browser workflows span a spectrum. A purpose-built automation knows its task in advance. At the other extreme, a just-in-time agent receives an arbitrary task from a user. Between them, the browser can serve a larger objective such as deep research or competitor analysis. This talk narrows to transactional workflows: completing a unit of work end to end, with enough flexibility to handle ambiguity along the way.
The difficult economic relationship is temporal. Each step incurs cost and adds another opportunity for failure, while the useful outcome arrives only after the whole task completes. Progress through a transaction can consume resources without delivering the requested result. A highly capable model improves individual decisions, but a long sequence still gives small risks many opportunities to accumulate.
How much does length matter? In Meegan’s thought experiment, every step succeeds independently with probability 99%. A 100-step workflow therefore succeeds with probability (0.99^{100} \approx 36.6%), roughly the 36% stated in the talk. The independence assumption makes this an illustration rather than a measured production success rate. Its useful lesson is that strong per-step accuracy can coexist with weak end-to-end reliability.
That calculation creates a reason to change the system, rather than abandon browser agents. The model’s accuracy is one lever. The number of decisions it must make, the way completion is verified, and whether a failed attempt can be retried also shape the outcome.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Define success with an artifact, then measure the transaction
The business test is deliberately simple: (\text{profit} = \text{revenue} - \text{costs}). An agent is another line item and should return more value than it consumes. In Browserbase’s customer context, that value often comes from completing work on another customer’s behalf. The product being purchased is repeated task completion.
Meegan orders the engineering priorities as performance, cost, and maintainability. Here, performance means completing the requested task. Establish that capability first; then reducing its cost and keeping it working become engineering optimization problems. Cheap attempts have little value if they consistently fail to produce the result.
Completion needs an observable artifact:
- Bill payment. A confirmation email can mark the completed payment.
- Order placement. An order ID or receipt can identify the result.
- Form submission. A new record can be queried deterministically in the destination system.
These artifacts give the system something concrete to inspect after the browser interaction. They turn success into a check on the outcome rather than a judgment about how plausible the agent’s actions looked.
A run and a transaction also need separate metrics. One run is one attempt; a transaction is the customer’s unit of work, which may allow several attempts. With a 50% success probability per independent attempt, four total attempts yield (1-(1-0.5)^4 = 93.75%) transaction success. The talk calls this roughly 94% while describing up to four retries; the arithmetic corresponds to four attempts in total. The distinction matters when setting a retry budget.
The customer primarily wants the workflow completed reliably. Measuring only per-run success hides the benefit of retries; measuring transaction success captures it. Retries still consume resources, so their benefit belongs beside the accumulated cost of the attempts needed to finish the work.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Budget for browsers, repairs, and a changing web
Operating cost extends beyond model calls. Per-run expenses include model usage, the infrastructure and compute needed to run browser sessions, and additional integrations and tools. Meegan expects downward pressure on model prices to make intelligence a smaller part of the total over time. That is his forecast; browser resources and integration work remain part of the economic calculation regardless.
Maintenance has its own recurring work:
- Observability. Capture both the agent’s decisions and what actually happened in the browser. A decision trace alone cannot explain every execution failure.
- Developer time. Repair broken behavior, reconfigure the harness, or add capabilities when the existing system stops working well.
- Reevaluation. Revisit the solution as sites change and models improve. A successful configuration at one point in time still needs ongoing evaluation.
Several kinds of change can undermine the system. Anti-bot measures can make the environment resist automation. The underlying task can shift, requiring the agent to interpret intent and recover from unfamiliar situations. A new API can offer a better implementation route than browser interaction. Even without an external change, a model can stray from the critical path between runs.
Model flexibility helps when the job contains ambiguity: better understanding of the user’s intention can support course correction. The same freedom also permits unnecessary detours. The architecture that follows tries to keep that flexibility where it contributes to the task while removing decisions whose answers are already known.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Turn an insurance-document download into a verified workflow
The concrete example is a health insurance portal. A request asks the agent to log in and download an explanation of benefits. The smallest starting system sends the request to an agent, which performs the browser actions. It can do the job, but the model initially carries responsibility for the whole sequence. Each architectural change should improve performance, cost, or maintainability.
The first change packages downloading into one tool. Obtaining the document requires interacting programmatically with the page, retrieving the downloaded file from storage, and bringing it into the agent’s runtime. A dedicated tool performs that sequence as one operation. The browser and storage work still happens, but the model no longer has to decide how to coordinate every part of it on each run.
The next change makes the result inspectable. A verification tool uses OCR to extract entities from the downloaded document and compares them with a system of record. The observable outcome becomes a document plus a comparison result, allowing the agent to distinguish a completed task from merely having downloaded a file. Meegan presents this as deterministic verification; its conclusion depends on the extraction and comparison correctly capturing the document requirements, whose fields and error behavior are not specified here.
Authentication then moves out of the model’s responsibilities. In this example, the login workflow is treated as stable even though the business operation can contain ambiguity. A separate authentication function can log in once through a defined procedure and be reused across several operations on the same portal. Removing those model-directed steps aims to reduce unit cost and failure opportunities, while placing login maintenance in one reusable function.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Give the remaining agent a critical path
After the predictable operations have moved into functions, the agent still needs guidance through the portal’s ambiguous business workflow. A skill supplies a standard operating procedure for that remaining work. If the agent knows it is at step two and the procedure identifies step three, the intended next action becomes more likely. This connects back to the opening probability model: instructions guide action selection, while deterministic tools execute the operations that have already been defined.
Where does the model still make decisions, and where does the document travel? The diagram follows the final explanation-of-benefits workflow. Authentication supplies an already authenticated browser to the runtime. The skill guides portal navigation, the download function delivers the file, and the verification tool compares extracted entities with the system of record.
The resulting system is more complex than the initial request-to-agent demo, but the model’s job is smaller. Login has a reusable implementation. Downloading has a defined operation. Verification supplies an outcome check. The agent handles navigation with a skill and invokes those capabilities. Fewer decisions remain exposed to model variation, while the system retains flexibility for the part of the portal workflow that needs it.
That is the production tradeoff the ending leaves intact: reducing the agent’s responsibilities requires owning the surrounding software. Tools, authentication, observability, and reevaluation become ongoing engineering work. The last joke lands because the architecture has made reliable operation a concrete responsibility. Welcome to production; now someone is on call.
Log in and obtain an explanation of benefits.
Defined functions handle authentication and file retrieval. The skill guides the agent’s navigation, and document verification checks the result against a system of record.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
Related talks
- Beating RL With Reflection: GEPA and Optimize Anything
Extends the discussion of skills and agent improvement by explaining how execution feedback can optimize prompts, programs, and reusable repository knowledge.
Read the complete timestamped transcript
- 0:12
All right.
- 0:16
Deploying browser agents. Uh, running a browser agent demo once is easy. Uh, it's on one site, it's one run, on a good day, with you watching and your fingers crossed. Uh, but we know production is vastly different. Thousands of runs, unattended, while the site can change or your model can just have an off day. So that's why we're here. What separates a demo from
- 0:46
production? My name is Derek Meegan. I'm a software engineer at Browserbase, and today we'll break up this talk into four parts. First, we'll bring order to chaos, and we'll understand concretely how agents interact with the web. Second, we'll define the impossible and explore why running them at scale is so hard. Third, we'll take a step back. We'll think about what's important and the value proposition that agents propose. And
- 1:16
then last, which is hopefully the reason why we're all here, we'll generate some shareholder value. So first, how agents interact with the web. I like to think about the output of a model as a probability distribution. You receive input tokens. They represent the page state, the overarching goal, or the steps that have already been taken. Then the model explores a distribution of possible actions, proposing the most likely one. Then that
- 1:46
action is transposed into a structured tool call and deterministically executed to interact with the browser.
- 1:54
But the interface of the browser is complicated. You have several distinct layers. You have requests flowing in and out of the browser. You have the document object model or the HTML representation of the page. You can screenshot the browser. The accessibility tree, which is a semantic textual representation of the page. And the browser state: storage, cookies, console URL, tabs, et cetera. Underneath all of this, you have a dynamic code
- 2:24
execution runtime where you can interact with the browser by writing arbitrary JavaScript or CDP.
- 2:32
How do you wire these together? Typically, in industry, there's three main strategies. Number one is creating a textual representation of the browser. This typically involves taking the HTML state of the page and the accessibility tree and creating a hybrid representation. Number two is computer use, where you represent the browser in a series of screenshots of the page. And finally, and becoming increasingly popular, is de-harnessing your agent, removing tailor-made tools in favor of
- 3:02
dynamic execution environments where the agent can write arbitrary code against the browser. In practice, most production systems lean on the first, the second, or a combination of the first and the second two strategies.
- 3:18
Now let's think about the type of agentic trajectories on the browser. And here I define three, spanning across a spectrum of agenticness. From least agentic, we have purpose-built browser trajectories that are specially made to complete a single task. To the most agentic, we have just-in-time browser automations, where we don't know what the task is beforehand, and the user provides us an arbitrary task to complete. In the middle, we have
- 3:48
something sort of between both. We're using the browser as an implementation detail to serve a larger goal, such as performing deep research or competitor analysis. At scale, I really want to hone in on these transactional workflows. Not because it makes things easy, but that's because we typically see at scale. These workflows are transactional by nature, where we want to complete a unit of work end to end. On the other side, browser agents provide us a flexible
- 4:17
execution environment to handle ambiguity in production. So we know how agents interact with the web and the different tripe-- types of trajectories we may see, and we're honing in specifically on transactional workflows. Why is it so hard in production? And I think it comes down to a key interaction. Cost accumulation is continuous, while value realization is terminal. Each step in your
- 4:47
trajectory, you're incurring some incremental cost. Additionally, you're incurring some incremental risk of failure. Yet I only receive value from this browser trajectory when all steps of the task have been completed. That's all to say that with browser agents, there is no partial credit, and the vast majority of web automations fall into this category. So you say, "Okay, well, each step, there's more risk and there's more cost. If I use a highly capable model,
- 5:17
I can sidestep this conundrum." Take this thought experiment for an example. You have a browser agent where each step has an independent success rate of ninety-nine percent. If your browser trajectory spans one hundred steps, the overall success rate at scale starts to look like thirty-six percent, only about a third of the time. That's not super great, and it's certainly not how we deliver reliable outcomes for customers in production. And so
- 5:47
you say, "Okay, well, we're continuously accumulating costs, we're tim-- continuously accumulating risk, and the percentage chance of value realization is very low." That would suggest to never use a browser agent. Uh, thankfully, that's not the subject of my talk today. I think this dynamic is real, but the math isn't the whole story. And so, it's time to take a step back. Let's think about what's important and how agents actually
- 6:17
deliver value. And to illustrate the value that I wish agents to deliver, uh, I'll propose a simple equation. Um, perhaps not used enough by many of us. And this equation is profit equals revenue minus costs. Yes, very simple. Um, but perhaps overlooked when thinking about browser agents. And this is all to say that your agent is just another line item, and that your agent should give you more than it takes.
- 6:47
But when we think about what your agent should give you, that begs the question of: what does a browser agent actually give you? And so, to answer this question, I like to think about the context of the customer. At Browserbase, we help our customers deploy browser automations for their users. Oftentimes, they're not automating the web for their own needs, but instead performing actions on the web on their customer's behalf. And so, a browser agent actually gives you a task completed for
- 7:17
your customer repeatedly. Well, uh, it gives you that ideally. And the question is, how you can create a dynamic system that affords the benefit of being able to wade through ambiguous problems, while also completing the task repeatedly every time when the environment stays the same? And so, we'll look across three dimensions, and in order of importance. First is performance, second is cost, and third is maintainability. And when we
- 7:47
think about performance, this is the primary measure of success that we want to focus on. Because once we know a browser agent is performance, it can complete the task we asked it to do, then cost and maintainability simply become optimization problems. And I like optimization problems because those can be addressed with engineering. So, the first thing we need to ask is what success actually looks like. And success in these transactional browser trajectories
- 8:17
look like some kind of concrete artifact that the run leaves behind. If you're paying a bill, that might look like a confirmation email. If you're placing an order, that mo- may look like an order ID or a receipt, or if you're submitting a form, that may look like a new record in a system that you can, uh, deterministically query. Second, how often is an agent correct? And there's two ways to measure that. Naively,
- 8:47
you may measure the success of an agent on a per run basis. Assume you have a browser agent. Its probability of success on any run is fifty percent. Yet, if you permit that agent to retry for a particular transaction, maybe up to four retries, your per transaction success rate quickly climbs to ninety-four percent. And so when you think about what metric matters, again, I'd implore that we ground ourselves in the customer. The customer does not
- 9:17
care necessarily how many retries you perform, but rather that the workflow was completed successfully and reliably. In short, permit retries and measure success against a per transaction basis.
- 9:34
Next, we'll explore cost, and we'll break cost down into two dimensions. The first is per run, which today is predominantly model cost. But as engineering evolves in the open source landscape for models, there is an incredible downward price pressure on the cost of intelligence. And I believe over time, model cost will become a smaller and smaller component of deploying browser agent trajectories. Next is infrastructure and compute. What are the real
- 10:03
resources you need to run your browser agent trajectories? And last is additional integration and tooling. Then we think about maintenance over time. We need to invest in observability, not only in how the agent made decisions, but what actually transpired in the browser. Second is developer time. When things break or things are not working optimally, people need to come in and reconfigure the harness or augment its capabilities. And last is
- 10:32
reevaluation as the problem changes. Sites change, models progress, and we need to reevaluate our point in time solution to last the test of time.
- 10:46
So, we talk about performance, we talked about unit costs, and we also talked about maintainability. What are the risk factors to the durability of these over time? Well, first and foremost, the environment can fight back. Anti-bots, and just generally, we understand the web as not the friendliest place for your agent. Second, the task changes. The underlying job shifts beneath the ag- shifts beneath the agent, and we ask the agent to wade through
- 11:16
ambiguity. Luckily, this also benefits from increase in model capability, as models are more able to intuitively understand the intention of the user, and therefore, in the face of something go wrong, course correct and ultimately complete the goal. Third is a better method appe- better methods appear. And this is sort of a canonical conundrum with automating the web. What happens if I'm automating the web today and that same website creates a API tomorrow? But for the vast
- 11:45
majority of browser agent use cases, uh, an API isn't coming anytime soon. And finally, the model strays off task. Models are inherently indeterministic. So from run to run, it may not id- follow the ideal path or the critical path to perform the action. So,
- 12:06
finally, we've come to the meat of the talk. How do we concretely generate shareholder value? This is the question. Uh, we understand that if a model performs well, then cost and maintainability are an optimization problem. So let's explore that in practice. Together, we're gonna go through a high-level architecture of a browser agent for real-world tasks. This task is automating a health insurance portal. Every step we take in creating this dynamic system should influence
- 12:36
one of these key levers: cost, performance, or maintainability. So first, we'll start with a lowly agent. The goal of this agent is to log in and download an explanation of benefits. A request comes in, the agent performs the action on the browser, and this is the minimum necessary system you need to just do something. Next, you realize that when you go to download that explanation of benefits, that downloading involves
- 13:06
programmatically interacting with the page. And then once you perform that download, you need to retrieve that file from storage and bring it into the agent's runtime. So you create a tool that does this all in one transaction. Instead of the model needing to make decisions on the fly, you encapsulate a complex operation into a single tool call. Next, you wanna give the agent an opportunity to verify its work real time. This can be done by giving it another tool.
- 13:36
This tool is an OCR tool that deterministically extracts entities from the document and compares that against a system of record. Great. Now the agent knows if it's did it-- if it really did its job right or not because it has a deterministic mechanism to verify that. Next, you realize that while this portal may have ambiguity in your business logic, that the authentication workflow is never changing, and you want to reuse the same
- 14:05
authentication mechanism for many business operations on the same portal. So you pull that authentication logic out from the responsibility of the model. Again, we're reducing the number of steps that the agent has to take to complete the task successfully. This means lower unit cogs, higher performance, and easier maintainability. As well as we can reuse this authentication function across multiple business logic workflows on the same portal.
- 14:35
Last, we've contained the responsibility of the agent to only the area of ambiguity necessary, so we draft a skill. These skills look very similar to a standard operating procedure that we would have humans use in the past, except this standard operating procedure is for an agent, and it allows the agent to follow the critical path more closely. At the beginning of the talk, we discussed how the output of an
- 15:05
agent can be thought of like a probability distribution. So if we introduce the critical path to the agent, it knows it's at step two and knows that it needs to take step three, because we have outlined it in this skill, then it's more likely to take that step three. It knows exactly where it needs to go, and we remove ambiguity from the agent's decision process. In the end, this is the final result. We have actually a pretty complex system. A request comes in, a
- 15:34
serverless authentication function authenticates the browser. That browser is provided to the agent runtime. The agent navigates the portal using the skill provided, calls a deterministic download function, provides that downloaded file to a verify tool, and can know authoritatively whether its trajectory succeeded or failed. In the end, this system takes steps out of the model's responsibility that simply do not need to
- 16:04
be in the responsibility of the agent. We reduce the number of steps it needs to take while still being it-- able to achieve the adap-- desired outcome. Thus, the system's more maintainable and it's more performant, and we can overcome the browser-agent conundrum. So here we are. Welcome to production, and congrats. Now you're on call. Thank you.