AI Engineer World's Fair 2026

Why 80% Reliability Isn't Good Enough — Felipe Blanes, Amazon AGI Lab

Read the talk

Why 80% Reliability Isn't Good Enough

Felipe Blanes explains how Nova Act’s customer feedback loop turns production failures into better evaluations, engineering fixes and product decisions—and why trust depends on more than a benchmark score.

From a talk by Felipe Blanes

At a glance

Ideas worth remembering

  • Static benchmarks leave a gap when customers attempt unfamiliar tasks. Production signals need to feed back into evaluation and development.

  • The eval flywheel defines customer success, captures telemetry and conversations, diagnoses model, harness and product gaps, then turns them into prioritized decisions.

  • Reliability matters through the work it removes: an agent that still needs monitoring and manual repair may increase the customer’s burden. Blanes’s 80% and around 92% examples describe an observed trust cliff, not a universal delegation threshold.

  • Explain what works, what is being improved and what is out of scope so customers can choose uses that warrant confidence.

  • Evaluations must evolve with customers: early exploration, recurring use cases and increasingly complex work each introduce new questions to test.

A browser-agent product built to learn from customers

A browser agent can perform well in testing and still disappoint the people who need it to do real work. Felipe Blanes, a member of technical staff at Amazon AGI Lab, approaches that problem through his work helping customers use Nova Act, the lab’s tool for building agents that automate browser tasks.

Nova Act’s research preview launched in March of the preceding year with a deliberate goal: get the product into customers’ hands and collect feedback. The first few months exposed gaps in the developer experience, leading to a new experience in July. The service then launched on AWS in December. Blanes worked with customers throughout that progression, connecting what they were trying to accomplish to what the product actually helped them do.

Selected presentation frame from Why 80% Reliability Isn't Good Enough — Felipe Blanes, Amazon AGI Lab at 91 secondsOpen full source frame
A timeline marks the March research preview, July developer experience improvement, and December AWS service launch.
0:120:42
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:12 · section reference included

The benchmark illusion leaves production to hope

The “benchmark illusion” starts with a familiar development sequence. A team receives requirements, builds a product, runs public benchmarks or evaluations against its own synthetic data, and improves the results until it feels ready to ship. Those tests can tell the team how well the product handles the scenarios it has chosen. Customers then introduce scenarios the team never anticipated.

Selected presentation frame from Why 80% Reliability Isn't Good Enough — Felipe Blanes, Amazon AGI Lab at 182 secondsOpen full source frame
The slide places an “Evals” example beside a “Production” example.

The failure is in keeping the evaluations static while actual use changes. After shipping, a fixed test suite leaves the team hoping that its examples represent customer behavior. Unexpected tasks break the product without necessarily changing what the team measures. A higher score on the same tests cannot, by itself, close that gap. Production experience needs a route back into evaluation and development.

2:132:43
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

2:13 · section reference included

Turn customer signals into decisions

The eval flywheel begins by defining success in the customer’s terms. Before choosing a metric, understand the problem the customer needs solved. That definition determines which signals matter: collecting large amounts of telemetry is less useful if the measurements miss the intended outcome.

Instrumentation supplies one kind of signal; conversations supply another. Meetings and visits help uncover what customers are trying to do and why the observed behavior matters. Blanes’s instruction is direct: “talk to your customer.” Data can be misleading when the team lacks that context. The loop therefore combines measured behavior with insights from working directly with users.

Diagnosis then separates three kinds of work:

  • Model failures: Identify where the model fails and feed those cases into model evaluations.
  • Harness failures: Find problems in the surrounding agent software and send them to engineering for conventional bug fixing.
  • Product gaps: Ask whether the product solves the intended problem, whether customers understand its purpose, and whether its positioning sets appropriate expectations.

This separation matters because a customer-visible failure does not automatically imply that the model needs improvement. It may require a software fix or a product decision. Once the gaps are categorized, the team prioritizes them and assigns work to science, engineering or product. The flywheel continues by repeating the process rather than treating the diagnosis as a one-time report.

Selected presentation frame from Why 80% Reliability Isn't Good Enough — Felipe Blanes, Amazon AGI Lab at 425 secondsOpen full source frame
A four-step flywheel labels Define Success, Capture Signals, Diagnose Gaps, and Feed Into Decisions.

How does customer experience reach the team that can change it? The cycle below shows the intermediate diagnosis step. Signals become useful when they lead to a specific kind of decision; repeating the cycle keeps the definition of success connected to what customers need.

How it fits togetherThe customer-driven eval flywheel

Understand the outcome the customer needs.

Customer signals pass through diagnosis before they become model, engineering or product work.

4:395:09
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

4:39 · section reference included

The questions change as customers adopt the product

Customers have their own progression through a product. Repeating the flywheel reveals different information at different stages, so the team should change what it is trying to learn as adoption grows.

  • Use-case discovery: Establish the customer’s problem and whether the product can solve it.
  • Early adoption: Explore what is possible. Customers try unexpected tasks, revealing both unsupported uses and surprising successes.
  • Scaling: Look for recurring patterns across more customers. These expose the “hero use case”—the problem that repeatedly brings people to the product.
  • Future product bets: Use gaps discovered during early adoption and scaling to decide what to build next.
Selected presentation frame from Why 80% Reliability Isn't Good Enough — Felipe Blanes, Amazon AGI Lab at 546 secondsOpen full source frame
A horizontal journey diagram shows Use Case Discovery, Early Adopters, Pattern Identification, and Product Bets.
7:137:43
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

7:13 · section reference included

Reliability changes whether customers can hand over work

An agent with 80% reliability can look ready to ship from the builder’s perspective. From the customer’s perspective, it can create another job. Someone must build the automation, monitor its execution and finish or repair the work when it fails. The agent’s successful attempts do not remove the need to supervise the workflow.

Blanes calls the change in willingness to delegate the “trust cliff.” In his customer experience, around 92% reliability changed the response from needing to watch the agent to feeling able to hand over tasks. These percentages illustrate his observed adoption pattern, rather than a universal threshold: the talk does not define a shared reliability metric or establish which workflows can safely be delegated at that level. The useful mechanism is the shift in supervision burden. Trust becomes a practical decision about whether the automation actually removes work.

Selected presentation frame from Why 80% Reliability Isn't Good Enough — Felipe Blanes, Amazon AGI Lab at 638 secondsOpen full source frame
The slide reads “80% reliability = 0% trust” and describes the extra checking and repair work when an agent works less than about 90% of the time.

That decision also needs honest information about the product’s limits. Customer discussions should distinguish three categories:

  • What works: Explain the tasks the product handles well and the reliability customers can expect for them.
  • What is being improved: Identify known gaps that the team is actively working on.
  • What is out of scope: Say plainly which tasks the product cannot handle, before customers spend time discovering that themselves.

Clear limits help customers choose an appropriate use and avoid predictable disappointment. With many competing tools available, an unsuccessful trial can be enough for a customer to move on. Blanes’s qualitative lesson is that transparency about limitations can build more trust than higher benchmark scores: it lets the customer understand where confidence is warranted.

9:3710:07
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

9:37 · section reference included

Customer work reveals changes a benchmark would miss

Amazon Leo, Amazon’s satellite internet business, supplied a concrete example of customer feedback changing the execution path. The resulting feature caches trajectories—the sequences of actions used to carry out a task—rather than requiring model inference every time. If the cached trajectory fails, execution falls back to the model.

The change moves repeated work onto a cheaper path while keeping model inference available when reuse does not work. Before caching, inference was needed on each use; afterward, a successful cached trajectory could avoid it. Blanes reports cost optimization for this customer, without quantifying the savings or specifying how trajectory failure is detected.

Selected presentation frame from Why 80% Reliability Isn't Good Enough — Felipe Blanes, Amazon AGI Lab at 790 secondsOpen full source frame
An “Amazon Leo” card describes caching trajectories and using a model fallback when something goes wrong.

Where does inference remain in the cached workflow? The diagram shows that it becomes the fallback after a failed cached attempt. The cost benefit comes from successful reuse; the fallback preserves a route forward when reuse fails.

Two other customers exposed different kinds of gaps:

  • Hertz: access for less technical QA staff. Some engineers using Nova Act for QA automation could write Python and other scripts. Others could not. That difference prompted a non-technical product experience intended to let more customers use the tool. The obstacle was how people built the automation, rather than simply how well the model performed it.
  • Sola: flexibility across varied automation tasks. The process-automation startup served many different customer use cases. Its needs pointed toward more customization, particularly in the actuation stack—the part of the tool that carries out actions. The talk identifies the direction of that flexibility without specifying the customization interface.

These examples connect customer contact to three distinct decisions: reuse work to reduce inference costs, broaden who can build an automation, and let an automation provider customize execution. None requires interpreting every customer problem as a request for a better model.

How it fits togetherReuse a trajectory, then fall back when it fails

Use Nova Act for the customer’s browser task.

Successful cached execution avoids repeated inference; failed cached execution returns to the model.

12:4313:13
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

12:43 · section reference included

Keep evaluations learning as customer work changes

Repeated use of the flywheel led to four operating principles:

  • Derive scenarios from production: Synthetic data can help at the beginning, but customer signals should feed the evaluations once real usage exists.
  • Categorize for action: Separate research, engineering and product issues so each can be prioritized and addressed by the appropriate team.
  • Share capabilities honestly: Include limitations in customer communication, especially tasks the product does poorly.
  • Run the loop quickly: Shorten the time between learning from customers and acting on what they reveal.
Selected presentation frame from Why 80% Reliability Isn't Good Enough — Felipe Blanes, Amazon AGI Lab at 941 secondsOpen full source frame
A numbered slide lists four principles: derive scenarios from production, categorize failures, share capabilities honestly, and run the flywheel fast.

The expectation is that evaluations should get “smarter every week.” Here, smarter means learning from the problems customers are actually encountering, rather than continuing to test only the team’s original assumptions. If that learning stops, evaluations become stale even while the product keeps running.

The ending ties freshness to the customer journey. As customers move from exploration into broader adoption, what they attempt changes and can become more complex. Evaluations must keep reflecting that work. The two lasting requirements are therefore connected: measure what customers care about, and keep updating those measurements as their use evolves. A suite built once cannot keep answering a question that keeps changing.

14:4315:13
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

14:43 · section reference included

Resources

Read the complete timestamped transcript
  1. 0:12

    Can everybody hear me? Okay, perfect. So first, thanks for joining, everyone. Uh, my name is Felipe, and I am a member of technical staff at the Amazon AGI Lab. Um, as part of my role, I help us to build the next generation... I manage the programs to build the next generation of products our lab is building. And today, I will talk to you about designing evals that earn customer trust. Before talking about

  2. 0:42

    evals, I just need to give you some context about what our lab was doing the past year. So everything started in March last year, where we launch our Nova Act research preview. For context, like Nova Act is a tool that we built as part of our lab, and it's a tool that helps you to build browser agents. So basically, everything that you do in a web browser, you can use Nova Act to automate those tasks. Uh, and as part of the research preview,

  3. 1:12

    our main goal was we wanted to ship a product to get as much feedback as we could from our customers. Uh, in March, we launched that. Then in July, we made one additional improvement in our product, which was basically develop a whole new developer experience. And this was because basically during the first couple months, we noticed, like, some gaps in the developer experiment-- uh, experience, and we wanted to make sure customer had the best experience building agents.

  4. 1:43

    Then in December, we finally launch our service on AWS. So today, Nova Act is a top-tier service on AWS that a- anyone can use to build, uh, production-level agents. And during this period, I was basically responsible for the customer enablement piece of that. What's customer enablement, just for context? It's basically I was working day to day with all of our customers to understand what was their problem and how we were helping them to solve that.

  5. 2:13

    So that was my main, uh, role, uh, during last year. And then let's go to, uh, the problem that we will be discussing today. So I like to call this is the benchmark illusion. So you all working on AI, you are probably working on and building things based on benchmarks, right? It could be public benchmarks that everybody's using and everybody's trying to optimize the products on that, and also benchmarks

  6. 2:43

    that you are creating, your own synthetic data that you are creating and you are evaluating your product. Uh, and the main goal here is you want to optimize, right? You want to be the best as you can there in those benchmarks. But then when you ship that to production and your customers start using that, what happen is, it's this, right? They start seeing problems, and it's basically everything that you were not expecting them to do, they try to do that, and they break your product.

  7. 3:13

    And this is happening because you basically-- most of us, we are doing-- we are working on evals that are, that are static. So what are static evals? So basically, let's talk a little bit about what's the standard development process of a product, right? First thing that you do, you receive, uh, the requirements from your product team, and you start building, right? Everybody does that, I hope. Then when you have a product in a good shape, the

  8. 3:43

    next thing that you will do is running your evals. Again, public benchmarks, synthetic data that you generated, and you optimize for that. When you start feeling comfortable, you're like, "Okay, now I will ship that," and you ship that to the customers. But if you have static evals, what's the only thing that you can do after that? It's basically this. It's hope, right? You just hope your evals were reflecting exactly what your customers will do with your

  9. 4:12

    product, and if not, it will start failing, right? And this is the gap. This is exactly what we will be discussing today and what we were doing in our lab to solve this problem. And just a spoiler, so this is not about better benchmarks. It's about closing the customer loop. So what we want you to do here is get the feedback from your customer and close the loop.

  10. 4:39

    So with that in mind, what we did in our lab was basically propose an eval flywheel. So if you are not aware about the terminology, but flywheel, it's basically a loop that you keep repeating, like, all the time. And the first step of this eval is define success. And very important here, so define success here, it's not what you think it's success. It's what your customer think it's success, okay? So

  11. 5:09

    next step, as soon as you understand what your customer really needs and the problem you are trying to solve, you go to capture signals. And we as engineers, what's the first thing that comes to your mind? It's like instrumentation, right? Let's get metrics. Let's get data, as much as you can, about what your customer is doing. But one very important piece of that, that we are missing, is you should talk to your customer. Like, talk to your customer. Go, schedule

  12. 5:39

    meetings, uh, visit them, and understand exactly what they are trying to do. Your data might be a little bit misleading. You might be missing exactly what's the problem that your customer is trying to solve. Third step, you go to diagnose gaps. That's easy. That's basically you get all the signals from customers, could be instrumentation, but also your inputs, your insights from the meetings that you had with them, and you just start categorizing those. And when I say categorizing, there are

  13. 6:09

    basically very high level, but there are three things that you could find from those discussions. First, it's basically inputs about your model. So where your model is failing there, and this feeds into evals, right, to your model. The, the second point, it's inputs to your engineering team. So basically, what, uh, is the gap in your harness that you should fix? This is more traditional bug fixing. And the third one is also on the product side. So

  14. 6:39

    what are you missing in terms of product? Like, do your customer understand exactly what your product is trying to solve? Is your product, like, properly positioned? So that's very important. So as soon as you have the gaps di- uh, the gaps categorized, you just feed that into decisions. So feed that into decisions is basically you prioritize that, and you ask your science, your engineering, your product team to address those. That's basic, right? And then you just continue

  15. 7:09

    repeating this loop over and over.

  16. 7:13

    So with that in mind, like in our lab, we repeated this flywheel several, several times since March last year, and what I also noticed is basically, as part of my journey as a product, each of my custom-- each customer will have their own journey using my product. What that means, so my customer evolves as they use my product, right? So there are basically four things that you can learn while

  17. 7:43

    your customer use your product. The first one is, that's the first thing that you will do the first time that you meet your customer. It's basically use case discovery. So what you want to do here at this first step, it's basically understand what's their problem. So what's the problem of your customer, and make sure what you are building solves that, right? Then after that, you will have couple early adopters, right? From these early

  18. 8:13

    adopters, the main signal that you want to get, it's basically what's possible. So that's the time where customer will try crazy things. You will find things that you cannot help them at a- at all, right? But you will also find things that your customers are using your product to solve. And many times you will be surprised because it's something that you would never imagine.

  19. 8:37

    Third point, this is when you start scaling, right? You get your early adopters, but now you start scaling. You are getting more customers, and you are able to identify some patterns, right? What are the patterns? The pattern is basically your hero use case, right? That's the case, that's the problem that everybody is using your product to solve. And last one, this is more future-looking, but these are your product gaps. So when you were working

  20. 9:07

    with early adopters, and also after you start scaling, you will also be finding some gaps on your product. Your product team can help you to prioritize those, and these are the bets for your future, right? What you should build next. Um, as I said, I work with several, several customers during this journey, and there are a couple learnings that I want to share that some might be obvious, but others, it's very, very insightful. So the first one is

  21. 9:37

    the t- the trust cliff that I f- I like to call. So what's the trust cliff? So imagine you build a product, you are ready to ship, and you s- and you measure the reliability of this product, and you found eighty percent reliability. That seems fine, right? You're like, "Okay, I can give that to my customer and they, they can start using it," right? But what your customer thinks when you say you are eighty percent of

  22. 10:07

    reliability is, "I cannot trust this product. I cannot trust." Uh, and why? Because when you show nine-- eighty percent reliability, what they will think is, "Oh, okay, so if it's eighty percent, I still have to monitor this agent. If they fail, I will need to do some manual work. So in fact, it's actually more work for me, not less work. I need to build this agent, plus

  23. 10:37

    monitor that, and also do some manual work after." But when you get to a point around ninety-two percent reliability, the message changes to your customer. And what they think, it's, "Okay, ninety-two percent, I can actually trust this, right, to handle my tasks. And I completely hand over, like, workflow to my customer." So it's basically reliability is a zero to one, right? Or you trust or you don't. There is

  24. 11:07

    nothing like in the middle there. Second learning, it's the eval transparency. So every time that you start working with a customer, you basically need to be transparent about three things that your product does. First thing is, that's easy. Everybody likes to do that. So what works? So you basically show to your customer, "I am really good doing this. That's the level of reliability. You can trust my product.

  25. 11:37

    Just use my agent and it will give you the reliability that you need." Second point is, "These are the things that I'm doing to improve my product," okay? So there are some things that I know it's a gap, but I am investing time right now to make this better. And the third thing, that's the tough one. Nobody wants to do th- that. It's what is out of scope. So you really need to be transparent. Like, if something doesn't work for your product, just go ahead and tell your customer.

  26. 12:07

    Don't let them try. Otherwise, like, there are so many tools in the market right now that if you do some-- if you show something for them that doesn't work, they will just forget about your product, and they will move to the next. So be transparent. And the main learning here is transparency on product limitation increase trust more than higher benchmarks. That's a very interesting finding, right? Be transparent about what your product is capable of doing and, more important, what it's

  27. 12:37

    not capable of. That's how you earn customer trust.

  28. 12:43

    Third learning is if you work, work very close with customers, you might find a lot of use case that you would never expected. So some examples here is Amazon Leo. It's the satellite-- internet satellite company from Amazon. Working with them, we actually learned some inputs, uh, about creating features to do caching of trajectories. So instead of doing inference all the time that they use our product, we cache trajectories. If it fail, we just fall back to the model.

  29. 13:13

    So we optimize the cost for them. Second learning is with Hertz, the car rental company. So they were also using our product for QA automation. And we learned that as part of their company, they had some QA engineers that were very technical and were capable of building, like, scripts with Python and other coding languages. However, some technical, uh, some people from QA, they are not that technical. So that

  30. 13:43

    was the input that we needed to actually build a whole non-technical experience for our product that would enable more customers to use it. Last one, this is a startup called Sola. Sola is basically a RPA company. If you are not familiar with the term, RPA is repetitive process automation, and they offer tools to automate process for other companies. So imagine the variety of use cases that they

  31. 14:13

    might have. It's a lot of different use cases. And what we learned from them is we could give some additional flexibility in our product for them to customize, especially the actuation stack of our tool. The learnings here are less important than keeping in mind that you will learn a lot about what your customers is trying to do here, right? And what should be the next features you should build for your product. And then,

  32. 14:43

    like, running this flywheel several times, we came up with four principles for this flywheel. First one is derive evals, uh, eval scenarios from production, not imagination. So again, like, creating synthetic data might be very important at the beginning of the process, but as soon as you start cu- getting customer signals, make sure that feeds your eval. Second, categorize into actionable areas. Again,

  33. 15:13

    product issues, engineering issues, research issues. Make sure you prioritize that well to your team. Third one, share capabilities honestly. Again, transparency builds trust. Make sure you share to your customer especially what your product is not good at. Third one is run the flywheel as fast as you can. So here it's basically the faster you run that flywheel,

  34. 15:43

    faster your product will improve, right? That's almost obvious, right? But keeping that in mind, like, run your flywheel as fast as you can. And then, like, with those four principles, what you will see is, uh, your evals will g- should get smarter every week. But very important, if you are not seeing your evals getting smarter every week, that means your evals are getting stale, okay? And

  35. 16:13

    just to finalize the talk here, there are two ideas that I want you guys to take home. Just two. So first idea is eval must reflect what customer cares about, not what you think they care about, okay? All about getting customer input here and make sure your evals reflects what's their problem. Second is evals are not built once. It's a flywheel that evolves as your customer

  36. 16:44

    evolves. One note here, remember that customer journey? As your customer moves through that customer journey, what they are trying to do with your product will also change. Like, it will get more complex, right? So your evals needs to be up to date and reflecting exactly what your customer is trying to do right now. And with that, what you have to do is simply, like, build the loop, ship evals that matter, and then

  37. 17:13

    you earn your customer trust. And that's all that I have for today. Uh, thanks very much for joining. Uh, I will stay in our booth after for more discussions, but again, enjoy the conference.