Felipe Blanes is a technical program manager whose work has connected railway-control engineering, Alexa quality assurance, and customer adoption of Amazon Nova Act, a system for automating browser tasks. His approach to AI reliability centers on whether automation succeeds in the work customers actually need to accomplish—and whether they can trust it enough to use it.
From railway controls to Alexa
Blanes studied control and automation engineering and moved from an internship at GE into railway-control work. His projects included signaling software for Waterloo’s light-rail system in Canada, the expansion of Brazil’s Carajás railway, and São Paulo’s Line 15 monorail. These systems required software to coordinate physical equipment and railway operations. He also co-founded Ride’O Caronas Inteligentes, a corporate carpooling startup, in 2017. His professional history spans these early ventures into transportation and automation.
He subsequently joined Amazon’s QA team supporting Alexa’s introduction in Brazil, bringing his automation background into testing a conversational product for local users. He later worked remotely with a Boston-based team before relocating to the area. That career progression took him from industrial controls to consumer software and technical program management.
Making browser automation useful to customers
Blanes worked directly with Nova Act customers as the service moved from research preview to general availability on AWS. Nova Act combines a specialized foundation model, orchestration, and browser controls to carry out browser tasks through natural-language instructions. The work extends automation into interfaces designed for people, making reliability in actual customer workflows a central concern.
His team’s work included browser-agent quality assurance and tools for putting automation into customers’ hands. The team released Nova Act Quick Deploy Studio, which lets customers author, run, and schedule browser automations in their own AWS accounts. Blanes’s release announcement credited his colleagues: this was a team contribution connecting the agent’s capabilities with the practical work of deploying and operating it.
Blanes describes the benchmark illusion: strong results on static evaluations can obscure failures that appear when customers try unexpected tasks. His response is an evaluation process that keeps learning from customer use, tying improvements to the workflows people are actually trying to automate.
Define success through the customer’s task. His evaluation flywheel starts with what the customer considers a successful outcome, then gathers signals—including direct conversations—to understand where automation falls short. Diagnosing whether a failure belongs to the model, the harness, or the product helps turn those signals into decisions about what to improve. Evaluation becomes a continuing development process rather than a score collected before launch.
Treat trust as a threshold. Blanes describes a “trust cliff”: at 80% reliability, customers can experience an agent as more work, while around 92% it can begin to feel trustworthy. He uses this contrast to explain why an apparently strong success rate may still be insufficient for adoption; those figures describe his customer-facing lesson rather than a universal threshold for every agent or task.
Make limitations visible. He argues that openness about what an agent can and cannot do builds trust more effectively than higher benchmark scores alone. Customers need to understand where they can rely on automation, and their experience should keep changing the evaluations used to guide its development.
These ideas extend the practical concern running through Blanes’s earlier work: software quality depends on how a system behaves within the operations people rely on. For browser agents, he makes customer feedback part of the mechanism for improving that behavior.
Felipe Blanes explains how Nova Act’s customer feedback loop turns production failures into better evaluations, engineering fixes and product decisions—and why trust depends on more than a benchmark score.
Static benchmarks leave a gap when customers attempt unfamiliar tasks. Production signals need to feed back into evaluation and development.
The eval flywheel defines customer success, captures telemetry and conversations, diagnoses model, harness and product gaps, then turns them into prioritized decisions.
Reliability matters through the work it removes: an agent that still needs monitoring and manual repair may increase the customer’s burden. Blanes’s 80% and around 92% examples describe an observed trust cliff, not a universal delegation threshold.