AI Engineer Code 2025

Why AI Didn't Actually Make You Ship Faster — Gabriel Spencer-Harper, Meticulous

Read the talk

Why AI Didn't Actually Make You Ship Faster

Selected presentation frame from Why AI Didn't Actually Make You Ship Faster — Gabriel Spencer-Harper, Meticulous at 61 secondsOpen full source frame
Slide reads “Engineering the future of AI” beside the speaker on the Expo Stage.

Gabriel Spencer-Harper explains how Meticulous moves frontend verification from handwritten assertions to recorded workflows and visual comparisons—and why network mocking, browser determinism and coverage-guided selection have to work together.

From a talk by Gabriel Spencer-Harper

At a glance

Ideas worth remembering

  • Faster code generation can leave verification as the limiting step: the team still needs to understand a change across roles, permissions, flags and configurations.

  • Meticulous compares recorded workflows before and after a change. It exposes visible differences; a person or agent decides whether they are expected.

  • Recorded network responses make replay repeatable and isolate tests for parallel execution. Browser-level timing control addresses variation that network mocking cannot remove.

  • Coverage-guided selection depends on frequent screenshots to turn executed paths into visually checked paths. At large screenshot volumes, controlling flakiness is essential to keeping the differences useful.

Writing code faster leaves the merge decision behind

An agent finishes a pull request. Would you feel comfortable merging and shipping it? That is the practical question behind Gabriel Spencer-Harper’s talk. The co-founder and CEO of Meticulous starts with a mismatch: AI can produce code faster than people can establish that the change is safe. Someone still has to verify it, so faster code generation can move the bottleneck rather than remove it.

A frontend change can behave differently under different feature flags, roles, permissions, settings and configurations. Reviewing the patch does not automatically reveal its effect across those combinations. The work before merge includes reaching the relevant application states, observing what changed and deciding whether those changes are acceptable.

Assertion-based tests encode expected behavior in advance. A human or agent chooses what should happen and writes checks for it. Spencer-Harper’s objection is the size of the remaining space: even diligent test authors cannot anticipate every possible regression. A passing assertion answers the question someone thought to ask; the merge decision often depends on effects they did not anticipate.

Selected presentation frame from Why AI Didn't Actually Make You Ship Faster — Gabriel Spencer-Harper, Meticulous at 153 secondsOpen full source frame
The slide “ASSERTION-BASED TESTS ARE NOT ENOUGH” shows a Claude Code-authored pull-request summary and a green “Merge & Ship” button.

The costs show up in several places:

  • Regressions: shipping faster can expose users to bugs with business consequences.
  • Verification work: manual validation, review, flaky-test debugging and suite maintenance consume engineering time. Spencer-Harper describes this as double-digit percentages of an organization’s time.
  • Changes deferred: dependency upgrades and sweeping refactors become easier to attempt when engineers can inspect their consequences with confidence.

Meticulous is presented as a way to obtain exhaustive or near-exhaustive frontend verification without developers maintaining the tests. That ambition needs a precise reading: the mechanisms described here compare recorded frontend behavior and select flows by executed lines of code. They do not establish that every possible state has been recorded or that every kind of correctness is visible in a screenshot.

0:120:43
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:12 · section reference included

Replay the same workflow before and after the change

The demo begins with one line of JavaScript injected into non-production environments: localhost, QA, development or staging. It instruments the browser and records workflows as people use the application. Clicking a login button, opening settings and visiting analytics becomes a sequence of events that the system can replay. The recorded collection can contain thousands or tens of thousands of flows.

When a pull request opens, the CI runner starts the application on localhost port 3000. Meticulous selects recorded workflows and dispatches their events one by one. Throughout each replay, it captures screenshots at what Spencer-Harper calls each atomic moment. Running the workflow before and after the code change produces two screenshot sequences; comparing corresponding states reveals the application’s visible differences before merge.

Selected presentation frame from Why AI Didn't Actually Make You Ship Faster — Gabriel Spencer-Harper, Meticulous at 306 secondsOpen full source frame
A GitHub pull request displays code changes with added and removed lines.

The concrete example contains two changes: a name has changed, and an error has been introduced into a dropdown. The comparison exposes the name change in both light and dark modes, along with a visible manifestation of the dropdown’s logical error. The causal path is straightforward: replay reaches the relevant state, the changed code renders it differently, and the before/after comparison puts that difference in front of the reviewer.

A pull-request comment leads to the differences, typically within a few minutes in the workflow described. Meticulous does not label each difference a bug. A reviewer decides whether the name change is intended and whether the dropdown’s changed behavior is acceptable. Business judgment moves from writing an assertion before the change to interpreting an observed difference afterward; Spencer-Harper also allows for an agent to perform that review.

Where does the reviewer enter the process? The diagram separates event replay and comparison from the decision about correctness. Both versions receive the recorded workflow, but the resulting diff still needs someone to interpret it. In the demo, that distinction lets an expected name change and an unexpected dropdown change appear through the same mechanism.

How it fits togetherOne workflow, two versions, one review decision

Browser events captured in a non-production environment.

Replaying the same events produces comparable screenshot sequences. The diff identifies visible changes; review determines whether they are intended.

4:044:34
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

4:04 · section reference included

Control the network and the browser’s timing

Comparing screenshots only helps if repeated runs of unchanged code produce comparable images. Meticulous addresses two sources of variation at different layers:

  • Recorded network traffic: at record time, it captures network requests and responses. At replay time, it stubs them in, giving the workflow recorded responses rather than relying on changing live network results.
  • Browser scheduling: it augments the browser from the scheduling-engine layer upward to control randomness that network mocking alone cannot remove.
Selected presentation frame from Why AI Didn't Actually Make You Ship Faster — Gabriel Spencer-Harper, Meticulous at 459 secondsOpen full source frame
Slide titled “KEY TECHNOLOGIES” lists exhaustive coverage, deterministic testing and automatic mocking.

Network mocking has two jobs in the design. First, it makes replay repeatable: Spencer-Harper describes running a test a thousand times and obtaining the same result. Second, each test is isolated from the others. Workflows need not compete through shared network state, so the system can run them horizontally in parallel and return results within a few minutes. The same decision that supports repeatability also supports speed.

The browser itself can still introduce differences. An animation spinner may render at different positions depending on the machine’s speed. Different interleavings of setTimeout and setInterval can alter what happens before a screenshot. If those differences survive replay, unchanged code can look changed, producing flaky results and wasting review attention.

Meticulous therefore aims for a deterministic or near-deterministic browser, controlling this variation below the test script. Spencer-Harper claims radically fewer flakes than Cypress or Playwright; the recording supplies the scheduling rationale rather than a quantified comparison. The useful architectural point is that visual comparison depends on controlling when the browser does work, as well as what data it receives.

7:107:40
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

7:10 · section reference included

Choose workflows by the code they exercise

A large recording library creates another problem: which workflows should run on each pull request? Meticulous replays each recorded session once against the main or master branch and monitors the lines of code it executes. That produces an index from workflows to executed lines. It then selects a subset of workflows that maximizes code coverage across the application.

This gives selection a concrete target. Instead of treating every recording as equally valuable, the system uses its execution history to choose flows that collectively reach code throughout the application. Feature flags, roles, permissions and configurations matter because they can lead execution through different branches. The index connects observed journeys to the implementation they actually exercised.

Why use code coverage when executing a line does not prove it behaved correctly? Spencer-Harper gives a sharp counterexample: a Cypress or Playwright test could execute 100% of a codebase while making only one assertion. Most of that execution would go unchecked. His phrase is the distinction to remember: “code covered is not the same as code tested.”

Meticulous tries to narrow that gap by comparing screenshots throughout the flow. The stated example is a 412-step workflow in which a single-pixel difference would be flagged. Executed code is therefore accompanied by much denser observation of its visible consequences than a single end-of-test assertion. That is the basis for Spencer-Harper’s claim that coverage and testing become approximately equivalent in this system: the approximation depends on the behavior producing an observable visual difference.

8:409:10
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

8:40 · section reference included

Dense observation makes determinism essential

The technical ending ties the design together. Coverage-guided selection supplies breadth across the recorded workflows. Screenshots throughout those workflows supply dense observation. But that density brings an enormous volume of comparisons: Spencer-Harper describes screenshot counts on the order of tens or hundreds of millions. Without controlling flakiness, the additional observation could drown the reviewer in noise.

What makes more observation useful rather than overwhelming? The stack below shows the dependency. Controlled replay makes differences more meaningful; frequent screenshots let selected flows expose changes along the way; coverage-guided selection determines which recorded paths contribute breadth. Removing browser variability is consequently part of the verification mechanism, not merely a convenience for running tests.

The merge decision still ends with judgment. In the demonstrated pull request, the name change and dropdown error both arrive as visual differences. The system’s contribution is to reach those states repeatedly and make the consequences inspectable before shipping. That is how the proposed verification layer supports the ambition from the opening: making larger changes, including dependency upgrades and refactors, without expanding manual exploration at the same pace as code generation.

How it fits togetherThe dependencies behind broad visual verification

Recorded network responses isolate tests; browser scheduling controls timing variation.

Frequent screenshots make coverage-guided replay informative, but only if repeatable execution keeps incidental differences from overwhelming the result.

9:4010:09
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

10:09 · section reference included

Read the complete timestamped transcript
  1. 0:12

    Hello, everyone. Hi. Um, my name is Gabe. I'm one of the co-founders and CEO of Meticulous. Um, I'll just wait one more minute for people to come in, and then we can, and then we can get started. Um, so in today's demo or talk, I'll talk about verification of code. Um, I'll spend one minute introducing Meticulous, and then one minute giving some background context on the company, and then sort of eight minutes giving a demo and technical, technical overview.

  2. 0:43

    Um, so if you have exhaustive verification, then you can ship code at the speed that your agents write it. If you don't, then someone, somewhere at your organization is spending time verifying code, and that is now your new bottleneck, and you're forced into a position where you're trading off velocity against bugs and user experience.

  3. 1:12

    So I'll give, um, super quick context on Meticulous. So Meticulous gives you exhaustive or near-exhaustive verification of your front-end code base with zero developer effort. Um, it's deployed across the entire engine organization at all of these companies like Discord, Wiz, Dropbox, Notion, ElevenLabs, LaunchDarkly, and many, many other great companies. So if an engineer touches front-end code at one of these companies, they use Meticulous every, every

  4. 1:42

    day. However, before we get into the product and talk more about verification, um, let's just cover the current, sort of, state of the world. So AI writes code faster than humans can review it, and review and verification is now the new bottleneck.

  5. 2:04

    The other state of the world is that assertion-based testing alone isn't enough. So no matter how hard a human or an agent tries to define the correct behavior of software, a priori, up-front in assertions, the space of possible regressions is too vast to exhaustively cover.

  6. 2:28

    And so a good clarifying question is, if an engineer in your organization, uh, AI-generated this pull request, would you feel comfortable hitting the, the merge, merge and ship button? And in most organizations, the answer is no. You would want to go and do the hard work of checking out all the different feature flags, all the different roles, permissions, settings, configurations, and the various different edge cases in order to understand the exhaustive

  7. 2:58

    impact of this, of this change and see if you can spot any issues.

  8. 3:04

    So there are three sort of like downstream consequences of that current state. The first one is bugs or regressions, uh, which have a, a business impact. The second one is engine organizations will spend double-digit percentages of their time maintaining end-to-end test suites. So across manual validation, review, uh, debugging flakes or updating or maintaining the test suite, um,

  9. 3:34

    all spend a lot of time maintaining, maintaining, maintaining tests. And then the third, like, problem or downstream, downstream consequence of, of the state of the world is that, um, if you, if you can get exhaustive verification, then you can program in a new and different way. So all of a sudden, you can bump all of your dependencies, you can do sweeping refactors, you can make any AI-generated change with complete, complete confidence.

  10. 4:04

    So then the question becomes, um, that sounds great. How do you, how do you get exhaustive verification? What does it look like, and how do you, how do you get there? Um, so I'll now give you a demo of how, how Meticulous, Meticulous works. So we give you-- Uh, give me one second to set this up.

  11. 4:34

    Great. Um, so we give you one line of JavaScript to inject onto non-production environments, so that's localhost QA dev staging. That JavaScript instruments the web browser and records thousands or tens of thousands of flows or workflows, so someone clicking the login button, clicking the settings panel, clicking the analytics panel. And then when you open a pull request on your CI runner, you spin up your web application on localhost three

  12. 5:03

    thousand, and Meticulous takes a subset of those recorded workflows, and it replays them. So it dispatches each event one by one. And as it dispatches those events, so click the login button, click the settings panel, click the analytics panel, it takes a screenshot at each atomic moment throughout time. And so it generates a sequence of screenshots for every workflow, one before your code change and one after your code change. It then diffs those two sequences together

  13. 5:33

    to show you what would be different about your application if you were to merge that code in before, before you do. So in this example, there's two changes here. One is someone's changed the name, and the second is they've introduced an error to a, to a dropdown. And so when you open a pull request, Meticulous posts a comment, typically within a few minutes, and when the developer clicks into this comment, the Wi-Fi is slow, so I'm just gonna switch tabs. But when a

  14. 6:03

    developer clicks into this comment, Meticulous will not tell you whether or not there's a bug. What it will show you is a set of diffs that show you what's different about that application if you were to merge that code in. And so each diff has a before state and an after state. And so by reviewing the diff, you can determine whether that change is expected or unexpected. So here's the name change on light mode, here's the name change on dark mode, and then this is a

  15. 6:33

    manifestation of that logical, logical error. So typically, when you write an assertion-based test, you, um, use your business context and judgment to encapsulate the correct behavior of software upfront and, and encapsulate that into your assertion-based test. In Meticulous, we just show you what's different, and you exercise your judgment at review time or, or, or an agent, agent does. Um, give me

  16. 7:03

    one second to switch this back.

  17. 7:10

    Cool. Cool. Um, so there are sort of three technologies that are really important for helping build a mental model of the tool and how it works, how it works and why it works. So the first one is that everything is mocked out. So, um, at record time, we record all the network requests and responses, and then at replay time, we stub those in. And we perform that mocking for two reasons. One is to make the test idempotent, so you can run them a

  18. 7:40

    thousand times in a row and get the same results each and every time. And the second reason is it means that every test is isolated from every other test, which eliminates the risk of race conditions and allows you to parallelize the test horizontally and get, get the results in a, in a few minutes. The second key technology is that you get radically less or orders of magnitude less flakes with Meticulous than any other tool, including Cypress or Playwright. And the reason why is we augment the

  19. 8:10

    browser from the scheduling engine layer up to be fully deterministic. So when browsers were first designed, determinism was not a design goal, and so there's many different sources of randomness. For instance, an animation spinner depends upon the clock speed of the CPU or the machine rendering the browser, or the interleaving between setTimeout and setInterval. And so we handle all of that randomness and, uh, make the browser deterministic or near deterministic so that you have radically, radically less flakes.

  20. 8:40

    And then the third technology, which is the most, um, important one for grokking, like, why this works, is if you want this to be exhaustive, it has to cover every feature flag combination, every permission, every role, every config, every possible branching through the application. So how do you, how do you do that? So one technique that we use is that for every session or flow that we record, we replay it once against main branch or master branch, and we monitor what lines of

  21. 9:10

    code get executed. And then we build a map or an index that maps each individual workflow to lines, lines of code executed and choose a subset of workflows that maximizes code coverage across your application. And typically, uh, give me one second. Typically, I would say that code coverage, in our opinionated view, is a terrible metric for every tool in the world, apart from Meticulous. And the reason why is you could write a Cypress test or Playwright test that covers a hundred percent

  22. 9:40

    of your code base but only makes a single assertion. And so code covered is not the same as code tested. But with Meticulous, it actually is approximately the same because we do this really unusual and radical thing, which is we're taking a screenshot at every possible moment throughout a flow. So if there's a four hundred and twelve-step flow and there's a single-pixel diff, Meticulous, Meticulous will, will, will flag it. And so you can start to see how all these technologies layer

  23. 10:09

    together. Um, you only get exhaustive verification because of this code coverage algorithm. That algorithm only really works because you take a screenshot at every possible moment. If you take a screenshot at every possible moment, you're taking on the order of tens, hundreds of millions of screenshots. And so traditionally, you would just drown in noise or drown in flakiness. And so you have to solve flakiness at the root level, um, which is, which is what we, what, what we, what we do by, by augmenting, augmenting the browser.

  24. 10:40

    Um, these are some nice things that our customers, customers have said about us. Um, and that concludes today's talk. If you're interested in finding out more, then come to booth LG6, and thank you so much for listening.