AI Engineer World's Fair 2026
The Death of the Code Review: What the Data Actually Says — Laurie Voss, Arize AI
Read the talk
The Death of the Code Review: What the Data Actually Says
Laurie Voss examines the gap between generating code and shipping software, the mechanisms behind automated review, and the human judgment needed to build systems worth trusting.
From a talk by Laurie Voss
At a glance
Ideas worth remembering
Faster generation shifts the constraint toward deciding which changes deserve to ship; human inspection has attention and burnout limits.
Mergeability includes regression safety, scope, test quality and maintainability. Passing existing tests establishes only part of that decision.
Automated reviewers combine repeated passes, selective findings, tool-assisted investigation and repair, while human acceptance supplies feedback for improvement.
Removing humans from individual approvals leaves human judgment in the harness, whose reviewers also need inspection for blind spots and manipulation.
Build review systems around explicit quality standards and domain context, then observe production trajectories to check what the shipped system actually does.
741% more code, 30% more software
Code can arrive much faster than a team can decide whether to trust it. Laurie Voss, Head of Developer Relations at Arize AI and co-founder of npm Inc., opens with a study tracking more than 100,000 GitHub developers and matching their activity against telemetry showing when they began using AI. Developers who enabled autonomous agents wrote 741% more code, while shipped software increased by only 30%. In Voss’s account, the study identifies review as the bottleneck between those outcomes.
Generating a change and accepting responsibility for it consume different resources. Agents increase the amount of code entering the pipeline; they do not automatically increase the capacity of the people checking it. Establishing trust remains expensive, particularly when a change touches sensitive code or can damage systems beyond the immediate task.
The migration examples make that imbalance concrete. Voss cites Stripe and Anthropic launch materials reporting a 50-million-line Ruby migration in one day, against an estimate of more than two months of team work. Bun reported moving more than a million lines from Zig to Rust in six days. Generating changes at that scale creates an urgent question: what would justify accepting them?
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A 10,000-line PR meets a human attention limit
The obvious response is to assign developers more review work, perhaps making it their entire job. But sustained inspection has limits. Voss cites a Cisco study conducted over ten months, covering 2,500 reviews and 3.2 million lines of code. Reviewers became less effective at finding defects beyond roughly 400 lines in one sitting, with effectiveness dropping sharply above 450 lines per hour. These figures describe that study’s conditions rather than a universal limit for every codebase.
Follow one 10,000-line agent PR through that constraint. At 450 lines per hour, reading it requires more than 22 hours; Voss estimates three or four working days for a real review. That is one change from one agent. Running several agents increases incoming work without increasing a reviewer’s attention, while turning developers into full-time reviewers adds the burnout problem. “Review harder” does not close the throughput gap.
An alternative changes the unit of human work. Peter Steinberger’s proposal is to design the loops that prompt coding agents, rather than continually prompting them yourself. Voss connects this with Andrej Karpathy’s argument for removing humans from the execution loop. The engineering question becomes whether the loop can inspect enough of the system to make a trustworthy decision.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Passing the grader does not settle “Would you merge this?”
OpenAI’s internal-product experiment takes loop design seriously. Starting from an empty repository, agents produced about a million lines of code and approximately 1,500 merged PRs over five months, with three engineers. Even the agent-review scaffolding was agent-written. Human PR review was optional, and most review work happened between agents. The product’s purpose and source were not disclosed in the account Voss describes, limiting what other teams can conclude about where the approach transfers.
The bet is that loop design can substitute for direct inspection. Its effectiveness depends on visibility: a check can judge only what it can observe. Tests supply an inexpensive proxy for quality, but passing the tests establishes less than a maintainer’s decision to merge.
METR tested that gap by hiring four active maintainers from projects represented in SWE-bench to inspect PRs that had already passed its grader. Maintainers considered only about half mergeable. Rejections included code-quality problems and changes that broke behavior outside the test suite. The agents could not revise their work after feedback, so this measures first-pass mergeability rather than the eventual outcome of an iterative contribution. For the question of removing humans entirely, however, requiring maintainer feedback would restore the checkpoint being tested.
Cognition’s FrontierCode is introduced around the maintainer’s actual question: “Would you merge this?” Voss describes more than 20 maintainers creating 150 tasks from their repositories, each representing more than 40 hours of expert work. The evaluation broadens the decision beyond the requested behavior:
- Behavioral correctness: Does the change do what the task requires?
- Regression safety: Does it preserve other behavior?
- Scope discipline: Does it stay within the intended task?
- Test quality: Are its checks useful?
- Maintainability: Is the resulting code suitable for the repository’s future?
The reported contrast is striking: 88% on SWE-bench Pro versus 29% on the hardest slice of a maintainer-oriented evaluation, with another model scoring under 6% on that slice. The benchmark attribution is unresolved because this passage alternates between FrontierCode and Frontier Bench; the model identities are also uncertain. These figures therefore support the reported contrast between test performance and mergeability, without establishing a verified model-by-benchmark comparison. They also describe different benchmark conditions, rather than a controlled test of one isolated review skill.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A review standard can become a training signal
A mergeability benchmark would do more than rank models. Voss draws on Sarah Guo’s explanation of why coding models improved quickly: compilers and test suites provide cheap verification. A training process can repeatedly generate candidates, check them and use the results to improve subsequent behavior. Making maintainability, scope discipline or regression safety similarly checkable would give model trainers additional targets.
That makes writing the rubric consequential. Behavior rewarded by today’s evaluator can become behavior produced by tomorrow’s model. Voss points to CriticGPT, trained by OpenAI in 2024 to find bugs in model-written code, as prior work: humans assisted by the model outperformed either humans or the model alone. Better review assistance can improve the feedback used to train coding systems.
Completeness remains the hard part. A machine-checkable rubric can capture useful parts of human review without capturing every reason a maintainer would reject a change. Voss treats a comprehensive mergeability standard as future work, then turns to the narrower review systems already operating in production.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Automated reviewers must earn attention before they can repair code
Automated review is already a substantial production workload. Voss reports 60 million reviews by GitHub Copilot’s reviewer, accounting for more than one in five reviews on GitHub. That establishes widespread use; useful individual findings still require careful engineering.
Cursor’s architecture exposes several mechanisms for improving those findings:
- Repeated inspection: Its first reviewer ran eight passes over each diff and shuffled reviewer order because order affected the results.
- Agreement filtering: Multiple passes can discard false positives. Voss cites a Peking University study reporting up to a 44% improvement in review quality from keeping findings on which passes agreed.
- Directed investigation: The rebuilt reviewer reasons over the diff, calls tools and chooses where to dig.
- Default suspicion: The model needed instructions to assume something might be wrong, rather than accept plausible-looking code at first glance.
False positives matter because every reported issue spends a developer’s attention. Repeatedly flagging correct code teaches people to ignore the tool. Default suspicion addresses a separate failure: declining to investigate because the change looks reasonable. A useful reviewer must search actively while keeping its eventual reports selective.
Review is also joining repair. Cursor’s reviewer can spawn a fix agent from a finding, write a patch and return a diff for human approval. Running the code to demonstrate that the reported bug is real is described as a next step Cursor wants to take. Repair reduces the work between finding and resolution, but approving the patch remains a human decision.
Where does the human remain when review produces a repair? The flow below follows a diff through investigation, a finding and a proposed patch. The checkpoint has moved: the developer receives a repair to approve, while the reviewer and fix agent perform the intervening work.
Other vendors address different parts of the problem:
- CodeRabbit: Voss reports more than 13 million requests reviewed, illustrating the scale of dedicated review services.
- Greptile: A repository-wide graph helps the reviewer follow a change into distant code.
- Graphite: Accepted and rejected suggestions form an evaluation set for improving the reviewer.
Human acceptance supplies the common feedback signal. Cursor calls its measure the resolution rate and reports raising it from 52% to over 70%. These companies tune their harnesses against developers’ judgments at scale. Acceptance tells them whether people take a suggestion; it does not independently prove the resulting software correct. Humans still review the reviews.
The proposed change enters review.
Investigation and patch generation happen before the human checkpoint; the developer still approves the returned diff.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Bun’s port exposes what a passing test suite leaves unresolved
Nicholas Carlini’s C compiler experiment separates being “in the loop” from being “on the loop.” Sixteen agents built a compiler in Rust over about 2,000 sessions, and it could compile the Linux kernel. Humans did not approve each code change, but a human built the test harness and feedback systems that judged the work. Removing per-change approval still left human judgment embedded in the machinery.
Bun’s migration develops the earlier throughput example into a review problem. Agents ported about a million lines from Zig to Rust in six days. The existing test suite served as the gate, and 99.8% of it passed. When the agents deleted the old Zig code in one giant PR, another robot flagged the deletion as “AI slop.” The mismatch is funny, but also useful: judging whether a diff looks plausible is different from checking whether a migration has completed.
Closer inspection changed what the passing tests meant. Voss reports 13,044 unsafe blocks in the Rust port, compared with roughly 74 in a comparable human-written codebase. In her explanation, these blocks place responsibility for memory-safety assumptions on the author. Their count does not prove that many defects exist; it identifies safety obligations that successful public-interface tests do not establish.
The sequence matters. Agents changed the implementation language. Existing tests checked much of the observable behavior, and those tests largely passed. Inspection then revealed thousands of internal safety assertions outside what the suite was designed to assess. Behavioral compatibility and unresolved safety obligations can coexist. The gate did real work, but it could not answer every question raised by the new implementation.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
OpenAI moved review into execution, logs and recurring cleanup
Returning to OpenAI’s experiment reveals the infrastructure behind optional human review. Codex reviews its changes, calls additional agents to review those reviews and continues until every agent reviewer is satisfied. The system can boot at every change, letting an agent run it, inspect the UI and check whether a bug was fixed. Access to the logging stack adds information that a diff alone cannot supply.
What can this review loop see? The diagram separates the proposed change from the running application and its logs, then shows additional agents inspecting the review. Agreement supplies the stopping condition. Execution and logging broaden the evidence available before the reviewers reach it.
Accumulated low-quality code required another mechanism. For a period, humans spent every Friday cleaning up “AI slop.” That recurring manual work did not scale, so agents were trained to identify and remove it. The response was to improve the inspection and cleanup system rather than ask the same people to try harder.
Experience can still force a reversal. Voss recounts Dexter Horthy spending six months advocating shipping without reading the code, then publicly retracting that advice after large parts of the system had to be ripped out and replaced. His plea, “Please, please read the code,” gives an operational counterexample to assuming that successful generation makes inspection unnecessary.
Some reasons to preserve code live outside its tests. Voss uses a module with three external users and an obscure cron job depending on that module as examples. A change can satisfy the visible task while deleting something another system still needs. Human checkpoints therefore tend to survive where correctness is expensive to establish, consequences are large or someone must accept responsibility for the result.
Codex reviews its own changes.
The running application and logs broaden what Codex can inspect; additional agents review its reviews until all reviewers are satisfied.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The code under review can influence the reviewer
Moving human judgment into a review system creates another review task: checking that system itself. Voss cites Anthropic’s automated security reviewer README warning that the action is not hardened against prompt injection and should review only trusted PRs. A reviewer reading attacker-controlled material can be influenced by that material while deciding whether it is safe.
In a study Voss describes, vulnerable code presented with an innocent-looking commit message fooled an autonomous review agent in 88% of attempts, compared with 35% for human reviewers receiving the same attempts. These rates belong to that attack setup, rather than automated review generally. Within it, the framing influenced acceptance despite the underlying vulnerability.
Confident presentation compounds the problem because coding agents can supply reassuring explanations alongside faulty code. The review system has to evaluate the change without inheriting its author’s confidence. Evaluating reviewers is itself unsettled: Voss notes benchmark scaffolds leaking answers and disagreement over how to measure review quality. The grader needs scrutiny alongside the code it grades.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Production is the last reviewer; the practical work is the harness
Once pre-merge review is automated, production behavior becomes the final check against reality. Voss calls it the “last reviewer standing.” The useful object of inspection is the trajectory: what the system actually did, step by step, while interacting with the real world. A passing pre-release test cannot substitute for observing those consequences after deployment.
Code review becomes an engineered system with several things to inspect: benchmarks, classifiers, rubrics, test suites and evaluations. Humans move from reading every line to designing and supervising those mechanisms. Each layer still needs someone to decide whether it deserves trust.
The closing instruction—“stop reviewing PRs”—is deliberately provocative after the failures just described. Voss’s practical recommendation is to spend human judgment building a reliable review harness: codify what good means, supply company context and domain knowledge, and give agents rules and evaluations they can work against. The earlier exceptions still matter. Expensive correctness checks, large consequences and accountability justify human intervention.
The proposed payoff is greater shipping capacity from judgment applied higher in the system, rather than repeated over every diff. This is a direction for teams to pursue, not a measured speedup guaranteed by the examples. The goal is to explain, with evidence, why the software you shipped deserves trust.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
Related talks
- Beating RL With Reflection: GEPA and Optimize Anything
Develops a practical counterpart to review-harness design: human annotations shape an evaluator, while execution traces help improve prompts, agent programs and repository skills.
Read the complete timestamped transcript
- 0:12
Hello, everybody. Thank you for following on a session about the death of the code review with a session about the death of the code review. Uh, who knows how scheduling decisions get made, but, uh, [REDACTED] decided it would be funny for those two to be back to back. Uh, hi. Uh, I'm Laurie. I'm Head of Developer Relations at Arize AI. Uh, some of you may remember me from when I used to co-found npm Inc. These days, I think about AI and how to test it. Uh, I'm here to talk about a problem that everyone is seeing right now.
- 0:42
The rise of AI agents has dramatically increased how fast developers can produce code, uh, but the speed at which humans can review code... Uh, yeah, I am having a talk right now. Thank you. That was very good, [REDACTED]. Uh, the te- the speed at which humans can review code has stayed exactly the same. Um, and that is creating a new bottleneck that everybody is feeling. Engineering teams across the industry are feeling it. Uh, so what can we do about it? Some
- 1:12
options are, uh, we could skip human review entirely. Uh, some folks are trying that. Can we automate reviews reliably? Some folks are trying that too. Uh, what we're doing-- going to do today, what I'm going to do for you, is I'm going to look at what the industry is really doing right now, uh, and try to figure out what you can do, uh, today when you leave the room. So let's start with two numbers. Uh, recently, three economists tracked more than a hundred thousand GitHub developers, uh, matched them against telemetry
- 1:42
that showed exactly when each one started using AI. Uh, the developers who turned on autonomous agents wrote seven hundred and forty-one percent more code, but only thirty percent more software shipped. Uh, that is the scale of the problem right there. They were writing code nearly eight times faster, uh, but their actual shipped software only rose by a third. Uh, and the authors of the study are blunt that review was the bottleneck. Um, the problem is that the route to production still runs through humans, uh,
- 2:13
and the human steps, uh, in particular code review, uh, choke everything downstream. Um, producing code is suddenly a whole lot cheaper. Knowing whether to trust it is still very expensive, especially if you have, uh, sensitive code with a blast radius. Um, how can we bring that second number down and thus bring this, the amount of software that we ship back into line with the amount that we can code? Uh, let's zoom out first. Let's be clear that the problem is real. Generation is no longer the
- 2:42
bottleneck. Uh, Stripe and Anthropic's launch materials for Fable this year reported migrating a fifty-million-line Ruby code base in a single day, uh, which is work that they'd estimated would take over two months for a team. Um, Bun, which is now part of Anthropic, reported that they migrated over a million lines of Zig, uh, to Rust in six days. Uh, and of course, all of us in this room are feeling this on a smaller scale. Uh, our agents are generating whole apps, and we're just sort of hitting the merge button and feeling guilty
- 3:12
about the fact that we haven't read these code diffs, and we're just sort of hoping that it works. Uh, we can't possibly keep on going like that. Um, so the obvious answer, and the answer some people are trying, is you just review more. You take people whose job was previously to review code as well as write code and just say, "Review code all the time. This is your job now." Uh, and that doesn't work, uh, partly because that's really boring and those people burn out really fast, um, but also the numbers say that we can't. Um, the
- 3:42
best study we have on this was done, uh, two decades ago at Cisco. Um, over ten months, they took two and a half thousand reviews. They took three point two million lines of code. Uh, and the study says that reviewers stop finding defects effectively if they try to read more than four hundred lines of code in one sitting, and their effectiveness completely falls off a cliff, uh, if they try to review more than four hundred and fifty lines of code in an hour. Uh, if you do the math on that mean-- that means that a
- 4:12
ten-thousand-line agent pull request at that pace would take three or four working days to get a real human review. That is one pull request from one agent, ten thousand lines is, uh, absolutely, you know, absolutely an expectable number of, of lines of code, uh, for an agent that is doing a lot of work for you. Uh, and a developer can now run a dozen agents at once, um, although I'm-- tend to be suspicious of the people who do. Um, so we can't just review harder, and we're all
- 4:42
feeling that too. The people who are trying to review harder are burning themselves out. So what is the alternative? Uh, some people have decided to just stop reading the code entirely. Uh, Peter Steinberger, who is the creator of OpenClaw, uh, says that you shouldn't be prompting coding, coding agents anymore. You should be designing the loops that prompt your agents. Uh, Andrej Karpathy, who is one of OpenAI's founding engineers, has made the same argument about taking yourself out of the loop, uh, because the human in the loop is holding the system back. Uh,
- 5:13
but this goes beyond bold claims on Twitter. Uh, in February this year, OpenAI published an account of building an internal product, uh, with, in their words, "no manually written code." They started with a completely empty repository, uh, and agents wrote everything. Five months later, they had about a million lines of code, about fifteen hundred merged pull requests, uh, pull requests, uh, and they'd done-- used three engineers to do this. Uh, even the scaffolding of the agent to review the agent and stuff like that was also
- 5:43
written by agents. What they said was, "Humans may review pull requests, but they are not required to. Uh, we've pushed almost all review effort towards being handled agent to agent." I think the interesting thing about the OpenAI experiment is that they did not tell us what the product did, and they did not review an open-- they did not release an open source product saying, "And this is how we did it, and this is how you should do it," which suggests to me that there are still holes in that strategy. Um, but it is one possibility. If OpenAI, OpenAI
- 6:13
says that you can do it, maybe you can do it. Um- But one poss-- so agent-- OpenAI says that's one thing you can do. The bet there is that loop design substitutes for inspection. Uh, whether you can build a loop that is good enough, uh, to review all of your code for you depends entirely on what the loop can see. Uh, but that is an enormous caveat. What exactly can the loop see, and is it reliable enough to do code reviews? For years, the industry's proxy for,
- 6:43
uh, reliable re-- for quality has been whether the tests passed, uh, because that is what the benchmarks measured. In March this year, uh, a research group called METR, who you've probably heard about before, uh, tested that proxy directly. They hired four active maintainers of open source projects, uh, from the same open source projects that SWE-bench tests against, and they got them to look at, uh, PRs that had already passed SWE-bench's grader.
- 7:13
So SWE-bench said, "This pull request is, is good enough to merge," and they got the open source reviewers to look at exactly the same PR and say, "Is it really good enough to merge?" And it was only good enough to merge about half of the time. Uh, and the failures weren't about correctness because obviously, the PRs were passing all of the tests. They were about code quality, and they were about changes that quietly broke other code, things external to the test suite. Uh, one caveat that METR mentions is that the agents got no chance to
- 7:43
iterate on feedback. So a human, uh, contributor obviously being told that their PR isn't gonna get merged, get another go at it, they can take another swing. Uh, but that caveat doesn't real- isn't really relevant to our purposes because that's just putting a human in the loop, and what we're trying to see is whether or not we can take the human out of the loop entirely. Uh, Cognition, the makers of Devin, uh, built a benchmark around the maintainer's actual question, which is, "Would you merge this?" Uh, they called it FrontierCode, and they launched it in
- 8:12
June. Um, more than twenty maintainers built a hundred and fifty tasks from their own repositories, each one over forty hours of expert work. Um, Frontier Bench grades behavioral correctness, regression safety, scope discipline, test quality, and maintainability. This is a human review rubric, uh, made machine checkable. Um, and what they found was that Fable 5, before it got pulled and then unpulled, uh, scores eighty-eight percent on SWE-bench Pro, but only twenty-nine percent on the hardest slice
- 8:42
of, uh, FrontierCode. So the same model on the same surface with the same job does fifty-one points less, uh, less well if you're asking not does it pass the tests, but whether or not, uh, you would actually merge the result. Um, and this isn't one model having a bad day. On the same set of tests, uh, GPT 5.5 scores under six percent. So the strongest models we have are nowhere near, uh, passing human review reliably. So somebody has to
- 9:12
write down what mergeable actually means. That is cl-- it's clear that we haven't done that yet. Uh, the moment they do, something will happen. Sarah Guo, who is a prominent investor in AI, uh, wrote about it and said that solving for-- that solving for that benchmark, a benchmark of genuine mergeability, uh, will be a critical turning point in the development of coding models. Um, in the essay where she was talking about this, she made clear why co-- models got so good so fast at
- 9:42
generating code, and that is because a compiler is a free verifier. A test suite is a free verifier, and anything that you can verify cheaply, you can train against until you beat it, and that is what has been happening, uh, at the major model trainers. Uh, so a benchmark of mergeability, if we could, if we could make one, wouldn't only be a measurement, it would immediately become a training signal for the Frontier models. Uh, which means maintainability, scope discipline, regression
- 10:12
saf-safety. The moment somebody comes up with reliable tests for those, the big models will be trained on them. Uh, and whoever writes today's review standard is writing next year's default model behavior. Uh, we have some prior art here knowing we have seen this happen already, uh, which is that in twenty twenty-four, OpenAI trained a model called CriticGPT to catch bugs in model-written code. Uh, human reviewers working on it beat the model alone, uh, and the human alone. Um, so OpenAI
- 10:42
built a model to review the models in order to clean the training, and it created a training signal that's now built into the models. It made the models better almost instantly. But mergeability standards are still in the future right now. We do not have a good rubric that captures everything that a human decides, uh, to, uh, consider when they are deciding whether something is mergeable. And I promised you something that you can do today. Um, so who is actually running automated
- 11:11
review right now? Uh, one company is GitHub. GitHub is-- uh, Copilot's reviewer has done sixty million reviews and now accounts for more than one in five, uh, code reviews on all of GitHub. So machine review of pull requests is the mainstream default, uh, on the world's largest code host. That is more than an experiment. That is a large production, uh, a large production deployment. Uh, Cursor is also doing an enormous amount of code review. Cursor, uh, has published its reviewer's architecture, so we
- 11:41
know how they're doing this. Um, and the details tell you what the job really is. The first version of their reviewer, uh, would run eight review passes over each diff, uh, and then shuffle the order of the reviewers, uh, to review the code again and again, uh, because the order in which it did those reviews changed the outcome of those reviews. What they were doing was filtering for false positives because a reviewer, an automated reviewer that flags something as bad when it's actually good,
- 12:12
uh, is going to get ignored. Um, and it's not just Cursor who's doing this, there's research about it. Uh, a team at Peking University tested the same idea independently. Run several review passes, see-- keep what they agree on, uh, and they found that it revie- raised review quality by up to forty-four percent. Uh, the multi-pass trick keeps getting rediscovered because, uh, false positives are the thing that kill a reviewer, and it actually works. So then Cursor rebuilt, uh, their code reviewer,
- 12:42
um, so that the model reasons over the diff, calls tools, and decides where to dig. Um, and my favorite detail from the rebuild, uh, was that they had to tell the model to be more suspicious of the code. Uh, the model tended to look at the code and say, "Well, that looks good to me, so ship it," which is exactly what a human would do in that situation. Uh, they had to tell it, "Don't trust the code by default. Assume there is something wrong with it." Um, and somewhere in that sentence is a whole, is a whole talk about what a good review actually is. Uh, it is about
- 13:12
being suspicious by default. Um, the next thing that's happening is that review is starting to fuse with repair. Cursor's reviewer now spawns a fix agent from its own findings, so it doesn't just flag the bug, uh, it writes a patch, uh, and it hands you back, uh, a diff to approve. Uh, you approving that diff is obviously another code-- is another human in the loop, but we don't want humans in the loop. Um, the next thing Cursor wants to do is the reviewer running the code to prove its own bug report is real.
- 13:42
Um, the line between reviewing and rewriting is getting very thin indeed. Um, and GitHub and Cursor are by no means the only vendors getting into this game. It is, uh, getting very crowded in there. There's a ton of companies, uh, in the field. Uh, CodeRabbit is the largest dedicated reviewer, has now reviewed over thirteen million code requests. Uh, Greptile builds a graph of your whole reposi- of your whole repository so that the reviewer can see how a change lands in distant code. Uh, Graphite builds its evaluation set
- 14:12
from which of its own suggestions developers accept or reject. Uh, but the thing to notice, uh, about what every one of them is doing is that all of them are using the same metric as their definition of success, which is, is a human accepting my answer? Uh, Cursor calls it the resolution rate, and it has driven it from fifty-two percent to over seventy percent. Um, the reviewers are trained on human accept or reject judgment every day at scale. They are training their harnesses to get better at the
- 14:42
definition of good as defined by human acceptance. That is a preview of what the models are going to do, except these companies have already shipped it. So the status quo is you can generate code automatically, you can review the code automatically, but humans still need to review the re- the reviews. Can you skip that part entirely? Can you get the human entirely out of the loop? Uh, there are two prominent projects that I've-- that have tried so far that I've heard of. Uh, in February this year, Anthropic's Nicholas Carlini had sixteen
- 15:12
agents build a C compiler from scratch in Rust. Uh, it was able to compile the Linux kernel, uh, across about two thousand sessions with no human in the loop. But if you read that experiment closely, uh, it's true that there was no human in the loop, uh, no human approving the code as it was written, but there was absolutely a human on the loop. The system, uh, that reviewed the code, the system that checked whether the code was doing what it was supposed to do, all of that, uh, was written by a human. Um, the test harness, the feedback systems,
- 15:42
all of that stuff, um, it was automated testing with tests that took a human to write them. Um, Carlini's own warning when he wrote about it was that it is easy to watch the tests pass and assume that the job is done, uh, and that it rarely is. Another very widely publicized experiment which I als- which I already mentioned briefly in passing was when Bun ported its entire, uh, runtime from Zig to Rust. That was about a million lines of code by agents in six days. Obviously, no human read that diff. Um, the
- 16:12
gate for that experiment was the existing test suite. Ninety-nine point eight percent of the test suite passed, um, so the test did real work. Uh, a fun fact about that experiment is that the PR, uh, where they, uh, where the agents decided that they, they were done with all of the Zig code and they deleted all of the Zig code in one giant PR was flagged by another robot as, "This is AI slop. You can't possibly delete all of your code." Uh, which I thought was a fun aside. Um, but there are some big caveats on that
- 16:42
Bun experiment. Somebody looked carefully at the ported code, and it has thirteen thousand and forty-four unsafe blocks. Uh, in a comparable human-written, uh, uh, Rust code base of that size, uh, you would see about seventy-four. So three orders of magnitude more unsaf-- memory unsafe blocks. An unsafe block is a place where the author asserts rather than proves, uh, that memory is being handled correctly. So the test suite can certify behavior at the public interface.
- 17:12
Um, it cannot certify thirteen thousand assertions that the test suite was never designed to look for. So they took humans out of the loop for sure, uh, but there's now no, no knowing what is lurking under the surface of their Rust as a result.
- 17:28
So let's go back to OpenAI, who have pushed this the furthest, because they show what skipping human review costs. They didn't delete review, they moved it. Uh, Codex reviews its own changes, uh, then calls in more agents to review those reviews, um, in a loop until every agent reviewer is satisfied. OpenAI made Codex bootable at every single change so that Codex can actually run a copy of Codex, look at it, look at the UI, see if the bug is being fixed. Uh, and they exposed the whole
- 17:58
logging stack to the agent. Um, and a line from their write-up says exactly what the Cisco study said, which is that when something failed, the fix was almost never to try harder. It didn't work sometimes. Uh, and one detail that you'll recognize from your own week is that for a while on this project, they had to spend every Friday cleaning up AI slop, uh, by hand as humans. Um- So that eventually didn't scale, and so they trained agents to look for AI slop and get rid of the AI slop.
- 18:30
Uh, review didn't disappear, is the lesson of this, of, of this experiment. Uh, it got rebuilt as a system, and that system is built by humans. So can you skip the human entirely? Uh, someone who tried it for real and then changed his mind is Dexter Horthy. Uh, he spent six months telling people not to review the code. He famously did that at AIE last year. Um, he told people, "Just ship. Let the agent do its thing." Uh, and in March this year, on stage, uh, he took that back.
- 19:00
Uh, he said, "I was wrong. Please, please read the code. We tried not reading the code for like six months. It did not end well. We had to rip out and replace large parts of that system." That is not a benchmark. That is somebody who ran this experiment for real, tried it with real code on a real system, uh, and retracted his remarks. Uh, OpenAI is running, uh, lived with the results, um, and reversed in public. Uh, so Sarah Guo, uh, in the same write-up that I mentioned earlier, talked about why
- 19:30
that happens. Passing the test never told you that the change was the right change. It never told you that this module exists because there are three external users of this module. Uh, it never told you that there's this cron job that nobody will admit to writing that relies on that module existing. Um, there is context outside of the test suite that the tests don't, uh, can't find. So in everything that we've looked at so far, the human, uh, checkpoint is continuing
- 20:00
to survive. It is moving around, uh, but it is surviving in predictable places. One is where correctness isn't cheaply checkable. Another is where, uh, the blast radius is large. In security-conscious environments, every time you tell people, "Oh, we can just get rid of human review," you would get an immediate no. Uh, wherever someone has to put their name on the result. Uh, but what's happening is the human role isn't disappearing. It is moving up the stack, possibly several levels up the stack, from inspecting the code directly to
- 20:29
designing and tuning the systems that inspect the code and designing the definition of good. Which raises the obvious question, who reviews those systems? If you've built a system that does your reviews for you, how do you do the meta review of your reviewer? Um, and the answer won't surprise you. It is humans. Uh, Anthropic ships an automated security reviewer, and in its README, it has a huge caveat, uh, which is, "This action is not hardened against prompt injection attacks, uh, and should only be rev- used to review trusted PRs." So
- 20:59
Anthropic's code reviewer, uh, can be talked out of its findings by the very thing that it is reviewing. Um, and that is a finding that has been reproduced. In study in March this year reinforced, uh, that vulnerable code dressed up in an innocent commit message, uh, fooled an autonomous review agent in eighty-eight percent of attempts. Uh, the same attempts, uh, sent to a human reviewer, uh, passed only thirty-five percent of the time. So you f- you take the human out of the loop, uh,
- 21:29
you don't just lose a reviewer, you lose the thing that was hard to fool. So automated reviews fall for confidently framed bad code, and confidently framed bad code is exactly the kind of code that agents are very good at producing. They're very like, "This is good, and I am ready." Uh, and they, uh, are wrong sometimes. Um, and the other problem is the field doesn't even agree yet on how to grade these graders. Uh, benchmark scaffolds have been caught leaking answers. Researchers disagree on how to review quality at all.
- 22:00
Um, which leaves one reviewer that you can't automate and you can't skip, which is production. Uh, once the pre-merge review is all machines, watching what the code actually does becomes the last reviewer standing. Um, once the code ships, the test result stops being the interesting thing. The trajectory does. What the system actually did step by step, uh, when it ran against the real world. Uh, I'm not going to give a pitch for Arize here. There are enough pitches for Arize at this conference.
- 22:30
Uh, but if you're shipping automatically reviewed code, then systems in production that review what it actually does in the wild become indispensable. Um, so after all that evidence, after all of that review of what the world is doing about automated code reviews, where I land is that code review isn't completely dead, but it is changing an enormous amount. Uh, it is being rebuilt as an engineered system. Humans are moving from being the engine that drives a code review, uh, reading code line
- 23:00
by line, uh, to its pilots. Uh, and given the numbers that we opened with, that is probably the right trade. Every layer of that system, the benchmark, the classifier, the rubric, the test suite, the evals, is itself unreviewed until somebody decides that that is their job. Um, and the teams that win the next few years won't be the ones that generate the most code. They'll be the ones who can say with evidence why they trust what they shipped. Uh, and I opened by promising you that I would give you something practical, that you could walk away with something
- 23:30
that you can do today. Uh, and that is to stop reviewing PRs. Um, it is the wrong level of abstraction for twenty twenty-six. Your human judgment is extremely valuable, but it can be made to scale much further than it is scaling right now. Uh, pour your precious time into building, uh, a reliable review harness. Codify your definitions of good, your company context, your domain knowledge, and then, uh, crank up the agents to work with that. You can go much faster than
- 24:00
one-third faster, uh, if you concentrate your efforts higher up the stack of reviewers, rules, and evals. Uh, I hope this look at what the industry is doing today has helped you what to... helped you decide what you should do and what to expect next. Uh, and thank you so much for your time and attention.