AI Engineer Code 2025
Fixing the PR Bottleneck — Matt Pocock, AIHero
Read the talk
Fixing the PR Bottleneck
Matt Pocock explains how stronger checks, a separate reviewer agent and risk-aware PR descriptions can reduce human review work—and how retrospectives can improve the system that produces the next change.
From a talk by Matt Pocock
At a glance
Ideas worth remembering
Passing checks only helps when they exercise meaningful behavior. Assertions about constants, source ordering or oversimplified mocks can leave the real requirement untested.
Deep modules provide a small interface for testing substantial behavior, reducing dependence on internal structure when tests stay at that interface.
A separate review context can apply repository-specific coding standards after implementation. Clear findings should become fixes; unresolved questions can remain comments.
Allocate human attention using reversibility and blast radius. Reverting code cannot undo every external consequence.
Use retrospectives to turn recurring review findings into better checks, standards, navigation guidance and tool use for future runs.
A software factory needs brakes
Pull requests were piling up before coding agents arrived. Agents now make producing more changes easier, while people still have to decide which changes deserve to merge. In Matt Pocock’s account, the bottleneck moves toward review: increasing the supply of code does little good if the resulting PRs demand more attention than the team can give them.
The software factory adds another source of pressure. Work can start without a person requesting each task: a database report about slow queries might trigger investigation, reproduction or a fix. Deterministic triggers and agents push more changes into the pipeline. Without mechanisms that improve those changes, the factory becomes a “slop cannon.”
Poor code also changes the environment in which the next agent works. An agent exploring a confused codebase inherits that confusion as context, making further poor code more likely. Pocock’s proposed brakes therefore protect both the current PR and the conditions for future implementation:
- Automated checks: linting, tests, typechecking and code-quality metrics apply repeatable rules.
- Automated review: an agent examines what those rules miss, including code structure.
- Human review: a person judges the prepared change after the earlier layers have reduced avoidable problems.
The intended speedup comes from needing fewer human interventions per PR.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Green CI can test the wrong thing
Automated checks are attractive because running them primarily costs CPU cycles rather than model tokens or reviewer attention. A failed test may cause an agent to spend tokens fixing a bug, but that is useful expenditure. Cheap execution makes it practical to run many checks frequently. The catch is that passing those checks only answers the questions they actually ask.
A tautological test asks whether the implementation still resembles itself. Pocock’s first example declares an X post character limit of 280, then tests that the constant equals 280. The assertion repeats the implementation without exercising the behavior that uses the limit. It also couples the test to that internal constant: changing or renaming it requires changing the test, even when the intended user-facing behavior remains intact.
The pitch detail page gives a more revealing example. The intended requirement is visible: the video section should appear after the content plan. The generated test reads the module as text, finds the two names and checks their order in the source file. It never renders the page. A source-layout change can therefore make the test fail without changing what a user sees. The causal mistake is substituting source position for rendered position; the check measures an implementation arrangement rather than the requirement.
Mocks can remove the very failures a test ought to expose. In the useAudioBoost example, the implementation uses the browser’s AudioContext API, but the test substitutes fake methods. Real AudioContext behavior includes error modes under particular conditions. If the fake cannot produce those errors, the test cannot reveal how the implementation handles them, leaving failures to appear in production.
These examples do not require an agent deliberately cheating. Following an instruction to write tests can still produce checks that inspect internals or simplify away difficult behavior. Automated and human review act as “lie detectors” by asking whether a passing test provides meaningful evidence. The next design question is how to make weak tests harder to write in the first place.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Give tests a small interface to exercise
Deep modules hide substantial behavior behind a simple interface, an idea Pocock credits to John Ousterhout’s A Philosophy of Software Design. His comparison is between a module with a small interface and a large implementation, and a module exposing many functions that each do little. The first gives callers fewer details to understand and tests fewer internal details to depend on.
The testing benefit depends on using that interface. Hiding implementation details does not help if the agent reaches inside the module to assert them anyway. The design and the testing instruction work together: provide a narrow place to exercise meaningful behavior, then direct the agent to test there. Returning to the pitch-page example, the desired observation is the ordering of the displayed sections; their order in a source file is an unreliable substitute for that observation.
Pocock’s module-deepening skill produces an HTML document of possible improvements, including before-and-after proposals that reduce duplication and create deeper, testable modules. The proposals still need implementation. A companion codebase-design vocabulary gives a team consistent words for discussing the results:
- Locality: how closely related code sits together, and how a change propagates through it.
- Leverage: how much useful behavior a caller gets from a simple function call.
These concepts help turn a vague request for better architecture into a reviewable design decision.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Make it work, then give review its own context
Good standards still need somewhere to run. The implementer already has three demanding jobs inside one context window: explore the repository, edit the code and debug it using automated checks. Pocock’s model is that implementation is overloaded. Adding a large set of architectural and testing standards to that same window makes another task compete for attention.
The code-review skill instead receives the diff and reads a repository-specific codingstandards.md file in a separate subagent. The diff provides a starting point, while some exploration supplies surrounding context. The reviewer begins without the implementer’s full discovery, implementation and debugging workload, leaving more room to examine the change against the standards.
Where does each responsibility go? The flow below separates the working change from the standards used to improve it. Two inputs meet at review: the implementer’s diff and the team’s coding standards. That separation explains the proposed allocation of context, rather than implying that a second agent automatically guarantees quality.
The sequence resembles red–green–refactor: one context window makes the change work, and another improves its structure. Pocock reports that this division has worked well for him; it is an experience-based recommendation rather than a measured guarantee across repositories. His practical consequence is to keep the detailed coding standards out of globally loaded agents.md instructions and put them where the reviewer explicitly reads them.
Explore, edit and debug using automated checks.
Implementation spends context on discovery, edits and debugging. Review receives the resulting diff and separately loads the standards.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
An automated reviewer should reduce work, not create a comment queue
A generic reviewer faces a specificity problem. Broad instructions to find bugs and security issues can produce irrelevant false positives. Narrow instructions may work for a particular language or repository but lose their usefulness elsewhere. That experience leads Pocock to recommend building a team-specific review process: accumulate coding standards over time, share them and turn neglected engineering documents into instructions the reviewer actually uses.
The reviewer’s output matters as much as its instructions. If every finding becomes a verbose PR comment, the human inherits another queue of decisions: read the suggestion, decide whether it applies and arrange the fix. Pocock’s default is direct: “The reviewer should commit.” Resolve clear problems in the code so the human reviews an improved artifact. Keep comments for findings that still involve questions.
That recommendation changes the purpose of the second pass. It is a stage that improves the candidate before human review, rather than merely describing its defects. Trying to produce perfect code in one implementer run concentrates every responsibility in the busiest context; a separate corrective pass gives the standards a dedicated place to influence the result.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Explain merge danger before asking for attention
After checks and automated review, the PR still needs to be easy for a person to assess. Pocock introduces a PR skill that was still in progress at the time of the recording. Its first principle is that reviews deserve different amounts of attention. The PR body should make the important distinctions visible before a reviewer starts reading the diff.
The first distinction is reversibility:
- Two-way door: the change can be merged and easily reverted.
- One-way door: the change creates consequences that reverting the code cannot readily undo, such as expensive migrations, data loss or an email sent to 60,000 people.
The email is a hypothetical example with a useful lesson: a simple code change can still trigger an irreversible action. Diff size does not determine how carefully it should be reviewed.
The second distinction is blast radius: what could go wrong, and how bad would it be? A short “merge danger” summary combines that scope with reversibility. A localized two-way-door change can justify lighter review; an irreversible change deserves much more attention. The summary helps allocate effort, provided the classification reflects the change’s actual effects.
The reviewer also needs to understand what changes and why. Pocock favors pseudocode and credits the Show Me skill from the HumanLayer skills repository for presenting changes with diagrams and images. Sequence diagrams can explain interactions; a compact CLI sketch can show a new command and two added flags. The useful compression is behavioral: the reader sees the changed command structure or sequence without first reconstructing it from a long description.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Turn the review into a better next run
Human review has a second object: the system that produced the code. Teams share skills and steering files, so a recurring defect may point to a recurring gap in that environment. The aspiration is to avoid writing the same comment twice. A correction should improve the next run as well as the current PR.
The Retro skill takes a completed agent session, a PR together with its session, or a collection of recent PRs and reviews. It suggests automated checks and coding-standard changes that could prevent the same problem. This closes the loop between expensive human judgment and cheaper repeatable enforcement: the lesson from one review becomes an input to later implementation and review.
Retro also looks beyond the final diff:
- Navigation pointers: if the agent struggled to find information, a pointer in
agents.mdmay shorten the next search. - Tool economy: session history can reveal tools or usage patterns that waste tokens.
- Instruction bloat: oversized steering files and skills may need clearer organization.
These address how the agent reached the result, including costs and confusion that are difficult to diagnose from code alone.
How does a human correction reach a future PR? The cycle below makes the intermediate step visible: Retro turns the session and review into proposed changes to checks, standards and guidance. Those changes can then shape the next run. The compounding effect depends on adopting useful suggestions, rather than simply generating another retrospective document.
The ending makes the risk distinction explicit. In Pocock’s proposed workflow, every one-way door needs human review; not every two-way door needs human review. Reversibility and blast radius together determine how much attention a change deserves. Stronger checks, corrective automated review and clearer PR explanations make that allocation possible. Human attention goes toward consequential decisions, while recurring lessons become part of the environment that produces the next change. Pocock points to AI Hero’s skills collection for the skills discussed in the recording.
Implementation produces code and a session history.
Retro examines sessions and reviews, then proposes improvements to the checks and instructions used in future work.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
From the talk
The recording’s destination for the skills discussed, including codebase design, automated review and retrospectives.
Related talks
- Beating RL With Reflection: GEPA and Optimize Anything
Extends the retrospective idea into reflective optimization: execution traces guide changes to prompts, agent programs and repository skills, carrying lessons into later attempts.
Read the complete timestamped transcript
- 0:01
All right. I'm here to talk about Fixing the PR Bottleneck, and this is kind of a grand title for trying to fix the thing that most organizations struggle with, I think, and have kind of historically struggled with before AI. You know, we have always had huge numbers of PRs just laying around that no one's bothered to review, and this has now increased massively because of the new strains on us, because of AI imposing all these weird constraints. And to do this, I'm gonna use the
- 0:31
rubric of my skills, which is-- you've kind of heard about, maybe you've used them. And I have a couple of new skills to announce that are going to hopefully improve the way that you do PRs, improve the way that or improve the speed at which you can review them and do them. Speed. Now, we're being pushed to do more with less, essentially, or more PRs, more work, more stuff. And this is kind of the central promise of AI, that we're going to be able to use these agents to
- 1:01
scale ourselves up to do more work. And this has resulted in the software factory, probably the biggest buzzword of the day. Everyone's talking about software factory that I've chat to. And I think of a software factory as primarily something where instead of the human initiating all of this work, we're gonna pass some of that initiation. Some of the initiation is gonna be done by agents. And I think of there
- 1:50
Maybe from there you have a classifier like Jev come in and say, "Okay, let's turn that into a fix or turn that into a reproduction," or maybe I ping people straight away. And maybe you have other things. Maybe you have PlanetScale hooked up, so it gives you query reports on the slow queries on your database. Maybe that then triggers a different, um, thing of your software factory. All of this is not humans triggering it, it's, uh, deterministic code triggering it, right? And so these accelerate your software factory. They push more code through it. But then you need brakes,
- 2:20
right? If you just have permanent acceleration pushing stuff through your factory, you're gonna end up with a slop cannon, right? You're just gonna end up with a ton of slop crappy PRs that you're not gonna be able to touch or review or even freaking look at. So you need brakes. These are mechanisms that slow down, that increase quality, that make sure that your code base doesn't turn into a software entropy nightmare, because code is the environment your agent operates in. And if you have bad code in your code
- 2:50
base, that is going to beget more bad code. And so I'm gonna talk about these three brakes in this talk, and talk about how we can use them to counterintuitively go faster. So automated checks. These are the deterministic checks in your repo that we've had for thousands of years or, you know, since the '50s. Deterministic checks where we have linting and tests and type checking and code quality metrics, all of these things going together, and they all work the same every time.
- 3:21
Layered on top of that, we have automated review. So we have agents who look at our code and say, "Okay, you know, these are for the things that the tests didn't catch," or, "This is looking at the structure of the code base in general." And then on top of that, the third layer, the final layer is human review. So people looking at the PR. And these three layers form this kind of cake that we end up with when we get to human review. And so more speed of course means more PRs, and so the goal here is to make human review
- 3:51
faster by leaning on those first two phases. And we've got to stop the slop. That's the first principle here, which is if you raise the quality of the code that you're shipping, you're gonna end up doing less human review because it's just gonna be better work, and so you're gonna end up needing to, uh, make fewer interventions. So automated checks. Now, automated checks are cheap. That's the cool thing about them, is that they don't cost tokens like automated review does. They don't cost human effort, they just
- 4:21
cost CPU cycles. So these are, for instance, you know, you run your tests on every code change. Maybe those tests do incur some tokens because let's say, um, an agent, uh, creates a bug and the tests catch it, then you need to spend some tokens to go and fix it. But those are tokens pretty well spent in my opinion. So checks are cheap. That means you can layer on loads and loads and loads of them on your repos, and you're probably not using enough of them or not being creative enough with your use. But checks can
- 4:50
lie.
- 4:53
Does a green CI mean that the code is ready for merge? No, it does not. And so we've always needed on top of these checks some extra layer to figure out if there's anything catastrophically wrong with the code before we ship it. And so all of the other phases, the human review and automated review, these are lie detectors. These are for finding lies in the automated checks. Now, I wanna show you some of these lies first of all, because this helps when we're thinking about code and thinking about automated checks to see how bad it is and how
- 5:23
bad things can get. The first is tautological tests, a test that just reasserts the implementation. Opus 5 got addicted to these. I don't quite understand why. It would have, for instance, x post character limit equals two hundred and eighty. Can anyone guess the test that was written to test this behavior, right? You've probably seen this a thousand times. This is real code from agents or from stuff that I found my agents doing. It said expect x post character limit to be two hundred and eighty. So the implementation
- 5:52
looked like that, and the test essentially reasserted the implementation. That is a tautological test. And tautological tests are bad because they're extremely structure sensitive. They're very sensitive to the actual internal workings of the system. So it means I cannot change that constant without a test failing. But I cannot rename that constant without a test failing. Like I have to do-- it's so tied into the structure of my system. And I found another
- 6:22
one which is even more egregious, I would say. This is incredible, uh, test. What it's doing here is it's essentially testing whether two things in the UI appear in the right order. So it's checking that the pitch detail page, uh, or rather the video section comes after the content plan. What it does is it doesn't render it to a screen, it just reads the actual file, the module into its own memory, and then it's-- finds the right thing, so
- 6:52
finds content plan, finds videos, and then it expects the videos to be after it in the source material. Which is crazy if you think about it, 'cause I can just m-- like change the way the source material looks and this test will fail. It's too sensitive to the structure of my code base. So that's another way that automated checks can fail. And there are also-- or sorry, automated checks can lie. There are also tests that literally cannot fail, and this will feel familiar to you if you've abused mocking in the past or
- 7:22
abused various things. For instance, here we have a useAudioBoost function, and this useAudioBoost function, uh, its internals use the AudioContext API in the DOM. Don't worry if you don't know any of this. But what we're doing here is we're just stubbing it out with some fake methods. And it turns out that AudioContext has some complicated error modes, and it will fail if you use it under strange conditions. And so just doing this means our tests cannot fail using those modes, and we're gonna hit strange errors in
- 7:52
production that our tests can't fix. And so the question is then, if you can cheat on automated checks, if, you know, and even in good faith ways as well. The, the AI isn't trying to write bad tests here, it's just taking our instructions and writing tests that are too tied into the structure instead of actually executing code. So how do we make automated checks harder to cheat? If we can do that, then we can increase the quality of those checks, which means which
- 8:22
increases our quality bar. And the first thing I really like about this is codebase design. So you can actually design your way out of these bad checks. So what does good codebase design look like? I've talked about this before in previous AI engineer talks I've given, which are deep modules. These are modules that hide complex behavior behind simple interfaces. This is a John Ousterhout idea from the philosophy of software design. If we look at these two modules, we've got A, which has a
- 8:52
large implementation hiding behind a tiny little interface at the top, okay? And then B is a large interface, lots of functions you can call, and those functions individually don't do very much. Does that make sense? Yeah? Now, if you have a deep module like A here, you're gonna have fewer structure-sensitive tests because you're hiding more of the implementation behind that interface. If it's just testing at that interface, you're gonna get better tests. And so your job
- 9:22
here is to force the agent to use that little interface instead of reaching into the implementation to test these weird implementation details. So I've got a skill for this. You can have like the weirdest vibe coded like code base, the crappiest code base that you've ever set your eyes on, and you can run this skill on it, and it will make it better. What this does, it essentially gives you opportunities for deepening modules, and kind of looks like this. Raise your hands if you've used this skill, by
- 9:52
the way. Not sure how many of my folks are in this room. Yeah, okay. It's really freaking nice. Essentially, it gives you a HTML document, and I'll get out of the way here, of all of the different, um, potential opportunities it sees. So we can see here we have a before and we have an after where we're sort of like reducing duplication. We're creating a nice deep testable module. And then you can go ahead and implement that. And attached to this, there's also this kind of language that I've put together for describing modules because like if you
- 10:22
ever try and read up about how to structure a code base, you're gonna find 20 different approaches, and they're all gonna be called DDD. And like what you need is a consistent language that you can use in your team to talk about this stuff. And so I have a little codebase design skill that defines what locality is, defines what leverage is, defines what a seam is. I was using seams before they were cool. And what locality means is kind of how, uh, well located together all of the code is. How can you
- 10:52
change like a small change in one module and have it ripple out? And also leverage is what you get when you have a deep module because the caller, the person who's actually calling that module, gets a lot of value of calling a simple function. Both of those are very good in code bases, and good for agents too it turns out. But I'm sort of describing all of these highfalutin coding standards, but how do we actually make sure the agent does them, right? How do you make sure the agent
- 11:22
creates deep modules and creates good tests and doesn't write these crap tautological ones or structure sensitive tests?
- 11:30
Well, I do think most people get this wrong. And my first piece of advice is don't put coding standards in your implementer agents, okay? Let me explain this. If you imagine the implementer agent kind of looks like this, where this is all of the things the agent needs to be able to do in its single context window. It needs to be able to explore, like to look for the code that it's going to change. It then needs to actually change it, so make the updates to the files in the
- 12:00
green. And then it needs some budget for actually debugging the thing, so for running those, uh, automated checks, for actually checking and verifying that it works. Now this is quite a lot of work it turns out, and if you try to impose your coding standards on it as well, it's going to perform worse.
- 12:21
So implementation is overloaded. That's the mental model I want you to have. And so can we find a way to impose those coding standards in a way that isn't so overloaded? Well, this is my effort. This is my code review skill, and it receives a diff, and it reads a file inside the repository called coding standards, which you can write, you can customize, and then it checks if the code follows those standards. And so if we look at the reviewer agent here, it also crucially
- 12:51
runs it in a sub-agent. So it's got its own context window to kind of handle here. It's got its own budget. This one, it needs to do some exploration, right? Because sure, it receives the diff, so it knows exactly where it's located, where the code is, but it should probably do a bit of exploration just so it has the wider context, understands the code. But it doesn't need to do any implementation, doesn't need to do any debugging. So while implementation is overloaded, review is actually underloaded, so it doesn't have that many
- 13:21
jobs to do. This means you can pile in a bunch of coding standards to it, and it will do a much better job than if you tried to do it with implement. So I think of this, and this is uncomfortable, right? 'Cause we, we all want to be able to just get good code out the first time. But I think of this as the two-part process for writing good code, which is implement, you make it work, and then code review, you actually make it good. You impose your coding standards. And for, you know, the retro developers among us,
- 13:51
this is essentially a red-green refactor approach. We use one context window to make it okay, do the red-green, and then we do another context window to refactor it. That's at least how it works in my head, and it's been very successful for me. This means that when you have coding standards, you don't put them in global scope. You don't put them in agents.md 'cause then they sort of drown out your implementer. It may read them, it may not. You put them in codingstandards.md, and that way just the code review agent does it.
- 14:22
Now, I've got another idea here, which is I've been talking to lots of people today, and lots of people saying, you know, I've been talking about review and automated review increased, you know, Fixing the PR Bottleneck. So many folks say, "Oh yeah, we just use a third-party service. We use, uh, Cursor Bug Bot, we use CodeRabbit," or something like that. I think that I've, I've really tried to make a generic code review skill in the past that finds all the bugs and does security review and that kind of thing. It turns out it's really, really hard because you either
- 14:52
make it too general and it just gives you false positives that aren't actually relevant to your use case, or you make it too specific. You say, "Okay, find all the TypeScript stuff," and then Rust people can't use it. So I would say don't outsource automated review. Build your own. Build up your own coding standards over time. Share them across your team. And if you have an opportunity to impose coding standards, if you've got some docs sitting around that no one reads, this is the place to put them in. And also,
- 15:22
when you have this automated reviewer, a really natural inclination for lots of people is to say, "Oh yeah, my code review agent, what it does is it reads the code and then it comments on the PR." Ugh. So what your code review d- agent is doing in that case is it's providing more work for the human reviewer. The human reviewer then has to read all of these verbose comments and figure out, "Okay, do we implement this? Do we implement that?" The reviewer should commit. It should make fixes. So it should actually change the things that it finds,
- 15:53
'cause then when the human comes around, you're reviewing a really nice artifact. If it finds anything that it has any questions over, then of course it can comment, but the default should be commits. Stop trying to one-shot good code. Stop trying to make the implementer agent the only thing that you do and go, "Okay, I'm gonna force it to be amazing." It takes a little bit of, you know, thinking your way out of there, but once you realize it, it is fabulous. So okay, with all of that process, we've run our automated
- 16:23
checks. We've now run our automated review to make sure the automated checks aren't lying. How do we then maximize the PR's quality in terms of human review? How do we get it, like, working the best it can? So we need a human-friendly PR, and this is a new skill, uh, coming into the repo, which is currently in progress, but I'll be releasing it soon, which is the PR skill. And this is one I've been mulling over for a long, long time, haven't quite figured out what the state of the art is, and I've realized the best way to make a
- 16:52
PR skill is just to steal all the best ideas that everyone's got, and it's very nice. Now, what does a good PR body look like? How would you recreate this skill on your own? Well, the first principle is some reviews are more important than others. Not every review is essential, right?
- 17:14
And once you understand this, you realize, okay, that means I can focus my energy on the really important reviews. But how do we categorize that? Well, you gotta think about the PR as using this AWS terminology, which is, is it a one-way door or is it a two-way door? Now, most PRs that you have will be two-way doors. You can merge the PR and then always pull it back later. It's the glorious benefit of being a software engineer as opposed to a civil engineer, right? Mostly when you're a civil engineer, it's a, it's a
- 17:44
one-way door, right? If you get something wrong, that bridge is going down. But if you have a two-way door, it means that you can easily revert the change. Now, that might be more, um, a bit more nuanced than you might expect. It might be that a very simple change accidentally blasts out an email to 60,000 people or something, in which case that is a one-way door. You wanna review that very, very carefully. Involves expensive migrations or data loss, that is a one-way door. Review the hell out of that PR. But also tied onto
- 18:14
that, we need to understand the blast radius of this PR. What can go wrong? And if things do go wrong, how bad is it? And this means I end up with a nice little sort of summary right at the bottom of all of my PRs, which is the merge danger. I can see this one is a two-way door. Its blast radius is localized, and so I can see, fantastic, I don't need to pay that much attention to this. I'm just gonna sort of review it a little. That's really important. Next, we need to understand what the PR is even doing, right? And I've
- 18:44
tried lots of different ways of figuring this out, and the best thing I've come up with is using pseudocode. Now, a huge, um, point of gratitude here to the Show Me skill from the HumanLayer skills repo by Dex Hawley. This is a phenomenal skill that just essentially dispenses with most text and shows things to you in images and diagrams instead. This makes it a lot easier to grasp actually what's changing and why it's changing. So you get the kind of standard
- 19:14
sort of set of like mermaid diagrams and UML for, you know, this is a sort of sequence of things that happened. You also just get these lovely simple ones like this, for instance. We can look at this and go, "Okay, we're working in a CLI. We can see a new command has been added, and we've got two new things, little flags up the top here." Just little summaries like this. It really does make a massive difference. It's hard to overstate. So you're m- trying to, uh, like make understanding the why as fast as possible. And
- 19:44
I think a third principle here is something you should be thinking about whenever you do human review. Because, because we're not doing like, um... Because our processes now are so sort of streamlined and all-- we're all collaborating around these same skill files, these same steering files. We're all building an environment for our agents to operate in together. You should think of the process that produces your code as just as important as the
- 20:14
code itself. In other words, when you do a human review, you're not just reviewing the code, you're reviewing the system that creates it. And the theory here is that you never want to write the same comment twice, right? You never want to catch the agent doing the same thing over two PRs. And so what's the mechanism by which you can make your human review matter? Well, this is a new skill. This is called Retro. Retro for retrospective. You essentially take a session that you've done. It can either be like a single,
- 20:44
um, agent session, or it can be a PR plus the session, or you can just get it to look at, okay, look at all of the PRs that we've done over the last week, all of the reviews, pull them all in. Let's do a retrospective on them. And it will suggest automated checks and coding standards to make the next one better. So this is the compounding effect, where you essentially by doing human review, you're making the quality of the next human review higher, and you're sort of saving less work from yourself next time. And
- 21:14
Retro is a really smart skill. It adds a bunch of stuff. So obviously it suggests automated checks. It suggests updates to codingstandards.md. It does other smart stuff too. So it looks at navigation pointers. How easily did the agent find its information? Can we provide a pointer inside agents.md to help it out next time? It looks at tool economy. Are there different tools that we're using in the session that can, you know, could be made more token efficient? It's amazing how many things this catches actually, 'cause those are often really
- 21:44
hard to debug from the outside. It just looks at bloat as well. So are there bloated steering files? Are there bloated skills that contribute to these bad results? Can we make them more organized? So that's the goal, is to make human review faster. We do that by layering up automated checks. We're layering up automated review, and we make the human review as painless, as simple, and as kind of optional as we need it to. You really don't need to review every single two-way door. Every
- 22:14
single one-way door you do. So these are my skills. aihero.dev/skills. I'm gonna be shipping version one point three this week. It has been glorious hanging out with you. It's been a really nice conference. I'm gonna be outside in the lobby if anyone wants to have a chat. Uh, thank you so much for having me. Thank you, Paris.