AI Engineer World's Fair 2026
Generation Is Cheap, Review Is Expensive: How to Stop Shipping AI Slop — Gabriel Martinez, G2i
Read the talk
Generation Is Cheap, Review Is Expensive: How to Stop Shipping AI Slop
Gabriel Martinez explains how generated work can hide unresolved decisions, shift effort onto reviewers and reward the wrong kind of speed—and how small changes, visible workflows and conventions help teams keep ownership.
From a talk by Gabriel Martinez
At a glance
Ideas worth remembering
Slop appears when polished output conceals unresolved choices, leaving engineers unable to recognize and own the decisions entering their system.
A prototype demonstrates visible behavior; integration, missing requirements, human testing and long-term maintenance still require judgment.
Generating a larger artifact without resolving ambiguity transfers thinking to reviewers and can make evaluation the team's bottleneck.
Small changes, visible workflows, conventions and clear ownership help humans review systems even as agents take on more implementation detail.
ORC is presented as a workflow for preserving planning, review and human responsibility, with coherent and maintainable software as the intended result.
When work looks finished before the thinking is
A feature can look complete while its important decisions remain unmade. Gabriel Martinez—Gabo, an engineering manager and self-described “resident slop warden” at G2i—calls this slop: “output that looks finished before the thinking is finished.” The problem predates AI. Generation has made it happen faster.
The failure begins when ambiguity gets resolved through guesses. A generated implementation chooses a behavior, an abstraction or a tradeoff, and its polished appearance conceals the fact that a decision was necessary. The engineer may never encounter the constraint or edge case that would have changed the design. A bad choice is visible enough to challenge; a choice nobody noticed can pass straight into the system.
Agents can write code, scaffold screens and, with suitable tuning, build features or systems. Those capabilities leave a separate set of responsibilities with the team: decide what matters, interpret ambiguous requirements, understand the resulting system and own it over time. Architectural restraint belongs on that list too. Knowing what to leave unbuilt matters more when generating another feature becomes easy.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Cheap code still creates a maintenance bill
Creating code and understanding a program have different costs. Generating more files lowers the first cost without necessarily lowering the second. Every addition becomes something the team may have to explain, change, test or maintain. Martinez invokes Dijkstra’s distinction between lines produced and “lines spent”: generation makes it easier to spend on abstractions and complexity, while ownership keeps the bill with the team.
Understanding does not require one person to read and bless every line. It requires a grasp of the main flows, system boundaries, data model and product behavior. Code should reflect the domain closely enough that these relationships remain understandable. Existing structure, patterns and rationale also give an agent a clearer path to follow. In a “big ball of mud,” adding code can simply make the ball bigger faster.
The cost eventually appears in the product. Brittle flows become difficult to change; features become places nobody wants to touch; screens accumulate escape hatches until ordinary tasks require training. The useful question is whether users needed that complexity or whether teams kept adding things because they could. Accountability makes someone responsible for answering it. Martinez proposes visible ownership and review of each chunk of work so that producing an artifact carries consequences beyond getting it merged.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A clickable prototype changes expectations before it proves readiness
A few prompts can produce screens and a working clickable prototype. That can be useful for exploring an idea. It also creates an expectation problem: stakeholders who see software appear quickly may conclude that the hard part of building it is over. What they can click is much more visible than the unfinished work outside the happy path.
Turning that prototype into part of an existing product requires several kinds of judgment:
- Integration: Fit the new feature into the system coherently, rather than merely making it work in isolation.
- Requirements: Discover missing requirements and settle the choices the prototype bypassed.
- Human behavior: Manually test whether the feature works as a person would expect and whether the UX makes sense.
- Acceptance: Read the code, own its tradeoffs and decide whether the addition deserves a place in the product.
A contribution also brings future obligations. Martinez recounts a quotation attributed to SQLite’s creator in which a supposedly free pull request asks its recipient to maintain, document and test the addition for the next twenty-five years. The point is the obligation transferred with the code. It is easy to “yeet a PR”; someone still has to own what accepting it means.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Merge counts can reward the people who skip judgment
If visible output becomes the measure of performance, careful engineers can appear slower than colleagues who skip specification work and review. That comparison creates pressure to merge faster than anyone can judge the changes. Martinez describes feeling this anxiety himself: watching others merge at an incredible pace led him to accept code with less scrutiny than he normally would.
The incentive changes what the team optimizes. Getting a PR in today can beat protecting code health over time. Martinez’s alternative is to reward delivered value, pragmatic and well-factored design, and clear test and evaluation harnesses. Those make the quality of a change easier to assess than a raw count of merged artifacts.
Change size matters because it determines how much a reviewer must understand at once. A multi-thousand-line PR can overwhelm that capacity. Breaking a need into simpler chunks gives reviewers smaller pieces of behavior and design to reason about, helping the team retain control of the codebase as generation accelerates.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Ten bullets become twenty pages—and the thinking moves to the reviewer
The review bottleneck becomes concrete in a document example. Someone receives ten bullets of scope and a few user journeys: a deliberately lightweight starting point. They return a twenty-page document with polished prose. The observable change is a much larger artifact. Martinez then spends hours trying to determine whether it contains real thinking.
Here is the causal chain: the scope contains ambiguity; the person doing the work leaves that ambiguity unresolved; generation expands it into a document that appears complete; other stakeholders must now read the expansion and settle the original questions. The document has moved work downstream. A vague coding task can undergo the same transformation into a large PR, or a small feature into a pile of files.
Where did the work go? The flow below follows the document from its small input to its expensive review. The important relationship is the continued presence of unresolved decisions: expanding the artifact does not resolve them, and the reviewer inherits both those decisions and the reading burden.
“Generation is cheap, but evaluation is expensive.” When output exceeds what humans can evaluate, more generation creates a bottleneck rather than team acceleration. Counting tokens burned repeats the mistake of counting lines of code as productivity. Martinez’s irritation with agentic hype—including his description of the Ralph loop as a while loop—returns to the same practical point: familiar software practices still apply. Small, composable pieces remain easier to understand and review than a whole feature requested in two sentences.
Ten bullets and a few user journeys.
The twenty-page document adds reading effort while leaving stakeholders to resolve the ambiguity in the original scope.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Keep the workflow visible as implementation moves out of view
If agents write the code, why care about its shape? Martinez takes the objection seriously. Software has long moved implementation details behind abstractions, from machine code and assembly to higher-level languages. AI continues that movement by taking on imperative implementation work. The remaining requirement is to understand how the system fits together.
Diagrams provide a compressed way to review those relationships. Useful diagrams expose service boundaries, data flows, state transitions and failure modes. These are the things a team must judge even when an agent generates thousands of lines underneath them. The main process or workflow can stay visible, with implementation details pushed into its leaves.
This explains the preference for declarative code and architecture with understandable layers. They let a human inspect what the system does and how its parts relate without having to reconstruct the whole design from implementation detail. The purpose of this structure is human knowledge and review. A diagram added after the fact as decoration would miss that purpose.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Conventions reduce the decisions a reviewer must reconstruct
Before making a meaningful change, an engineer has to discover where code belongs, which pattern to follow and what a proper implementation looks like. That is cognitive overhead. An agent can increase it by producing five plausible patterns in five different styles. The reviewer must then reverse-engineer both the code and the intent behind each choice.
Several practices address different parts of this burden:
- Conventions and standards: Make the expected implementation path recognizable.
- Documented flows: Explain how the system behaves so reviewers can compare a change with that behavior.
- Small reviewable changes: Limit how much new material someone must hold in mind at once.
- Ownership and definitions of done: Identify who is responsible and what must be established before the work is accepted.
Rails supplies the example Martinez cares about most. Its conventions narrow the number of recurring structural decisions. A shared approach may not be everyone’s personal favorite, but people understand it and can work within it. After a four-year gap between working with Rails projects, he returned to familiar patterns and de facto libraries. He could understand intent without first decoding a new architecture.
The tradeoff is flexibility for familiarity: fewer competing approaches mean fewer decisions to revisit and less ramp-up between projects. Martinez connects this to Rails’ convention-over-configuration and programmer-happiness principles, and extends the reasoning to agents. A more obvious implementation path gives them fewer competing approaches to choose among. His claim that such environments are superior for agentic development is an engineering judgment illustrated by this experience, rather than a measured comparison in the recording.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
ORC puts the surrounding discipline into the workflow
ORC grew out of a repeated failure pattern at G2i: AI could generate code, while planning, review, adversarial checks and context management still needed deliberate attention. The team kept repeating the same processes—break down the work, manage context, weigh options and make sure additions entered the codebase in a form people could still reason about.
Martinez presents ORC as a way to build those lessons into the workflow while agents do more of the implementation. Correctness remains a human responsibility. The recording explains this purpose without detailing the product’s internal implementation, so the useful lesson here is the responsibility it aims to preserve: after generation, the software must still be coherent, tested, maintainable and owned.
That ending gives the title its practical force. The objective is software a team can understand and stand behind after the code arrives. More generated code is useful only when the surrounding workflow can turn it into that result. Martinez’s name for the future is “orchestration with judgment.”
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
Related talks
- Beating RL With Reflection: GEPA and Optimize Anything
Explores how execution traces and detailed human feedback can improve agents and evaluators—a concrete companion to this talk's concern with the cost and quality of evaluation.
Read the complete timestamped transcript
- 0:12
Hi, everyone. My name is Gabriel Martinez. I go by Gabo. I'm an engineering manager and resident slop warden at G2i, and I'm here to talk about slop. I'm here to talk about slop because I'm rather passionate about the topic, and by passionate, I mean angry. The term slop has been beaten to a pulp, but the thing it points at has long been around, only now it happens at warp speed.
- 0:39
All right. Slop is what happens when judgment-- when generation outruns judgment. Slop is not AI-generated code. I don't mean to define it as some concrete, objective thing that can easily be sniffed out, although there are tells. Slop is output that looks finished before the thinking is finished. It is work product where ambigu- ambiguity was cleared with guesses, where trade-offs were left to chance, where ownership was not taken. Output was
- 1:09
produced, sure, but understanding was not, and the real danger is not simply that the machine made a bad choice. The danger is that the engineer never noticed that there was a choice to make. They never hit the ambiguity. They never felt the constraint. They never discovered the edge case that would have changed the design. The work arrived already smoothed over, and with it, the unknown unknowns stayed unknown. But software is a record of choices, technical choices, product
- 1:39
choices, operational choices, user-facing choices. If we do not see those choices, we cannot own them, and if we do not own them, we are not really engineering. We are just accepting output. We are just a code monkey hitting the merge button. I'm not the purist that's come here to te- to tell you that AI is bad. It's not bad, okay? I'm here to describe some of the patterns that have led to our drunken AI stupor. It's absolutely understandable, given the advancements
- 2:09
made since December '25, how we got to where we are. But we need to dial it in. We need to slow down. We need to sober up. Agents can seem to make magic, and the internet would make you believe it is so. They can write code, scaffold screens, and with right tuning, can build whole features or systems. But they cannot replace the human work of deciding what matters, understanding what is being built, interpreting ambiguity, owning the system for the long haul, exercising judgment and
- 2:39
restraint at the system architecture level, and arguably the most important, important, knowing what not to build. Slop is not about AI as a tool. Rather, it is about the absence of judgment. So let's talk a little bit about what got us here.
- 2:56
In a world where anything is possible, deciding what to build is more important than ever. People have gotten overly excited and built behemoths of spaghetti product nonsense. It's almost as if they've forgotten that any line of code you write is a liability, that the best code is no code at all. The cost of creating code is falling, but the cost of understanding a software program is not. And I wanna be clear here. I'm not saying every line of code needs to be personally read, blessed, and understood in isolation. That has never really been the
- 3:26
standard. Large systems have always contained more detail than any one person can hold in their head, but the logic has to be understood. The main flows, the boundaries, the data model, and product behaviors have to be understood. Good software has a kind of harmony to it, not in some precious aesthetic sense, in a practical sense. The code should reflect the shape of the domain. AI can generate code, but it cannot magically give a messy system coherence. An agent can operate with much more
- 3:55
effectively inside a code base that has structure, patterns, clear boundaries, and rationale. But if the system is already a big ball of mud, the agent does not rescue you from the mud. It just builds the ball bigger, faster. Just because you can generate another component does not mean the product needs another component. Dijkstra had this great line: "If we could count lines of code, we would not think of them as lines produced, but as lines spent." It is the cost of building. AI makes it easier to spend lines. It makes it easier to
- 4:25
spend on files, abstractions, and complexity. But at the end of the day, you own the bill. That bill shows up as confusion, brittle flows, features nobody wants to touch because they don't fully understand. It shows up as product surfaces that technically do everything, but feel like a maze. Think about that special kind of b- bloated B2B software where every task requires training, where every screen has 12 escape hatches, and nobody can explain whether the complexity is there because users needed it or
- 4:55
because teams just kept adding things. Teams should be h- held accountable for work they are producing. Accountability is the key word here, taking ownership over the product you're building and having a way to tie that to positive or negative reinforcement. Just in the mere fact of knowing that someone else will review your work will force engineers to pay just a little bit more attention to the finer details and be a bit more meticulous, especially if there are accountability metrics that can be made visible for each chunk of work product. We're entering the age of bloated,
- 5:25
bloated products with spaghetti UX on the order of worst kinds of B2B ERPs you may have ever been trained to use. Today, anyone can vibe code an application in an hour and end up with a working clickable prototype, and that has changed expectations. Because once people see software appear that quickly, they start to believe software can be built that quickly. They see a few prompts turn into screens. They see a terminal session produce something that looks l- real. And to be fair, sometimes it is real enough to be
- 5:55
useful. Sometimes that prototype is a great way to explore an idea. But the problem is that stakeholders do not live in the code base of-- that stakeholders, uh, do not live in the code base often experience that prototype as evidence that the hard part is over. They're not thinking about the parts where real battle-hardened engineers earn their keep. It's outside the happy path. Figuring out cohe- how to cohesively integrate a new feature set, the w- the work involved in clarifying requirements, the work of actually manually
- 6:25
testing a feature set to ensure that it works as a human should expect, and that the UX isn't complete garbage. They're not at all thinking about the most time-consuming bits, the human parts, the last mile where the judgment lives. This is where someone has to read the code, test the behavior, own the trade-offs, notice the missing requirements, and decide whether this thing deserves to be a part of the system. Here's a tweet that emphasizes the point I'm trying to make here. It quotes the creator of SQLite, who on the
- 6:55
topic of pull requests says, "You say, 'Oh, it's free.' No, it's not free. What you're doing is asking me to maintain it for you, to document it for you, to test it for you, to maintain it for you for the next twenty-five years. That is not free." And this really struck a chord with me. I couldn't have said it better. It's easy to yeet a PR, but who has to own that thing? Who's responsible for the repercussions of what that means? So external s- pressures from stakeholders, coupled with pressures of competing with
- 7:24
slopmeisters, begets mucho slopo. I'm sure most of you have seen this in the wild, where the bad engineers, the ones who don't review the details, that don't refine the spec, that aren't detail-oriented, end up producing more faster. The folks who pay attention to the detail and go above and beyond are the ones that might seem like they're doing less and moving slower. How this propagates is it yields performance anxiety, where in team environments can pressure engineers to merge PRs faster than they can
- 7:54
judge, leading to code being merged without full understanding. Short-term thinking, getting that PR in, wins over long-term code health. I've personally felt this anxiety where engineers around me were merging at an incredible pace, and I felt constant anxiety that I was not moving fast enough. And I can admit that it led me to merge code that I didn't necessarily review with a level of detail that I normally do or should have. I've seen horror stories of people merging code they don't fully understand or haven't even tested themselves.
- 8:25
At the end of the day, we're all human. We are driven by incentive, and if incentives are driving people to get PRs in, then we're gonna be in for a rude awakening. Rewards should be centered around producing value, elegant, well-factored, pragmatic system design, and around clarity of test and evaluation harnesses. There's just something that really grinds my gears about seeing a multi-thousand line PR. There's no way in hell that a human can fully understand all the important bits of a change
- 8:55
that big. It is our job as engineers to optimize for this new world, to decompose needs into its simplest chunks, which lends to understanding, understanding simpler code surfaces and controlling of our code bases.
- 9:12
AI h- is supposed to have been eradicating jobs, and engineers are supposed to be moving at the rate of token production, so there's a lot of pressure to be in ch-- producing more faster. What AI is really doing is producing more faster, more characters that eat more human time. I've seen this happen in the most mundane ways. You give someone ten bullets of scope, a few user journeys, something intentionally lightweight, and they come back with a twenty-page document. Maybe it looks polished with
- 9:42
fancy prose, but now I have to spend hours figuring out whether there's any real thinking inside it. That is AI-generated work in the worst sense. It does not clarify the problem. It expanded it. It took ambiguity that should have been resolved by the person doing the work and handed it back as a larger artifact for other stakeholders to do it. In code, a vague task becomes a large PR. A small feature becomes a pile of files. The output arrives with the shape of
- 10:12
completion, but the real work has just moved to the reviewer, and that is the part that we need to be honest about. Generation Is Cheap, but evaluation is expensive. If AI produces more than humans can possibly evaluate, it has not accelerated the team. It has created a bottleneck.
- 10:32
One thing that I find particularly annoying is Twitter. You've got all these self-proclaimed experts and clickbaiters summarizing some talk or promoting some truths that we've always known as if they're some novel genius idea. The Ralph loop was one of those. It was made to sound like a novel agentic breakthrough when it's just a while loop. Somehow, we've completely forgotten that lines of code is a garbage metric for productivity, but then we decided that it's different than deciding the engineer who burns the most tokens
- 11:02
must be producing the most. We used to break a feature into small atomic pieces we could reason about, small enough to hold in your head and actually review. Now we type two sentences and expect a whole not fully thought out feature set back in one shot without really understanding it i-- what it is we're asking for. The key point I wanna highlight here is that we do not forget the ground truths we've always known. Sixty years of software wisdom hasn't been invalidated by AI.
- 11:32
Composable code, good patterns, and solid practices remain essential.
- 11:40
A friend recently challenged me on why I spend so much time thinking about guardrails for agents, coding conventions, architectural styles, and, and constraints around how the system should be shaped. His question was basically, "What's the point? If the agent writes the code, why should we care about what the code looks like?" And to an extent, I agree. There will eventually be a world where much implementation detail is completely abstracted away from humans. That has been the soft-- the direction of
- 12:10
software from the beginning. We went from machine code to assembly, which another level of abstraction up got us to C, then Java, then languages like Ruby, where the primary design goal was programmer happiness. Just like in languages, we built abstractions for garbage collection and manual pointers, AI is abstracting away the imperative bits. It does not remove the need to understand how a system comes together cohesively. This is why diagrams matter so much. Not as decorative documentation after the fact,
- 12:40
but as a way of making systems legible. When a-agents can generate thousands of lines of implementation detail, humans need a compressed review surface, service boundaries, data flows, state transitions, and failure modes. A good diagram lets us see the shape of the system in a way code often obscures. We can leave the main process or workflow as the reviewable surface with implementation details pushed to the leaves. That is why I vouch for declarative code and architecture patterns that create
- 13:09
layers humans can actually reason about. The goal is not just code generation. The goal is human knowledge, human understanding, and human review. We need to optimize for that human layer.
- 13:23
Cognitive overhead is a tax every engineer has to pay before they can make a meaningful change. It is the time spent figuring out where code belongs, which pattern to follow, and what proper implementation actually means. AI does not remove this tax. Unless tamed, it exacerbates it. An agent can generate five plausible patterns in five different styles, and now the human reviewer has to reverse engineer not just the code, but the intent behind it. With
- 13:53
proper conventions, intent is more easily wrapped. The answer to this is conventions, standards, documented flows, small reviewable surfaces, and clear ownership boundaries and definitions of done. The more obvious the path, the less room there is for slop to hide.
- 14:15
I wanna give special recognition here to Rails and its place in my heart. Rails matters here because it shows what happens when a framework aggressively reduces decision surface, when conventions reduce your overhead. Rather than ending up with ten different ways to structure the same conte-cept, you end up with one. It may not be perfect in everyone's subjective opinion, but it's understood and it works. In my career, between my first and second forays with Rails, I had a gap of four years.
- 14:45
It was beautiful to come back after that gap to then see that the same patterns and de facto libraries were still the same patterns and de facto libraries. I could hop right back into a code base and immediately understand intent without needing to parse through architecture. It allows engineers to easily move between projects laterally with minimal ramp-up. The same can't be said about other, uh, language ecosystems. JavaScript. Environments with a single correct way
- 15:15
of doing things are superior for agentic development because agents aren't confused by dozens of competing approaches. This is well supported by the official Rails doctrine. It explicitly names optimize for programmer happiness and convention over configuration as core pillars, and describes convention as a way to eliminate recurring decisions and lowering the barrier to getting started. It is the reduction of cognitive overhead.
- 15:44
Everything I've talked about today is what led us to build ORC. Not only because we think the ans-- not only because we think the answer is more agents, and definitely not because we think engineering judgment can be automated away. We built ORC because we kept seeing the same failure pattern. AI could generate code, but the surrounding discipline was still missing. The planning, the review, the adversarial checks, the context management. Over time, we found ourselves repeating the same human processes to fight slop. Break
- 16:14
down the work, manage context, weigh out your options, and make sure the output ge- enters the code base in the way the team c-can still reason about. ORC is taking those lessons and baking them into the workflow. It is not a magic box that makes correctness disappear as a human responsibility. It is a system for preserving that responsibility while letting agents do more of the lift. The goal is not to ship more generated code. The goal is to ship software that is still coherent, tested,
- 16:44
maintainable, and owned after the generation happens because the future is not vibe coding better, faster. The future is orchestration with judgment. Thank you all for your time, and remember, don't get sloppy.