AI Engineer World's Fair 2026
The Z/L Continuum: Should AI Engineers Still Read Code?
About this talk
Alex Volkov, host of ThursdAI, presents the Z/L Continuum, a task-specific framework for deciding how closely engineers should inspect AI-generated code. Contrasting Mario Zechner's emphasis on careful review with Ryan Lopopolo's agent-forward approach, he examines Claude Code, METR capability measurements, rising development throughput, and reliability risks. He argues that traces, evals, shadow mode, observability, rollback, and increasingly autonomous loops should shift oversight toward system-level verification without eliminating human judgment.
Chapters
- 0:00The code-review debate and the Zechner–Lopopolo Continuum
- 0:46Agent capability growth, METR, and Claude Code
- 3:06Contrasting agent-first and review-first engineering
- 6:00Alex Volkov introduces the continuum and reliability tradeoffs
- 15:11Verification, observability, and engineering durable safeguards
- 17:17Capability drift, agent loops, token costs, and human judgment
Talk transcript
- 0:00
[on hold jingle] Two talks at AI Engineer Europe.
- 0:15
One guy is saying code is free and deleted his IDE, and the other one is saying, "Read every effing line of code." So should AI engineers still read code their agents output in twenty twenty-six?
- 0:31
I named this the Zechner-Lopopolo Continuum, and you guys probably have argued about this in Slack. You probably talked about this in the hallway track. So let's talk about this here because code got cheap, attention didn't.
- 0:46
As you may know, back in December twenty twenty-five, something big changed. AI engineering has changed forever, and it broke its own trend line. Actually, Swig, the organizer of AI Engineer, is collecting evidence to that single moment in time at the website called wtfhappened2025.com.
- 1:05
I recommend you go and check it out. It's really, really funny. Uh, this is just one example from METR, the Machine Evaluation Center, and it shows that models, for the first time, started completing tasks that would take engineers over sixteen hours to do.
- 1:21
And in fact, we've gone way up the curve, way up the trend line after that. This is the backdrop to everything that AI engineering is experiencing.
- 1:32
Because we don't write code anymore, most of us at least. I want to see one-- Can you guys, uh, give me a raise of hands if you still handcraft and write code, most of your code?
- 1:42
Anybody here, most of your code is written by hand? Amazing. This is the token maxing track after all. [laughs]
- 1:48
I, I think the one person here who still writes code is maybe a little shy of raising their hand. That's okay 'cause we don't type code anymore. We're not handcrafters.
- 1:55
We supervise. I like to say we babysit agents. And the greatest example for this, obviously, is Boris Cherny. You guys know Boris, the creator of Claude Code, uh, at Anthropic.
- 2:07
A hundred percent of his code is written and authored by Claude Code at this point. And he didn't stop being an engineer. He moved up the layer. He still ships twenty to thirty PRs, maybe more.
- 2:18
And recently, he talked about he deleted his IDE. [laughs] I found it really funny. Just, just no reason to just hand-type code anymore. And in fact, eighty percent of Anthropic's code is now AI-written, and this is-- this stat is at least a few months old.
- 2:32
It's likely more right now. And he's not the only one. Some of you have seen this chart from GitHub. [laughs] Some of you have maybe remembered this chart while GitHub was down for you.
- 2:44
Uh, the reason is GitHub is on track to, to, to get fourteen billion commits this year. All of twenty twenty-five, all of yester-ye-last year was one billion. They're fourteen x-ing the number of commits.
- 2:58
They're seeing fourteen x the number of commits, which is insane, and most of this is AI-assisted, and it's a lot of code.
- 3:06
And so the engineering has changed forever, and I want to tell you about AI Engineer World's Fair. I've been to every single one, and I'll tell you about this later.
- 3:15
And, uh, AI Engineer is a great place to get the zeitgeists of where our career is going and how is it changing. Okay? This one obviously is three x bigger than last year.
- 3:25
This is just one of the rooms. There's like a bunch of rooms. Seven thousand people I think we clocked in, thirty-six tracks. And if you wanna know what happens in AI engineering, you're gonna have to be here.
- 3:36
So this will be a little bit of a meta talk. So one of the guys at AI Engineer EU talked about code is cheap. The other one talked about, uh, we should read every line of code.
- 3:46
Let's listen to them for just a second, okay? This is Ryan Lopopolo from OpenAI. I don't think he's-- he made it here, but, uh, this is Ryan Lopopolo from OpenAI.
- 3:54
The models at this point are good enough where they're isomorphic to human eyes.
- 3:58
Can you guys-
- 4:00
Use code at high quality that solve real user problems in real code bases.
- 4:07
Code is free. It's free pro-to produce, free to refactor, and it is not a thing to get hung up on anymore. Humans no longer need to concern themselves with implementation.
- 4:20
The important thing is not the code, but the prompt and the guardrails that got you there. You can just simply say, "Do not produce slop." Don't accept slop. You won't get slop in your code base.
- 4:30
But to do that requires taking short-term velocity hits in order to back up or double-click into a task to figure out what it is the agents are struggling with.
- 4:42
So this is Ryan Lopopolo, okay? He came up on stage at the AI Engineer, and he opened with like, "Hey, I'm a token billionaire, and I want you to be as well."
- 4:50
In fact, the token billionaire lounge that's in front of the leadership track that you guys see, that's because of him. He came up with this concept, uh, and he got the golden card and everything.
- 4:57
Uh, on the other side, the same conference, the other side, Mario Zechner, creator of Pi.
- 5:02
Slow the fuck down. [laughs] Everything's broken. And then there's people that say, "Our product's been a hundred percent built by agents." Yes, we know. It fucking sucks now. Congratulations. [laughs] [audience applauds]
- 5:21
Our agents are actually compounding booboos, which is my word for errors, with zero learning and no bottlenecks and, uh, delayed pain. The delayed pain is for you. Those are my most beloved people.
- 5:32
I don't even read the code anymore. Congratulations. Something is broken, and your users are screaming. So who you gonna call? Not yourself, because you haven't read the code. Non-critical code?
- 5:42
Sure. Write slop ahead. Critical code? Read every fucking line.
- 5:46
So two folks, same conference, day after day, talking about the one anxiety that we all feel. Should we all still be reading code in 2026? By the way, these two folks are the number six and number seven most watched YouTube videos from AI Engineer from all time.
- 6:00
So they're obviously representing something that we're feeling, we're talking about, and this is being the leadership track, something that folks that report to you are talking about, okay? Should they still be reading code, and what's the, what's the level of quality?
- 6:11
So they named the same anxiety from both ends. Uh, at this point, I probably should introduce myself. Uh, hi, I'm Alex Volkov. I'm the host of ThursdAI Podcast. It's a podcast and a newsletter.
- 6:20
We go live every week to talk about AI. For the past three and a half years, we've been tracking every change in AI engineering, every release from every lab, every model, and I'm also an AI evangelist with Weights & Biases and CoreWeave.
- 6:33
Um, what also should I tell you about myself? That I've been covering AI engineers specifically since the first one in 2023, and oh boy, has it changed. And so you can treat this as a dispatch from the front line because all of these people now are my friends, and we constantly talk about this in the speakers room,
- 6:47
in the hallway track. I couldn't stop thinking about that tension. I couldn't stop thinking about that kind of disparity between the two folks, okay? And I put them both on the line, Zechner from one end, Lopopolo on the other end.
- 7:02
I called it a continuum, and I basically started asking people, "Hey, where are you on this line? Where... Are you a Zechner? Do you still read every line of code?
- 7:11
Are you a Lopopolo? Do you just YOLO and don't even look at code and think agents are good enough, et cetera?" And I [laughs] I, I got the framing wrong.
- 7:20
But I'll tell you about this in just a second, okay? So before this, I want you to be ki- honest with yourself. And again, if you don't write code, or let me say this, if you don't babysit your own agents, but you, you have reports that babysit agents for you, uh, think about them when you answer this,
- 7:33
okay? And be honest. On the Z/L Continuum,
- 7:37
where are you? And let's take, um, let's take by vote of hands, who here has committed code that they've never looked at before?
- 7:48
Amazing. Love that. Uh, who here still reads every line of code, of at least critical code? I see one cowboy over there. I love that, man. I'm gonna talk to you afterwards, okay?
- 7:57
I wanna understand exactly why you do this. Um,
- 8:02
and so who's right? Let's talk about who's right. Let's talk about where we are right now, and we start with Ryan Lopopolo. [laughs] If you get to meet Ryan over here, he is very AGI-pilled.
- 8:11
I think even within OpenAI, the AGI organization, Ryan is kind of like the more AGI-pilled person. Uh, if you had the chance to go downstairs and grab the AGI pills that Sue ex- prescripted, I think Ryan had all of them. [laughs]
- 8:24
Uh, he works at OpenAI, where he says code is literally free.
- 8:28
So are tokens. Uh, it's really fun. We renamed Ryan, uh... Do you guys know the dash, dash YOLO in Codex? It's kinda like the skip dangerous permissions in Claude Code.
- 8:36
So we renamed Ryan Lopopolo, [REDACTED:username]. He's okay with it, by the way, I asked him. So if we check his kinda side, the, the folks like him against the data, they're actually right.
- 8:48
The optimists are right, at least about output. This is from Ferrous AI. I think I'm not the only speaker at this conference who cites this essay. It's, uh, sorry, this survey.
- 8:57
It's new from April 2026. I think it's, uh, one of the best kinda evidence of where we're going that we can now cite, okay? Uh, 22,000 engineers were surveyed about code.
- 9:07
They call this the acceleration whiplash. And they're talking about, my favorite stat on here, and you can read this yourself, 861% increase in code deletion per PR. So us, together with, with AI agents, we love deleting code. [laughs]
- 9:22
Uh, Anthropic also said that they are shipping eight times more code per quarter than in 2025.
- 9:30
But is it all good code, okay? Let's play a game, and if you know the answer, you... Let me have my moment on here on stage, okay? But if you don't know the answer, let's guess.
- 9:41
Whose status page is this? Claude. I think [laughs] I, I hear a few answers. I think most of us guessed it. This is Claude. Uh, in fact, as you can see on the o- on the right, it was down when they took the screenshot.
- 9:55
It was really funny. Uh, uh, this-- Anthropic is the company that probably uses the most AI-generated code, and their status page looks like a Christmas tree. Now, I'm not here to dunk on Anthropic.
- 10:06
Farik just did an incredible job back on stage, uh, talking about Claude and et cetera. Um, this may be due to scale. This may be due to the other factors.
- 10:13
Uh, I'm not here to dunk on them, but it just goes to show that they're not the only ones like this. Obviously, GitHub famously also suffers from a little bit of both.
- 10:20
Um, output does not mean stability, okay? So maybe this is a good example of what? Same essay, 31% increase in PRs merged with no review at all, human or agentic.
- 10:34
Don't do this. I beg of you. Don't. [laughs] It's, it... We'll talk about how to fix this in a second. So when you ship this fast and this much, something gives, and usually it's quality.
- 10:44
So maybe Mario's right, yeah? Maybe the bill does come due in production. Same study, 242% increase in incident per PR.
- 10:54
This is kind of scary. The second stat is also scary. Bugs per developer is up six times than 2025.
- 11:02
So even Anthropic concedes this. I don't know if you guys read the RSI essay they posted, the recursive self-improvement, where they talked about, "Hey, what does the future hold?"
- 11:14
They outlined two scenarios. One of them says, "Maybe the acceleration will stop, and we're gonna get used to this." They a- they say that's actually not likely to happen.
- 11:24
We just added this, uh, eventuality for clarity. We don't think that's likely to happen. What we think is going to happen is, uh, engineers and companies 10X-ing to 100X-ing to 1000X-ing their output and productivity, and then they say this.
- 11:38
"We, as we began to push more code around the organization, human code review has become a new bottleneck." They're citing Amdahl's law that shows that if you have an explosion of productivity in one area, another area g- gonna gets blocked.
- 11:53
And- Nobody removes the human in these organizations. In fact, careers in Anthropic and careers in OpenAI, they're still hiring humans. So nobody's removing the human, and they're both saying that human code review is still a concern.
- 12:08
And here's my mea culpa. I promise you I'll tell you where I got it wrong, the framing. My mea culpa is the continuum is real. The Z/L continuum is real, but it's not about the people.
- 12:17
It's about the tasks. The continuum is real. It's not about the people. It's about the task. Same engineer could be a Ryan Lopopolo on one piece of code and has to be Mario Zechner and read every line of other pieces of code.
- 12:35
Different tasks just need different proof. If we look at them closely, I obviously character- characterize them. [laughs] I've practiced this word multiple times, and I still get it wrong. Characterize them.
- 12:47
They're a character, uh, on both ends for the Z/L continuum. Uh, but if you look at them closely, what they're saying closely, they're actually not that different. Ryan's mechanism is moving attention up the layer.
- 12:58
He's saying humans are unreliable at catching repeated mistakes of the same time, repeatedly catching the mistakes of the same time. So when they... you do catch a mistake during the PR review, write the documentation, the linter, and the reviewer needs to remember this once, so the system will catch this type of bugs.
- 13:15
He's not saying don't inspect your code. He's saying inspect the system, not every line.
- 13:20
Mario, from the other end, is saying route by task. If it's not critical, let it rip. He said it. [laughs]
- 13:27
And if it's critical, you read every fucking line. How do you know what's critical? Well, his answer is easy. You read the effing code. Uh, my answer to add to this is also you ask your clankers.
- 13:38
They're great at looking at a large repository and telling you, and telling you, "Hey, this line is actually critical. [laughs] You should look at this area. These primitives over here are critical."
- 13:47
So you ask your clanker. So they agree more than I kind of gave him credit for.
- 13:53
And so I think at the beginning of this, the wrong question is should I still be reading code in 2026? I think the better question right now for all of us is what proof does this specific change need?
- 14:06
What proof does this specific change need? And so I took Mario on the left, obviously. I took Ryan on the right, and then I took a bunch of other, uh, great AI engineers, friends, some friends, uh, many of them speakers at this conference, and kind of distilled their advice down to a routing table.
- 14:22
And, uh, they told me, I think Swix told me on, on, on Twitter, there's gonna be one slide that I will... that people need to take a screenshot of.
- 14:29
It's gonna be this slide. You don't have to read it with me, but at the end, you're welcome to take a, a, a picture of this. This is the, your Monday artifact.
- 14:36
Routing the change where the proof needs it. Routing the change to the proof that it needs. You read every line of authentication, money movement, permissions, and irreversible data. You inspect the critical paths yourself, and then obviously you keep going.
- 14:50
Uh, decomposing I think is very important. The more code is getting written, the more it's hard. Your eyes are starting to glaze over a, a very long p- pull request.
- 14:59
So splitting into atomic reviewable PRs. You know who's good at it? Agents. They're great at decomposing code. Ask them to do it. You verify. That doesn't go away. This has been with us in the, in engineering, software engineering, and AI engineering.
- 15:11
It doesn't go away. Traces, evals, shadow mode. Come talk to me after, [laughs] after this talk. I don't have enough time, but shadow mode is a really cool one that I learned about preparing this talk.
- 15:19
And then I think the most important one is separating. Many people have the same agent that writes the code, also inspects the outputs, and writes the test. Separating is very important.
- 15:28
If you don't separate, it's kind of like if I came up with an exam, and then I took an exam, and then I scored myself on the exam. It's not, not really productive, right?
- 15:36
And then last one is engineer. Rails, observability, rollback. This is what Ryan Lopopolo talks about. Build the system that builds the system because read spends your attention once. Engineer makes the system remember,
- 15:50
right? And you might sitting-- m-might be sitting there and saying, "Hey, [laughs] did you hear the news, Alex? Fable is back. What about Fable? What about Mythos? Is this still relevant at this next scale of capability?"
- 16:02
Because when I coined the Z/L Continuum, it was only eighty-two days ago. Mythos has just been announced. We weren't sure, like, what's going on. Only the people in Anthropic got access to it.
- 16:11
And Derek, uh, [REDACTED:username], that was on stage from Anthropic, he said about Mythos and s- and Fable, "We used to check if Claude is doing the work right, and with Fable 5, I instead check if Claude is doing the right work."
- 16:29
Let it land for a second. I don't know if you read the sen- se-statement. When I read the statement, I felt like little chills at the back of my neck about the next, like, level of capability, okay?
- 16:38
We used to check if Claude is doing the work right. With Fable, we check if Claude is doing the right work.
- 16:44
And our favorite senpai, who recently joined Anthropic and is getting unnecessary heat on Twitter, uh, said this. Uh, Andrej Karpathy has said, "It's never felt so tempting to stop looking at code at all.
- 16:56
But don't do this in production." Senpai is great for the sole reason... Do you guys know the sentence, uh, this meeting could have been an email? So [laughs] this presentation could have been Andrej Karpathy's one sentence, okay?
- 17:08
He's naming the anxiety from both ends. It's never been so tempting to stop looking at code. Don't do this in production, even with Fable.
- 17:17
And so if you guys noticed, uh, I have a little thingy here.
- 17:22
This. Uh, it's so white you can see my little, uh, laser pointer. Do you guys see the arrow, the capability drift arrow? This thing? [laughs] When I wrote the continuum, I realized that it, it's only a temporary place and time.
- 17:37
Capability increases move us towards Lopopolo. So we're gonna talk about capability increases as well because the review layer moves. If yesterday we inspected the outputs and this, we read the code, uh, and today we inspect the task direction and kind of like direct it to the right proof, maybe tomorrow we're inspecting the loops.
- 17:56
Capability drift changes where proof belongs. It doesn't remove the requirement of proof.
- 18:02
Talking about loops, is that the next primitive? I think most of this conference, I think the zeitgeist for this one is going to be, is token factories and co-factor is real, and is loops is a real thing that I need to be doing at this point.
- 18:14
By raising of hands, who here heard of loops? Keep your hands up please, and take them down if you are not running loops right now and you have no idea what they are.
- 18:24
There's a good perception... There's a good, uh, number of people here who heard about loops, and they started with both these folks, Peter Steinberger, creator of Open Cloud, uh, Open Claw, and now is OpenAI, and, uh, Boris Cherny.
- 18:36
And pretty much within the span of two days, both of them started talking about loops that became kind of the zeitgeist. And loops are moving us from prompting each turn to designing the system that writes the actual prompts.
- 18:48
By the way, do you guys know what's common between these guys and what's different between me and these guys? Their tokens are free. So when they talk about loops and their tokens are free, uh, they're not telling you, "Hey, you should be doing this right now specifically."
- 19:02
But because they work at bigger labs, you can treat them as kind of a lighthouse that's pointing where we're all going, kind of like, uh, Gretzky's skate where the puck is going to be.
- 19:10
They're gonna tell us what, uh, all of our enterprises are gonna get, get up on. And if it's, if it's loops, then let me g- at least give you a TLDR, okay?
- 19:19
Loops are basically fancy cron jobs that run on a schedule, but what they do is they discover a task and kind of start writing a prompt for this task from the plan.
- 19:27
They run the plan, they execute, and most importantly for my talk here, they verify themselves, and if it doesn't work, they try again. So an agent that loops grades its own work against a goal with less human intervention.
- 19:40
But if the builder grades itself, you didn't remove the review, you hid it. Okay? This, this connects to my routing table. This comes from Adi Osmani recently at Google.
- 19:50
He's also at this conference, a great engineer. Uh, he said, "If, if I wro- if I wasn't reviewing the code myself or relied entirely on automated loops to fix my code, let's say a bug comes up in Jira and my loop picks it up and starts fixing this, my product quality would suffer.
- 20:05
I'd likely end up in a downward spiral, digging myself into a deeper hole." So again, loops don't remove judgment, but they do raise the stakes on where you put it.
- 20:15
So what about the future, folks? Nobody fucking knows. Anthropic did not know that Claude Code is gonna explode on them and this is gonna be a billion-dollar product. Uh, nobody knew the coding agents and Harness are gonna be the generalized agent, and now everybody's pursuing them, folks at OpenAI with Codex, Elon with, with Grok Code, uh, Google
- 20:37
with, with Antigravity. Model capability is jumping at an insane pace, and what I implore and to tell you here is that flexibility is required. You need to keep... be nimble to keep up with, with the trends.
- 20:50
This is why you're the engineer. And by the way, I told some folks here about my podcast, ThursdAI News. If you want to keep tracking where that line moves, feel free to scan this QR code, join a, uh, a, you know, newsletter, et cetera.
- 21:02
Um, and I'll leave you with this, because it's my time, I'll leave you with this. Not every line in 2026 needs your eyes.
- 21:11
Every system still needs your judgment. Thank you. [audience applauding] [outro music]