AI Engineer World's Fair 2026
AI-Generated Code Is Already Competing With Human Code — Daksh Gupta, Greptile
Read the talk
AI-Generated Code Is Already Competing With Human Code
Daksh Gupta examines Greptile’s enterprise pull-request data, finds similar quality signals for human and agent code, and explains why growing code volume changes what validation must accomplish.
From a talk by Daksh Gupta
At a glance
Ideas worth remembering
Agent authorship needs multiple signals. Author fields alone identified fewer than 1% of PRs; adding co-author footers and branch prefixes raised the estimate to roughly a quarter, with full autonomy still inferred.
Greptile’s observational comparisons found broadly similar human and agent results across reverts, PR-size analysis, flagged issue severity, and review rounds before merge.
Overall quality can conceal distinct failure patterns. Comment-based comparisons put Claude’s SQL-injection frequency at about 1.5× the human baseline and Devin’s authentication-bypass frequency at about half.
Validation should examine existing user behavior, future failure risk, and author intent. Greptile combines inspection of changed and related code with sandboxed browser interaction to support that work.
From completing a line to producing a pull request
An agent opening a hundred pull requests a day makes an impressive story. It also raises a practical question: can a company with real customers safely use those changes? Daksh Gupta, co-founder of Greptile, approaches that question with skepticism earned from programming before autonomous agents. Producing code quickly and producing code a business can merge are different accomplishments.
Greptile already sits between a proposed change and its acceptance. Its agents inspect changed files and related code to look for bugs. They can also install dependencies, start the application in a sandbox, mock inputs, and click through the running app. That position provides both code-review observations and a reason to care about whether the author is a person or an agent.
The shift in coding tools explains why the question has become urgent. Gupta traces a progression from tab completion, through multifile editing in 2024, to agents in 2025 that could take a task and create an entire pull request. Each step expands the unit of work delegated to software: first a continuation, then coordinated edits, then a proposed contribution ready for review. The remaining question is whether that contribution holds up in an existing enterprise codebase.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
What would make an agent-written PR worse?
Authorship answers who likely produced a change. Quality requires a separate definition. Gupta starts with reverts: a change that a team subsequently undoes is a useful candidate for a bad PR. He tracks GitHub’s revert branch naming pattern, which includes the original PR number and name, to connect an undo operation to its source change.
The reported rates are about one revert per thousand Codex PRs, about 3.5 per thousand Devin PRs, and about 2.5 per thousand human PRs. Humans sit between the two agents rather than clearly outperforming both. These are observational comparisons using reverts and automated review findings as quality proxies; they do not establish equivalent task difficulty or general superiority.
One plausible explanation was task allocation. People might give agents small, well-scoped changes while retaining complicated work themselves. More complicated PRs expose more opportunities for failure, so a raw revert comparison could penalize humans for doing harder jobs. Gupta therefore examines PR size alongside reverts. He reports that the size comparison did not provide strong evidence that human PRs were better. This addresses the proposed size explanation, although size alone cannot capture everything that makes a task difficult.
The next signal is the severity of issues Greptile flags. P0, P1, and P2 are the severity labels used in the reported review findings; the talk does not define their thresholds. Counting findings in these categories asks whether similar revert rates conceal different amounts of problematic code caught before acceptance. Three of the four tested agents produced fewer flagged P0 issues than humans. Across the severity categories, Gupta again reports broadly comparable quality.
Finally, review rounds measure the work needed to make a proposal mergeable. In the workflow described here, an agent reads review comments, addresses them, and commits a revision to the same branch. A higher-quality starting point should, in principle, need fewer cycles. Tracking the rounds between opening and merging a PR produced another comparison with little reported difference between human and agent contributions.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Similar overall quality, different ways to fail
A similar total number of problems can hide a different mix of problems. Greptile makes about four comments per PR on average, giving Gupta several million comments across the preceding months. Searching that corpus for phrases such as SQL injection and N+1 query lets him compare specific failure categories across authors.
The comparison sets human frequency to 1× separately for each error category:
- SQL injection: Claude’s reported relative frequency is about 1.5× the human baseline.
- Authentication bypass: Devin’s reported relative frequency is about half the human baseline.
The ratios compare each agent with humans within a category. They do not say that SQL injection and authentication bypass occur equally often, or that one agent is uniformly safer. Because the analysis searches review-comment language, these are patterns in flagged findings rather than independently confirmed incident rates.
This changes the useful question from “Are agents bad at coding?” to which failures a particular tool tends to introduce. The aggregate comparisons weakened Gupta’s initial skepticism about enterprise use; the category comparisons preserve a reason to inspect the resulting code carefully. Agents can make meaningful contributions while retaining distinctive weaknesses.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
More code turns validation into the bottleneck
Once agents can contribute useful code, the next constraint is the capacity to validate it. Gupta describes a large gap between typical Greptile users and the most prolific users, whose monthly output reaches hundreds or thousands of changes. Exact activity figures are uncertain because the recording and its description use different units and percentile figures. The operational concern is the same: manual review, testing, and external QA must keep pace with a much faster stream of proposed changes.
Greptile’s response starts with what must be established between a PR and a safe deployment. Gupta proposes three questions:
- Present behavior: Does the change violate the application’s user contract—the behavior users should be able to rely on?
- Future risk: Does the change make a later violation of that contract more likely?
- Author intent: Does the implementation accomplish what its author meant to change?
The questions separate three ways a change can disappoint. It can break existing behavior now, leave behind a risk that becomes a failure later, or preserve existing behavior while failing to deliver the requested improvement. This gives validation a purpose beyond accumulating review comments: establish whether the proposed change is safe and whether it does the intended job.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Inspect the surrounding code, then try the running app
How can validation connect a changed file to what users actually experience? Greptile combines two avenues described across the talk. Code analysis follows the change into related files, extending inspection beyond the diff. Sandbox execution installs dependencies, starts the app, supplies mocked inputs, and lets browser agents interact with it. One avenue examines the implementation and its relationships; the other looks for failures in running behavior.
The flow below answers where a proposed change becomes evidence for a revision. Findings from code inspection and browser interaction become review comments. The coding agent can then address those comments and commit a new version to the PR branch. The observable change is a revised proposal, and further review determines whether more work is needed before merge.
Gupta expects this combination to discover many issues and increase confidence in merging. The talk supplies the approach rather than a measured coverage guarantee, and it does not spell out how each of the three validation questions is conclusively answered. Its closing adoption signal is that almost a fifth of the PRs Greptile reviews are merged without human review or human testing. Greptile wants that share to grow while maintaining code quality. The ambition is to make validation scale with code production, so generating a useful change does not simply move an unmanageable workload onto reviewers.
A proposed code change enters review.
Code inspection and sandbox interaction provide findings for the same review-and-revision loop. The loop can repeat before merge.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
From the talk
- GreptileReference
The code-review product discussed in the talk; link supplied in the recording description.
Further reading
The complete AI Engineer recording, with paragraph citations tied to its source timestamps.
- Daksh GuptaReference
The speaker website supplied in the recording description.
Read the complete timestamped transcript
- 0:12
How's everyone doing?
- 0:14
Awesome. Uh, my name is Du. I'm one of
- 0:16
the co-founders of a company called
- 0:18
Gravile. And at Grapile, we're working
- 0:21
on AI agents that validate poll requests
- 0:25
with full context of the codebase. The
- 0:27
way Graval does this is every time
- 0:28
there's a pull request, we have a swarm
- 0:30
of agents and they go and look at every
- 0:32
file that's changed, every file that's
- 0:35
related to the files that changed to
- 0:36
figure out if there's bugs. And it also
- 0:38
spins up your code in a sandbox,
- 0:40
installs the dependencies, spins up
- 0:42
local host, clicks around to try to
- 0:44
break things, mocks inputs, and
- 0:47
everything else to figure out if the
- 0:48
code is broken. But I'm going to talk
- 0:50
about something a little bit different,
- 0:52
which is fully autonomous coding agents.
- 0:56
So
- 0:57
I first moved to San Francisco three
- 0:59
years ago to work on AI coding. And this
- 1:00
is because GBT 3.5 which came out in
- 1:03
2022 was the first model that seemed
- 1:06
like it was really good at programming
- 1:08
at the time. Code complete tab complete
- 1:11
was like the main paradigm for coding.
- 1:13
The most exciting thing that was
- 1:14
happening was you had tab complete from
- 1:15
cursor and you had co-pilots code
- 1:17
complete and that was like the way in
- 1:19
which AI coding works. In 2024, for the
- 1:23
first time, multifile editing started to
- 1:26
work. Cursor was the first to do this
- 1:28
and then a couple of other products
- 1:29
followed. But now, for the first time,
- 1:31
AI could simultaneously edit multiple
- 1:33
files at once. But things got really
- 1:36
interesting in 2025 because for the
- 1:37
first time, we had autonomous agents.
- 1:39
They could be given a task and they
- 1:40
could go and fulfill the task and create
- 1:42
entire pull requests all at once. And
- 1:43
this came to a precipice in December of
- 1:45
last year when coding agents became
- 1:48
literally completely autonomous. And
- 1:50
anyone that's working in AI coding knows
- 1:51
that December when the new models came
- 1:52
out was sort of this like watershed
- 1:54
moment in the history of AI coding
- 1:56
because these products for the first
- 1:57
time were like truly autonomous in how
- 1:59
they programmed. And all over Twitter
- 2:02
were all these crazy stories of people
- 2:04
that were having these agents spin off
- 2:05
and open a 100 pull requests a day. They
- 2:08
were experimenting with polyphasic sleep
- 2:10
so they could stay awake to re-trigger
- 2:12
their agents from time to time. And as
- 2:14
someone that started programming before
- 2:17
agents, I was very skeptical that real
- 2:19
companies could program this way. So I
- 2:21
was really really curious, are these
- 2:23
fully vi PRs really actually good? Is
- 2:25
this something that's sort of a Twitter
- 2:27
hype and you can have independent
- 2:28
developers and and maybe like small
- 2:30
startups with no real customers do this?
- 2:32
But anyone with real customers with an
- 2:35
actually commercially viable code base,
- 2:38
could they still use these end-to-end
- 2:40
coding agents?
- 2:42
And at Grubile, we're fortunate to work
- 2:44
with some very large companies. We work
- 2:46
with Nvidia, with Coinbase, with Scale,
- 2:48
with Data Dog, with American Express.
- 2:49
And we got really excited to see if
- 2:53
these coding agents were being used
- 2:55
really at these large companies and if
- 2:56
they were any good at doing real world
- 2:59
coding.
- 3:02
So, as an amateur data scientist, I
- 3:04
decided to go and dive into our data. We
- 3:06
review more than a million poll requests
- 3:07
a month and we tend to get a lot of
- 3:10
interesting data on what's good and bad
- 3:11
about these poll requests. We do this
- 3:12
for thousands and thousands of companies
- 3:14
and they're usually enterprises or at
- 3:16
least companies with serious products
- 3:18
with real customers. And so we figured
- 3:20
actually studying this data would yield
- 3:22
some really interesting results. Was a
- 3:24
surprisingly hard to figure out from all
- 3:25
the poll requests that GR was reviewing
- 3:27
which of the poll requests was actually
- 3:29
generated by AI. Surprisingly difficult
- 3:31
to do this. The first thing that I tried
- 3:34
was I started looking at the GitHub
- 3:35
author field. GitHub every commit has an
- 3:39
author field and so went through the
- 3:41
authors and turns out that less than 1%
- 3:43
of pull requests had codeex or clawed or
- 3:45
cursor as the author. But it wasn't
- 3:48
intuitive to me that only but 1% of all
- 3:50
code was fully AI generated. It just
- 3:51
seemed intuitive that that number would
- 3:52
be much higher. So I started to look for
- 3:54
other signals. Now thankfully Claude and
- 3:58
Codex and some other products leave a PR
- 4:00
description in the footer. In the PR
- 4:02
description, they put something along
- 4:03
the lines of co-authored by claude or
- 4:04
co-authored by cursor and so on. And
- 4:06
this gave me a little bit more signal, a
- 4:08
little bit more data on which PRs were
- 4:09
ostensibly largely, if not completely AI
- 4:12
generated. The third thing I looked at
- 4:14
was branch name prefixes. If you use
- 4:16
Codex, you know that Codex names as
- 4:17
branches. And it seemed reasonable that
- 4:19
if someone was using these coding agents
- 4:21
in such a sense that they were writing
- 4:23
the branch names or writing the PR
- 4:24
descriptions that the PRs were largely
- 4:26
AI generated. That seemed like a
- 4:27
reasonable assumption to me. So now I
- 4:29
had a good set of signals to indicate
- 4:30
whether a poll request was in fact
- 4:32
completely or largely AI generated or
- 4:34
not. And I came to the result that about
- 4:37
a quarter of all the poll requests that
- 4:38
Graal was reviewing in any given month
- 4:40
were completely or at least largely
- 4:43
generated by AI.
- 4:46
Then I backtracked this data to the last
- 4:48
12 months and it turns out that this
- 4:49
number is growing really fast. In fact,
- 4:52
early last year, fewer than 1% of pull
- 4:54
requests had any evidence of being
- 4:55
completely AI generated. So this number
- 4:57
is going very very fast as model
- 4:59
performance is getting better. What's
- 5:00
also interesting is that you can't tell
- 5:03
when new models came out on this chart.
- 5:05
It seemed like the progress is very
- 5:06
continuous. These things are just
- 5:08
generally diffusing into the economy at
- 5:10
a generally high pace.
- 5:13
So then came the next question.
- 5:14
Everyone's vibe coding. Everyone's
- 5:16
producing these PRs that are
- 5:17
antiendentic. Not human plus AI but
- 5:19
literally AI. The question is are these
- 5:21
PRs any good? First question, what does
- 5:23
it mean for PR to be good? That seemed
- 5:25
like actually a very important question
- 5:26
to answer before we got into judging
- 5:28
these AI generated PRs. And so I took a
- 5:31
few methods to try this. The first one I
- 5:32
tried was revert rates. If a pull
- 5:34
request reverted, it was probably pretty
- 5:36
bad. And so that seemed like a pretty
- 5:40
sincere sort of guess at what a bad PR
- 5:42
could be. And so I started tracking,
- 5:44
okay, which PRs were in fact reverted.
- 5:46
GitHub conveniently names its branches
- 5:48
with a revert dash PR number and PR
- 5:51
name. And so I was able to track the
- 5:52
rate at which these things were being
- 5:53
reverted. And I found some interesting
- 5:55
data. Codex PR is reverted about one out
- 5:58
of every thousand poll requests. Devon
- 6:00
once every three and a half times every
- 6:02
thousand poll requests. Humans right in
- 6:04
the middle at about two and a half. So
- 6:05
there doesn't seem to be very big
- 6:07
difference between the rate at which
- 6:08
poll requests were reverted from people
- 6:09
versus agents in my study. Now I was
- 6:12
very skeptical of this and I figured
- 6:14
okay this is probably because humans are
- 6:16
making agents do the easier work the
- 6:17
simpler well scoped pull requests and
- 6:20
the more complex work was being done by
- 6:21
humans and of course complex work had a
- 6:23
higher propensity to be reverted because
- 6:24
the poll requests were more complicated
- 6:26
had larger surface areas risks and so
- 6:28
on. So I decided to measure the average
- 6:30
size of PR and whether the revert rates
- 6:33
were sort of aligned for both of those
- 6:35
and I found really interesting data.
- 6:36
Turns out there's actually very little
- 6:38
correlation, if any, between the size of
- 6:40
PRs of humans were getting reverted
- 6:41
versus size of PRs from agents that were
- 6:43
being reverted. So this does not seem
- 6:45
like actually very strong evidence that
- 6:47
the human PRs are any better than the
- 6:49
agent PRs.
- 6:52
The second set of signals I started
- 6:53
looking at was Graile comments. Now, now
- 6:54
Grapple is reviewing all these poll
- 6:56
requests and Graile finds P 0's, P1's,
- 6:59
P2s across all these changes and it
- 7:00
seemed reasonable that if Grapile was
- 7:03
finding more bugs in the code that the
- 7:05
code was most likely worse. So then I
- 7:07
started tracking the counts of P 0, P1's
- 7:10
and P2s in all these changes. And
- 7:12
interestingly, once again, there was not
- 7:14
that much of a difference. In fact,
- 7:16
three out of the four agents that we
- 7:17
tested performed better than humans in
- 7:20
terms of the rate at which they were
- 7:21
producing P 0. they were producing fewer
- 7:23
P 0 than humans were. Similarly the case
- 7:26
with P1's and P2s. Broadly speaking,
- 7:29
human generated PRs were about equal in
- 7:32
quality to agent generated PRs based on
- 7:33
this data.
- 7:43
have the agent go and look at those
- 7:46
comments, address them, and make a new
- 7:47
commit on the pull request branch. That
- 7:49
is the most common way to use guptile.
- 7:51
And so it seems reasonable that if the
- 7:52
pull requests were higher quality, it
- 7:54
would require fewer iterations before
- 7:56
they would be merged. And so I started
- 7:57
tracking the number of iterations, the
- 7:59
number of review rounds between when the
- 8:01
poll request was opened and when it was
- 8:02
merged.
- 8:04
Sure enough, very little difference.
- 8:06
Devon's PRs 2.1, Codex PR is 2.45, four,
- 8:10
five number of review cycles to merge
- 8:12
and humans right in the middle. Once
- 8:14
again, very little if any statistical
- 8:17
difference between human generated and
- 8:18
AI generated pull request in terms of
- 8:20
how many iterations before they were
- 8:21
ready to merge.
- 8:24
So now that there wasn't any strong
- 8:26
evidence that agent generated PRs were
- 8:28
worse than human generated PRs, I got
- 8:30
curious if there were qualitative
- 8:31
differences. Maybe the types of ways in
- 8:33
which agents failed were different from
- 8:34
the types of ways that humans failed.
- 8:36
And so then I looked at the corpus of
- 8:38
reptiles comments. It makes on average
- 8:40
about four comments per pull request. So
- 8:42
we had this corpus of several million
- 8:44
comments across the last several months.
- 8:46
And so I started scanning them for
- 8:47
specific phrases and words. For
- 8:49
instance, SQL injection or N plus1
- 8:52
query. And I plotted the frequency with
- 8:54
which these terms occur in graphile
- 8:55
comments for these various agents.
- 8:58
And so I made this chart of patterns of
- 9:01
failure.
- 9:03
To interpret this chart, you can assume
- 9:05
that 1x is the human propensity for
- 9:07
producing that type of error across the
- 9:08
entire chart. And you can kind of see
- 9:10
that there's actually quite a lot of
- 9:12
variation in the types of failures that
- 9:14
these agents seem to produce. For
- 9:16
instance, Claude is one and a half times
- 9:18
more likely to produce a SQL injection
- 9:20
error than humans. Devon is about half
- 9:23
as likely as humans to produce a off
- 9:26
bypass issue.
- 9:29
I found it very interesting that there
- 9:30
was this much variation in how these
- 9:31
agents were performing and how different
- 9:34
their failure modes were from humans.
- 9:38
So it turns out that in spite of my
- 9:40
initial skepticism around the enterprise
- 9:43
usability of endto-end coding agents,
- 9:46
the evidence seems to suggest that
- 9:47
they're here and they probably can
- 9:49
contribute in real meaningful ways to
- 9:51
enterprise coding environments. And so
- 9:54
we started to think a little bit more
- 9:55
about what code review would look like
- 9:56
in such a world. Here's an interesting
- 9:59
stat. Today, Graal is used by several
- 10:01
tens of thousands of engineers every
- 10:03
single week to review all of their code.
- 10:05
So, we also know generally speaking the
- 10:08
number of pull requests that each of
- 10:09
these people are writing. The median
- 10:11
grubile user writes 50 pull requests a
- 10:14
month. So, call it about two plus per
- 10:16
workday. The 90th percentile writes 500
- 10:19
poll requests per month. That is a
- 10:21
drastic difference between the median
- 10:22
and the P90. The P99 is in the thousands
- 10:25
of pull requests. That is very
- 10:27
interesting in its own sense because
- 10:28
that means that the people on the
- 10:30
margins are actually producing poll
- 10:31
requests at the rate at which they're
- 10:32
coming up with new ideas. And beyond
- 10:35
being interesting, it also makes you
- 10:37
wonder well there existing systems for
- 10:38
validating this code which is manual
- 10:40
code review of course testing maybe you
- 10:43
have a QA firm that you work with
- 10:44
naturally can scale to that same degree.
- 10:47
And at Grav, we decided to take sort of
- 10:48
a first principles view at what really
- 10:50
good validation could look like. Instead
- 10:52
of saying that we wanted to automate QA
- 10:53
or automate testing or automate code
- 10:55
review, we took a step back and said,
- 10:57
what would we need to do for anyone to
- 10:59
be able to merge hundreds of pull
- 11:01
requests a month in an enterprise
- 11:02
environment where it really matters that
- 11:04
the code is correct? What would need to
- 11:06
happen between when the code was
- 11:07
expressed into a pull request and when
- 11:09
it was merged and deployed safely? We
- 11:11
figured we only actually had to answer
- 11:13
three questions. The first one is does
- 11:15
this change violate the user contract?
- 11:17
The second one, does it increase the
- 11:19
propensity of a future violation of the
- 11:21
user contract, whatever the user
- 11:22
contract might be for that application?
- 11:24
And third, does it fulfill the intent
- 11:26
that the author described? Did the poll
- 11:28
request do the thing that the author
- 11:29
wanted to do? And so we started
- 11:32
approaching this problem from base
- 11:34
ground level and said, okay, agents can
- 11:36
probably figure out if something's going
- 11:38
to violate the user contract and detect
- 11:40
bugs. If you let it spin up the code in
- 11:42
a sandbox, have it install the
- 11:44
dependencies, mock the inputs, run the
- 11:46
browser agents, you can probably start
- 11:47
to discover most of the issues that
- 11:50
could occur, you get a pretty high
- 11:51
degree of confidence on merge. Today,
- 11:54
almost a fifth of all the poll request
- 11:55
or gravile reviews are merged without
- 11:57
any human review or without any human
- 11:58
testing, which I find very interesting
- 12:00
and that is a number that we care a lot
- 12:01
about and want to bring higher and
- 12:03
higher, of course, within the guard
- 12:04
rails of producing really high quality
- 12:06
code.
- 12:10
Thank you so much. My name is D, one of
- 12:12
the co-founders of Graptile. We have a
- 12:13
booth here which you should come to. And
- 12:15
if you're interested in trying Graptile,
- 12:17
you can find us at gretile.com and try
- 12:19
it today for free. And we'd love to hear
- 12:22
your feedback. Thank you so much.