AI Engineer World's Fair 2026

AI-Generated Code Is Already Competing With Human Code — Daksh Gupta, Greptile

Read the talk

AI-Generated Code Is Already Competing With Human Code

Daksh Gupta examines Greptile’s enterprise pull-request data, finds similar quality signals for human and agent code, and explains why growing code volume changes what validation must accomplish.

From a talk by Daksh Gupta

At a glance

Ideas worth remembering

  • Agent authorship needs multiple signals. Author fields alone identified fewer than 1% of PRs; adding co-author footers and branch prefixes raised the estimate to roughly a quarter, with full autonomy still inferred.

  • Greptile’s observational comparisons found broadly similar human and agent results across reverts, PR-size analysis, flagged issue severity, and review rounds before merge.

  • Overall quality can conceal distinct failure patterns. Comment-based comparisons put Claude’s SQL-injection frequency at about 1.5× the human baseline and Devin’s authentication-bypass frequency at about half.

  • Validation should examine existing user behavior, future failure risk, and author intent. Greptile combines inspection of changed and related code with sandboxed browser interaction to support that work.

From completing a line to producing a pull request

An agent opening a hundred pull requests a day makes an impressive story. It also raises a practical question: can a company with real customers safely use those changes? Daksh Gupta, co-founder of Greptile, approaches that question with skepticism earned from programming before autonomous agents. Producing code quickly and producing code a business can merge are different accomplishments.

Greptile already sits between a proposed change and its acceptance. Its agents inspect changed files and related code to look for bugs. They can also install dependencies, start the application in a sandbox, mock inputs, and click through the running app. That position provides both code-review observations and a reason to care about whether the author is a person or an agent.

The shift in coding tools explains why the question has become urgent. Gupta traces a progression from tab completion, through multifile editing in 2024, to agents in 2025 that could take a task and create an entire pull request. Each step expands the unit of work delegated to software: first a continuation, then coordinated edits, then a proposed contribution ready for review. The remaining question is whether that contribution holds up in an existing enterprise codebase.

2:022:04
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:12 · section reference included

An author field misses much of the agent activity

Greptile reviews more than a million pull requests a month across thousands of companies, usually enterprises or businesses with established products and customers. Before comparing quality, Gupta had to divide that stream into likely agent-written and human-written changes. The first attempt—looking for agent names in GitHub author fields—identified fewer than 1% of pull requests.

Source frame: An author field misses much of the agent activity
Source frame: An author field misses much of the agent activity

That result prompted a wider search for traces left by the tools:

  • Author fields: An agent named as the author provides a direct metadata signal.
  • Co-author footers: PR descriptions may credit Claude, Cursor, or another tool even when the author field does not.
  • Branch prefixes: Agent-created branch names provide another indication that a tool participated in producing the pull request.

Consider the classification change implied by the footer check. A PR whose author field names no agent would escape the first detector. If its description credits Claude, the expanded detector now has evidence of agent involvement. Nothing about the code has changed; the observable change is in how the PR is classified. Gupta takes an agent writing the description or branch name as a reasonable sign that it also wrote much of the implementation. That remains an inference: these signals do not establish full autonomy, and an unmarked PR need not be entirely human-written.

With the combined signals, roughly a quarter of reviewed PRs appeared largely or entirely AI-generated. Looking back over the preceding twelve months, Gupta found fewer than 1% showing such evidence early in that period. He describes the adoption curve as continuous: individual model releases did not stand out as obvious discontinuities. The striking observation is the spread of the workflow through companies, rather than a single release-day jump.

3:023:04
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

3:02 · section reference included

What would make an agent-written PR worse?

Authorship answers who likely produced a change. Quality requires a separate definition. Gupta starts with reverts: a change that a team subsequently undoes is a useful candidate for a bad PR. He tracks GitHub’s revert branch naming pattern, which includes the original PR number and name, to connect an undo operation to its source change.

Source frame: What would make an agent-written PR worse?
Source frame: What would make an agent-written PR worse?

The reported rates are about one revert per thousand Codex PRs, about 3.5 per thousand Devin PRs, and about 2.5 per thousand human PRs. Humans sit between the two agents rather than clearly outperforming both. These are observational comparisons using reverts and automated review findings as quality proxies; they do not establish equivalent task difficulty or general superiority.

One plausible explanation was task allocation. People might give agents small, well-scoped changes while retaining complicated work themselves. More complicated PRs expose more opportunities for failure, so a raw revert comparison could penalize humans for doing harder jobs. Gupta therefore examines PR size alongside reverts. He reports that the size comparison did not provide strong evidence that human PRs were better. This addresses the proposed size explanation, although size alone cannot capture everything that makes a task difficult.

The next signal is the severity of issues Greptile flags. P0, P1, and P2 are the severity labels used in the reported review findings; the talk does not define their thresholds. Counting findings in these categories asks whether similar revert rates conceal different amounts of problematic code caught before acceptance. Three of the four tested agents produced fewer flagged P0 issues than humans. Across the severity categories, Gupta again reports broadly comparable quality.

Finally, review rounds measure the work needed to make a proposal mergeable. In the workflow described here, an agent reads review comments, addresses them, and commits a revision to the same branch. A higher-quality starting point should, in principle, need fewer cycles. Tracking the rounds between opening and merging a PR produced another comparison with little reported difference between human and agent contributions.

5:135:19
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

5:13 · section reference included

Similar overall quality, different ways to fail

A similar total number of problems can hide a different mix of problems. Greptile makes about four comments per PR on average, giving Gupta several million comments across the preceding months. Searching that corpus for phrases such as SQL injection and N+1 query lets him compare specific failure categories across authors.

Source frame: Similar overall quality, different ways to fail
Source frame: Similar overall quality, different ways to fail

The comparison sets human frequency to 1× separately for each error category:

  • SQL injection: Claude’s reported relative frequency is about 1.5× the human baseline.
  • Authentication bypass: Devin’s reported relative frequency is about half the human baseline.

The ratios compare each agent with humans within a category. They do not say that SQL injection and authentication bypass occur equally often, or that one agent is uniformly safer. Because the analysis searches review-comment language, these are patterns in flagged findings rather than independently confirmed incident rates.

This changes the useful question from “Are agents bad at coding?” to which failures a particular tool tends to introduce. The aggregate comparisons weakened Gupta’s initial skepticism about enterprise use; the category comparisons preserve a reason to inspect the resulting code carefully. Agents can make meaningful contributions while retaining distinctive weaknesses.

8:248:26
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

8:24 · section reference included

More code turns validation into the bottleneck

Once agents can contribute useful code, the next constraint is the capacity to validate it. Gupta describes a large gap between typical Greptile users and the most prolific users, whose monthly output reaches hundreds or thousands of changes. Exact activity figures are uncertain because the recording and its description use different units and percentile figures. The operational concern is the same: manual review, testing, and external QA must keep pace with a much faster stream of proposed changes.

Greptile’s response starts with what must be established between a PR and a safe deployment. Gupta proposes three questions:

  • Present behavior: Does the change violate the application’s user contract—the behavior users should be able to rely on?
  • Future risk: Does the change make a later violation of that contract more likely?
  • Author intent: Does the implementation accomplish what its author meant to change?

The questions separate three ways a change can disappoint. It can break existing behavior now, leave behind a risk that becomes a failure later, or preserve existing behavior while failing to deliver the requested improvement. This gives validation a purpose beyond accumulating review comments: establish whether the proposed change is safe and whether it does the intended job.

9:549:55
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

9:54 · section reference included

Inspect the surrounding code, then try the running app

How can validation connect a changed file to what users actually experience? Greptile combines two avenues described across the talk. Code analysis follows the change into related files, extending inspection beyond the diff. Sandbox execution installs dependencies, starts the app, supplies mocked inputs, and lets browser agents interact with it. One avenue examines the implementation and its relationships; the other looks for failures in running behavior.

The flow below answers where a proposed change becomes evidence for a revision. Findings from code inspection and browser interaction become review comments. The coding agent can then address those comments and commit a new version to the PR branch. The observable change is a revised proposal, and further review determines whether more work is needed before merge.

Gupta expects this combination to discover many issues and increase confidence in merging. The talk supplies the approach rather than a measured coverage guarantee, and it does not spell out how each of the three validation questions is conclusively answered. Its closing adoption signal is that almost a fifth of the PRs Greptile reviews are merged without human review or human testing. Greptile wants that share to grow while maintaining code quality. The ambition is to make validation scale with code production, so generating a useful change does not simply move an unmanageable workload onto reviewers.

How it fits togetherFrom proposed code to a revised pull request

A proposed code change enters review.

Code inspection and sandbox interaction provide findings for the same review-and-revision loop. The loop can repeat before merge.

0:280:30
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

11:32 · section reference included

Resources

From the talk

  • GreptileReference

    The code-review product discussed in the talk; link supplied in the recording description.

  • The complete AI Engineer recording, with paragraph citations tied to its source timestamps.

  • Daksh GuptaReference

    The speaker website supplied in the recording description.

Read the complete timestamped transcript
  1. 0:12

    How's everyone doing?

  2. 0:14

    Awesome. Uh, my name is Du. I'm one of

  3. 0:16

    the co-founders of a company called

  4. 0:18

    Gravile. And at Grapile, we're working

  5. 0:21

    on AI agents that validate poll requests

  6. 0:25

    with full context of the codebase. The

  7. 0:27

    way Graval does this is every time

  8. 0:28

    there's a pull request, we have a swarm

  9. 0:30

    of agents and they go and look at every

  10. 0:32

    file that's changed, every file that's

  11. 0:35

    related to the files that changed to

  12. 0:36

    figure out if there's bugs. And it also

  13. 0:38

    spins up your code in a sandbox,

  14. 0:40

    installs the dependencies, spins up

  15. 0:42

    local host, clicks around to try to

  16. 0:44

    break things, mocks inputs, and

  17. 0:47

    everything else to figure out if the

  18. 0:48

    code is broken. But I'm going to talk

  19. 0:50

    about something a little bit different,

  20. 0:52

    which is fully autonomous coding agents.

  21. 0:56

    So

  22. 0:57

    I first moved to San Francisco three

  23. 0:59

    years ago to work on AI coding. And this

  24. 1:00

    is because GBT 3.5 which came out in

  25. 1:03

    2022 was the first model that seemed

  26. 1:06

    like it was really good at programming

  27. 1:08

    at the time. Code complete tab complete

  28. 1:11

    was like the main paradigm for coding.

  29. 1:13

    The most exciting thing that was

  30. 1:14

    happening was you had tab complete from

  31. 1:15

    cursor and you had co-pilots code

  32. 1:17

    complete and that was like the way in

  33. 1:19

    which AI coding works. In 2024, for the

  34. 1:23

    first time, multifile editing started to

  35. 1:26

    work. Cursor was the first to do this

  36. 1:28

    and then a couple of other products

  37. 1:29

    followed. But now, for the first time,

  38. 1:31

    AI could simultaneously edit multiple

  39. 1:33

    files at once. But things got really

  40. 1:36

    interesting in 2025 because for the

  41. 1:37

    first time, we had autonomous agents.

  42. 1:39

    They could be given a task and they

  43. 1:40

    could go and fulfill the task and create

  44. 1:42

    entire pull requests all at once. And

  45. 1:43

    this came to a precipice in December of

  46. 1:45

    last year when coding agents became

  47. 1:48

    literally completely autonomous. And

  48. 1:50

    anyone that's working in AI coding knows

  49. 1:51

    that December when the new models came

  50. 1:52

    out was sort of this like watershed

  51. 1:54

    moment in the history of AI coding

  52. 1:56

    because these products for the first

  53. 1:57

    time were like truly autonomous in how

  54. 1:59

    they programmed. And all over Twitter

  55. 2:02

    were all these crazy stories of people

  56. 2:04

    that were having these agents spin off

  57. 2:05

    and open a 100 pull requests a day. They

  58. 2:08

    were experimenting with polyphasic sleep

  59. 2:10

    so they could stay awake to re-trigger

  60. 2:12

    their agents from time to time. And as

  61. 2:14

    someone that started programming before

  62. 2:17

    agents, I was very skeptical that real

  63. 2:19

    companies could program this way. So I

  64. 2:21

    was really really curious, are these

  65. 2:23

    fully vi PRs really actually good? Is

  66. 2:25

    this something that's sort of a Twitter

  67. 2:27

    hype and you can have independent

  68. 2:28

    developers and and maybe like small

  69. 2:30

    startups with no real customers do this?

  70. 2:32

    But anyone with real customers with an

  71. 2:35

    actually commercially viable code base,

  72. 2:38

    could they still use these end-to-end

  73. 2:40

    coding agents?

  74. 2:42

    And at Grubile, we're fortunate to work

  75. 2:44

    with some very large companies. We work

  76. 2:46

    with Nvidia, with Coinbase, with Scale,

  77. 2:48

    with Data Dog, with American Express.

  78. 2:49

    And we got really excited to see if

  79. 2:53

    these coding agents were being used

  80. 2:55

    really at these large companies and if

  81. 2:56

    they were any good at doing real world

  82. 2:59

    coding.

  83. 3:02

    So, as an amateur data scientist, I

  84. 3:04

    decided to go and dive into our data. We

  85. 3:06

    review more than a million poll requests

  86. 3:07

    a month and we tend to get a lot of

  87. 3:10

    interesting data on what's good and bad

  88. 3:11

    about these poll requests. We do this

  89. 3:12

    for thousands and thousands of companies

  90. 3:14

    and they're usually enterprises or at

  91. 3:16

    least companies with serious products

  92. 3:18

    with real customers. And so we figured

  93. 3:20

    actually studying this data would yield

  94. 3:22

    some really interesting results. Was a

  95. 3:24

    surprisingly hard to figure out from all

  96. 3:25

    the poll requests that GR was reviewing

  97. 3:27

    which of the poll requests was actually

  98. 3:29

    generated by AI. Surprisingly difficult

  99. 3:31

    to do this. The first thing that I tried

  100. 3:34

    was I started looking at the GitHub

  101. 3:35

    author field. GitHub every commit has an

  102. 3:39

    author field and so went through the

  103. 3:41

    authors and turns out that less than 1%

  104. 3:43

    of pull requests had codeex or clawed or

  105. 3:45

    cursor as the author. But it wasn't

  106. 3:48

    intuitive to me that only but 1% of all

  107. 3:50

    code was fully AI generated. It just

  108. 3:51

    seemed intuitive that that number would

  109. 3:52

    be much higher. So I started to look for

  110. 3:54

    other signals. Now thankfully Claude and

  111. 3:58

    Codex and some other products leave a PR

  112. 4:00

    description in the footer. In the PR

  113. 4:02

    description, they put something along

  114. 4:03

    the lines of co-authored by claude or

  115. 4:04

    co-authored by cursor and so on. And

  116. 4:06

    this gave me a little bit more signal, a

  117. 4:08

    little bit more data on which PRs were

  118. 4:09

    ostensibly largely, if not completely AI

  119. 4:12

    generated. The third thing I looked at

  120. 4:14

    was branch name prefixes. If you use

  121. 4:16

    Codex, you know that Codex names as

  122. 4:17

    branches. And it seemed reasonable that

  123. 4:19

    if someone was using these coding agents

  124. 4:21

    in such a sense that they were writing

  125. 4:23

    the branch names or writing the PR

  126. 4:24

    descriptions that the PRs were largely

  127. 4:26

    AI generated. That seemed like a

  128. 4:27

    reasonable assumption to me. So now I

  129. 4:29

    had a good set of signals to indicate

  130. 4:30

    whether a poll request was in fact

  131. 4:32

    completely or largely AI generated or

  132. 4:34

    not. And I came to the result that about

  133. 4:37

    a quarter of all the poll requests that

  134. 4:38

    Graal was reviewing in any given month

  135. 4:40

    were completely or at least largely

  136. 4:43

    generated by AI.

  137. 4:46

    Then I backtracked this data to the last

  138. 4:48

    12 months and it turns out that this

  139. 4:49

    number is growing really fast. In fact,

  140. 4:52

    early last year, fewer than 1% of pull

  141. 4:54

    requests had any evidence of being

  142. 4:55

    completely AI generated. So this number

  143. 4:57

    is going very very fast as model

  144. 4:59

    performance is getting better. What's

  145. 5:00

    also interesting is that you can't tell

  146. 5:03

    when new models came out on this chart.

  147. 5:05

    It seemed like the progress is very

  148. 5:06

    continuous. These things are just

  149. 5:08

    generally diffusing into the economy at

  150. 5:10

    a generally high pace.

  151. 5:13

    So then came the next question.

  152. 5:14

    Everyone's vibe coding. Everyone's

  153. 5:16

    producing these PRs that are

  154. 5:17

    antiendentic. Not human plus AI but

  155. 5:19

    literally AI. The question is are these

  156. 5:21

    PRs any good? First question, what does

  157. 5:23

    it mean for PR to be good? That seemed

  158. 5:25

    like actually a very important question

  159. 5:26

    to answer before we got into judging

  160. 5:28

    these AI generated PRs. And so I took a

  161. 5:31

    few methods to try this. The first one I

  162. 5:32

    tried was revert rates. If a pull

  163. 5:34

    request reverted, it was probably pretty

  164. 5:36

    bad. And so that seemed like a pretty

  165. 5:40

    sincere sort of guess at what a bad PR

  166. 5:42

    could be. And so I started tracking,

  167. 5:44

    okay, which PRs were in fact reverted.

  168. 5:46

    GitHub conveniently names its branches

  169. 5:48

    with a revert dash PR number and PR

  170. 5:51

    name. And so I was able to track the

  171. 5:52

    rate at which these things were being

  172. 5:53

    reverted. And I found some interesting

  173. 5:55

    data. Codex PR is reverted about one out

  174. 5:58

    of every thousand poll requests. Devon

  175. 6:00

    once every three and a half times every

  176. 6:02

    thousand poll requests. Humans right in

  177. 6:04

    the middle at about two and a half. So

  178. 6:05

    there doesn't seem to be very big

  179. 6:07

    difference between the rate at which

  180. 6:08

    poll requests were reverted from people

  181. 6:09

    versus agents in my study. Now I was

  182. 6:12

    very skeptical of this and I figured

  183. 6:14

    okay this is probably because humans are

  184. 6:16

    making agents do the easier work the

  185. 6:17

    simpler well scoped pull requests and

  186. 6:20

    the more complex work was being done by

  187. 6:21

    humans and of course complex work had a

  188. 6:23

    higher propensity to be reverted because

  189. 6:24

    the poll requests were more complicated

  190. 6:26

    had larger surface areas risks and so

  191. 6:28

    on. So I decided to measure the average

  192. 6:30

    size of PR and whether the revert rates

  193. 6:33

    were sort of aligned for both of those

  194. 6:35

    and I found really interesting data.

  195. 6:36

    Turns out there's actually very little

  196. 6:38

    correlation, if any, between the size of

  197. 6:40

    PRs of humans were getting reverted

  198. 6:41

    versus size of PRs from agents that were

  199. 6:43

    being reverted. So this does not seem

  200. 6:45

    like actually very strong evidence that

  201. 6:47

    the human PRs are any better than the

  202. 6:49

    agent PRs.

  203. 6:52

    The second set of signals I started

  204. 6:53

    looking at was Graile comments. Now, now

  205. 6:54

    Grapple is reviewing all these poll

  206. 6:56

    requests and Graile finds P 0's, P1's,

  207. 6:59

    P2s across all these changes and it

  208. 7:00

    seemed reasonable that if Grapile was

  209. 7:03

    finding more bugs in the code that the

  210. 7:05

    code was most likely worse. So then I

  211. 7:07

    started tracking the counts of P 0, P1's

  212. 7:10

    and P2s in all these changes. And

  213. 7:12

    interestingly, once again, there was not

  214. 7:14

    that much of a difference. In fact,

  215. 7:16

    three out of the four agents that we

  216. 7:17

    tested performed better than humans in

  217. 7:20

    terms of the rate at which they were

  218. 7:21

    producing P 0. they were producing fewer

  219. 7:23

    P 0 than humans were. Similarly the case

  220. 7:26

    with P1's and P2s. Broadly speaking,

  221. 7:29

    human generated PRs were about equal in

  222. 7:32

    quality to agent generated PRs based on

  223. 7:33

    this data.

  224. 7:43

    have the agent go and look at those

  225. 7:46

    comments, address them, and make a new

  226. 7:47

    commit on the pull request branch. That

  227. 7:49

    is the most common way to use guptile.

  228. 7:51

    And so it seems reasonable that if the

  229. 7:52

    pull requests were higher quality, it

  230. 7:54

    would require fewer iterations before

  231. 7:56

    they would be merged. And so I started

  232. 7:57

    tracking the number of iterations, the

  233. 7:59

    number of review rounds between when the

  234. 8:01

    poll request was opened and when it was

  235. 8:02

    merged.

  236. 8:04

    Sure enough, very little difference.

  237. 8:06

    Devon's PRs 2.1, Codex PR is 2.45, four,

  238. 8:10

    five number of review cycles to merge

  239. 8:12

    and humans right in the middle. Once

  240. 8:14

    again, very little if any statistical

  241. 8:17

    difference between human generated and

  242. 8:18

    AI generated pull request in terms of

  243. 8:20

    how many iterations before they were

  244. 8:21

    ready to merge.

  245. 8:24

    So now that there wasn't any strong

  246. 8:26

    evidence that agent generated PRs were

  247. 8:28

    worse than human generated PRs, I got

  248. 8:30

    curious if there were qualitative

  249. 8:31

    differences. Maybe the types of ways in

  250. 8:33

    which agents failed were different from

  251. 8:34

    the types of ways that humans failed.

  252. 8:36

    And so then I looked at the corpus of

  253. 8:38

    reptiles comments. It makes on average

  254. 8:40

    about four comments per pull request. So

  255. 8:42

    we had this corpus of several million

  256. 8:44

    comments across the last several months.

  257. 8:46

    And so I started scanning them for

  258. 8:47

    specific phrases and words. For

  259. 8:49

    instance, SQL injection or N plus1

  260. 8:52

    query. And I plotted the frequency with

  261. 8:54

    which these terms occur in graphile

  262. 8:55

    comments for these various agents.

  263. 8:58

    And so I made this chart of patterns of

  264. 9:01

    failure.

  265. 9:03

    To interpret this chart, you can assume

  266. 9:05

    that 1x is the human propensity for

  267. 9:07

    producing that type of error across the

  268. 9:08

    entire chart. And you can kind of see

  269. 9:10

    that there's actually quite a lot of

  270. 9:12

    variation in the types of failures that

  271. 9:14

    these agents seem to produce. For

  272. 9:16

    instance, Claude is one and a half times

  273. 9:18

    more likely to produce a SQL injection

  274. 9:20

    error than humans. Devon is about half

  275. 9:23

    as likely as humans to produce a off

  276. 9:26

    bypass issue.

  277. 9:29

    I found it very interesting that there

  278. 9:30

    was this much variation in how these

  279. 9:31

    agents were performing and how different

  280. 9:34

    their failure modes were from humans.

  281. 9:38

    So it turns out that in spite of my

  282. 9:40

    initial skepticism around the enterprise

  283. 9:43

    usability of endto-end coding agents,

  284. 9:46

    the evidence seems to suggest that

  285. 9:47

    they're here and they probably can

  286. 9:49

    contribute in real meaningful ways to

  287. 9:51

    enterprise coding environments. And so

  288. 9:54

    we started to think a little bit more

  289. 9:55

    about what code review would look like

  290. 9:56

    in such a world. Here's an interesting

  291. 9:59

    stat. Today, Graal is used by several

  292. 10:01

    tens of thousands of engineers every

  293. 10:03

    single week to review all of their code.

  294. 10:05

    So, we also know generally speaking the

  295. 10:08

    number of pull requests that each of

  296. 10:09

    these people are writing. The median

  297. 10:11

    grubile user writes 50 pull requests a

  298. 10:14

    month. So, call it about two plus per

  299. 10:16

    workday. The 90th percentile writes 500

  300. 10:19

    poll requests per month. That is a

  301. 10:21

    drastic difference between the median

  302. 10:22

    and the P90. The P99 is in the thousands

  303. 10:25

    of pull requests. That is very

  304. 10:27

    interesting in its own sense because

  305. 10:28

    that means that the people on the

  306. 10:30

    margins are actually producing poll

  307. 10:31

    requests at the rate at which they're

  308. 10:32

    coming up with new ideas. And beyond

  309. 10:35

    being interesting, it also makes you

  310. 10:37

    wonder well there existing systems for

  311. 10:38

    validating this code which is manual

  312. 10:40

    code review of course testing maybe you

  313. 10:43

    have a QA firm that you work with

  314. 10:44

    naturally can scale to that same degree.

  315. 10:47

    And at Grav, we decided to take sort of

  316. 10:48

    a first principles view at what really

  317. 10:50

    good validation could look like. Instead

  318. 10:52

    of saying that we wanted to automate QA

  319. 10:53

    or automate testing or automate code

  320. 10:55

    review, we took a step back and said,

  321. 10:57

    what would we need to do for anyone to

  322. 10:59

    be able to merge hundreds of pull

  323. 11:01

    requests a month in an enterprise

  324. 11:02

    environment where it really matters that

  325. 11:04

    the code is correct? What would need to

  326. 11:06

    happen between when the code was

  327. 11:07

    expressed into a pull request and when

  328. 11:09

    it was merged and deployed safely? We

  329. 11:11

    figured we only actually had to answer

  330. 11:13

    three questions. The first one is does

  331. 11:15

    this change violate the user contract?

  332. 11:17

    The second one, does it increase the

  333. 11:19

    propensity of a future violation of the

  334. 11:21

    user contract, whatever the user

  335. 11:22

    contract might be for that application?

  336. 11:24

    And third, does it fulfill the intent

  337. 11:26

    that the author described? Did the poll

  338. 11:28

    request do the thing that the author

  339. 11:29

    wanted to do? And so we started

  340. 11:32

    approaching this problem from base

  341. 11:34

    ground level and said, okay, agents can

  342. 11:36

    probably figure out if something's going

  343. 11:38

    to violate the user contract and detect

  344. 11:40

    bugs. If you let it spin up the code in

  345. 11:42

    a sandbox, have it install the

  346. 11:44

    dependencies, mock the inputs, run the

  347. 11:46

    browser agents, you can probably start

  348. 11:47

    to discover most of the issues that

  349. 11:50

    could occur, you get a pretty high

  350. 11:51

    degree of confidence on merge. Today,

  351. 11:54

    almost a fifth of all the poll request

  352. 11:55

    or gravile reviews are merged without

  353. 11:57

    any human review or without any human

  354. 11:58

    testing, which I find very interesting

  355. 12:00

    and that is a number that we care a lot

  356. 12:01

    about and want to bring higher and

  357. 12:03

    higher, of course, within the guard

  358. 12:04

    rails of producing really high quality

  359. 12:06

    code.

  360. 12:10

    Thank you so much. My name is D, one of

  361. 12:12

    the co-founders of Graptile. We have a

  362. 12:13

    booth here which you should come to. And

  363. 12:15

    if you're interested in trying Graptile,

  364. 12:17

    you can find us at gretile.com and try

  365. 12:19

    it today for free. And we'd love to hear

  366. 12:22

    your feedback. Thank you so much.