The State of AI in Software Development: Data from 400+ Orgs — Justin Reock, DX
Read the talk
The State of AI in Software Development: Data from 400+ Orgs
Justin Reock walks through DX’s findings on delivery speed, quality and AI use, then shows why faster code generation needs better measurement—and improvements across the rest of the development workflow.
From a talk by Justin Reock
At a glance
Ideas worth remembering
Deployment frequency and PR throughput describe work moving through the system. Read them alongside failure rates, review burden and business value.
Fast generation beside slow builds can encourage larger PRs. Improving validation helps preserve small, understandable changes.
AI utilization is not the same as efficiency: senior engineers report similar time savings to juniors while spending fewer tokens.
Measure utilization, impact and cost together, and use agent feedback to locate friction in specific use cases.
Look beyond generation for the limiting step. Legacy-code discovery, coordination, review and incident preparation are all concrete targets in the closing examples.
More deployments do not yet tell us how much value improved
Deployment frequency is rising in DX’s data, but that is only the opening question. Justin Reock, DX’s deputy CTO, presents research on AI’s relationship to developer experience and productivity, drawing on a study covering about 200,000 engineers. The useful tension is already visible: organizations are shipping more often, while engineers’ perceived rate of delivery has increased by only about 4.5% over a year.
DX gathers organizational data through a platform whose research lineage includes DORA metrics, the SPACE framework and the DevEx framework. Here, deployment frequency and pull-request throughput serve as proxies for how work moves. They help answer whether more changes reach production; they do not directly measure the value those changes create.
The deployment trend is steadily upward, with some tapering. North America continues to rise, while Europe pulls back in the most recent quarter discussed. Reock offers differences in working practices, spending and regulation as possible explanations for regional variation. Even where frequency improves, it leaves separate questions unanswered: how many changes fail, how many defects appear and how often teams revert a release.
Perceived speed deserves its own measurement. Reock contrasts DX’s modest increase with a small METR study that, in his account, involved 16 engineers: productivity fell by about 19% while perceived productivity rose by about 20%. He disputes that study’s methodology and describes a subsequent reconsideration of its signals. The comparison illustrates the question to ask rather than settling it: do engineers feel faster, and does observed work actually move faster? Those are different measurements.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Quality becomes more volatile as maintainability and confidence diverge
Change failure rate adds the missing quality dimension. In the company-level comparison Reock describes, some organizations improve and others worsen substantially. An increase of two percentage points is large against the approximately 4% industry benchmark he cites: moving from 4% to 6% would mean a 50% relative increase in the failure rate. That arithmetic concerns failed changes; it should not be read as a direct count of additional defects.
Release pipelines and automated tests already produce differences between organizations without AI. Reock’s interpretation is that AI increases the amplitude of those differences. The observations do not isolate AI as the sole cause, so the practical response is to measure the organization’s own failure trend alongside its release and testing conditions, rather than assume a universal quality effect.
Two developer-reported measures expose another split. Perceived code maintainability rises by almost 4%, while change confidence falls by 6%. These usually move together: modular, understandable code tends to make developers more comfortable changing it. Here, assistants can make code easier to understand and modify while leaving developers less certain that the resulting change is safe to ship.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Instant functions meet a 45-minute build pipeline
Average pull-request size grows from around 44 lines to 72 over roughly a year. Reock connects that growth to falling change confidence and makes PR size a metric to watch closely. His complaint carries the perspective of someone who has written code professionally since the late 1990s: implementing a use case with fewer lines used to be a mark of skill. Generating more code cheaply changes the incentive, but the review burden remains.
The clearest mechanism is his hypothetical build pipeline. An assistant generates four functions almost instantly, but a build takes 45 minutes or an hour. The developer could submit four separate PRs and face the build delay for each, or bundle the functions into one PR. Fast generation beside slow validation encourages the second choice. The observable change is a larger unit of work arriving for review, even though the original speedup happened during coding.
What changes when generation becomes fast but validation stays slow? The diagram follows the same four-function example. It shows how the unchanged build delay can encourage batching; it does not imply that every pipeline runs builds serially or that every developer makes this choice.
The cost travels downstream. More lines mean more material to review and more opportunities for bugs or vulnerabilities. Larger changes also make it harder to understand what changed and to roll it back. DX’s reported sentiment about being able to work in small, incremental changes falls by 10%. The pipeline example explains why improving generation alone can work against the small-batch delivery practice that makes releases manageable.
AI removes much of the wait involved in producing the code.
In Reock’s example, four quickly generated functions meet a 45–60-minute build. Bundling avoids repeated waits in the developer’s decision, while increasing the amount of code in one review.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The heaviest AI users do not necessarily save the most time
Junior engineers use AI the most in DX’s findings. Reock’s explanation is that newer entrants have less to unlearn and may already have used these tools during school. Usage, however, does not map directly onto the amount of time saved. Staff-plus and other senior engineers save about the same amount of time while using fewer tokens.
The comparison becomes useful at the level of a shared use case. Junior developers spend more tokens on the same kind of task, while senior engineers can more readily identify hallucinations and understand the architecture around a proposed change. Experience can therefore improve how efficiently someone directs and checks the tool. Counting interactions alone would miss that difference.
Smaller companies also lead in reported time savings. Reock attributes this to simpler release pipelines and less organizational complexity. That explanation connects back to the build example: the same coding capability has more room to affect delivery when fewer delays surround it.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Measure utilization, impact and cost together
Measuring developer productivity was difficult before AI. Adding token spending and tool telemetry does not resolve that problem by itself. DX’s AI Measurement Framework keeps existing developer-experience and productivity measures, then asks how AI use relates to them. Tool data identifies who uses what and where; delivery and quality measures help determine whether that use improves work.
The framework separates three questions:
- Utilization: Who uses the tools, how often and for which use cases? Daily and weekly active users describe adoption.
- Impact: Which outcomes should improve as use increases? Compare speed, quality and business value rather than treating adoption as the outcome.
- Cost: What does that use cost? Token spending belongs beside the benefits it is supposed to produce.
Organizations often begin with utilization because it is the first thing they can observe. The next step is to connect it to impact. Reock proposes separating users into cohorts and comparing PR cycle time, PR size and pushback during review. A cohort might move work faster while producing larger PRs or attracting more review corrections. Keeping those measures together makes the tradeoff visible; cohort comparisons establish associations rather than automatically proving that AI caused the difference.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Good developer experience becomes good agent experience
AI readiness shifts attention from the model to the environment it must work in. Reock compresses the progression into assistants first, agents next, and then the realization that infrastructure was not ready. If an agent struggles to find accurate context or obtain dependable test feedback, token spending becomes a poor substitute for fixing the platform.
The readiness criteria are familiar developer-experience investments:
- Understandable context: Clear, accurate, well-structured documentation and data structures with straightforward relationships.
- Manageable code: Modular code that is practical to understand and change.
- Dependable feedback: Reliable local CI and non-flaky test suites.
These conditions help both people and agents. Reock’s wry observation is that AI may finally motivate investments teams should have made over the preceding decades. The earlier pipeline problem fits here: faster generation does not remove the need for a development environment that can validate small changes efficiently.
Agent experience adds another source of feedback. DX asks agents about problems with human steering, the context supplied and the feedback cycles they went through. Breaking those reports down by use case helps identify where collaboration consumes tokens and where it produces useful results. These are qualitative reports from the agents, so their role is to help diagnose friction alongside measured cost and outcomes.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Code generation addresses only part of the value stream
The fifth trend widens the scope to the software and product development lifecycles. Reock estimates that even perfectly accurate, instant code generation would address only about 14–16% of the overall value stream. That is his framing of the opportunity, not a claim that current models produce perfect code. Requirements, coordination, review, release and operational work still surround generation.
The reported throughput results are much smaller than the familiar productivity multiples. In the study interval Reock describes as beginning in November 2024 and ending in February, median PR-throughput growth is 7.7%, with a 13% average. Even the top performers are in the 70% range. None reaches 2×. These are changes in a velocity proxy, not equivalent increases in business value.
Meetings, context switching, interruptions and development-environment friction can consume the time that AI saves. The four-function example makes the constraint concrete: code arrives immediately, yet validation still determines how the developer packages and delivers it. Reock invokes Eli Goldratt’s theory of constraints to make the operational point: improving a step that does not limit the flow may leave overall throughput unchanged. Find the limiting step before assuming that faster coding will fix delivery.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Agents remove discovery and coordination work across the lifecycle
Morgan Stanley’s example targets the work before implementation. An agent interprets legacy code, including COBOL and Perl, and creates product requirements documents for engineers. That handoff removes a reverse-engineering step: engineers receive a description of the existing behavior instead of having to reconstruct it all themselves. Reock reports approximately 300,000 hours saved per year at Morgan Stanley.
Zapier applies an ecosystem of agents to administrative work and coordination. Agent summaries help reduce stand-ups from five times a week to two. Engineers onboard in about two weeks, compared with DX’s cited industry benchmark of usually more than a month. These interventions attack time around coding—the same time that can otherwise swallow generation savings.
Reock reports about 15% more value creation per engineer at Zapier and says the company is hiring more than at any previous point in its history. His preferred interpretation is increased capacity: if each engineer produces more value, adding engineers can improve the company’s competitive position. He calls it a “throughput story.” The examples’ savings and value figures are reported organizational outcomes; the recording does not define a shared evaluation method that would make them directly comparable.
A further review example handles about 3,000 code reviews a week, triggered by pull requests. The agent checks superficial issues and leaves its findings in PR comments. Human reviewers still perform the substantive review, but they can see what the agent already examined. Keeping the output in the existing record makes the handoff inspectable and can reduce repeated work.
Spotify’s SRE agent targets discovery during an incident. It gathers remediation steps from runbooks, combines them with incident information and context, and posts the assembled material into SRE communication channels. The immediate change is that the responder starts with relevant context instead of spending several minutes finding it. The mechanism described is preparation for human incident response, with no claim here of autonomous remediation.
These closing examples give the bottleneck argument practical shape: interpret legacy behavior before implementation, reduce coordination overhead, make review findings reusable and assemble incident context before a responder needs it. Reock closes by pointing listeners to DX’s quarterly research reports. The work to carry forward is the same measurement loop: identify where time goes, apply AI to that use case, and check whether quality, cost and delivered value improve together.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
Related talks
- Beating RL With Reflection: GEPA and Optimize Anything
Explains how execution traces, tool errors and repository knowledge can guide agent improvement—a technical companion to Reock’s discussion of agent feedback and effective token spending.
Read the complete timestamped transcript
- 0:12
Uh, I'm Justin Reock. I'm the deputy CTO at DX. I get to do a lot of our research around, uh, developer experience and its relation to developer productivity. And certainly, over the last year, a lot of that research has focused on AI's impact on the developer experience and basic aspects of productivity, uh, within an organization. Uh, so we produce these quarterly reports, uh, our, our State of AI and AI Impact reports, which looks at really a lot of raw data that we're able to get from our platform. Our
- 0:42
platform is a research-backed platform, uh, built from the same folks who worked on DORA metrics and the SPACE framework and the DevEx framework. Uh, and really, at the end of the day, it's a data gathering platform that can look at a lot of different trends across an organization with, you know, how are we focusing on developer experience? Are we seeing outcomes and improvements with developer experience? And certainly, are we seeing outcomes with the way that we're using AI in our workflows? This session is gonna be a little bit different than the three forty-five
- 1:12
session, I think, that's up in the leadership room later today, which just looks at the raw data from our newest report. It's a little bit of a preview. This is looking at more, like, organizational trends and how our workflow's changing and, and things like that based on the data that we're seeing. Uh, so first things first, let's kinda look at what we're seeing in terms of the velocity impact. Now, if anybody's familiar with looking at various velocity metrics like PR throughput and deployment frequency, like, hopefully you understand the caveat that these metrics are not perfect. They are proxy metrics for understanding the way
- 1:42
that work flows through an organization. They're not fully representative of the value generation. We're gonna look at some metrics that are closer to that as well. But let's just try to understand what this technology is doing in terms of just the raw velocity metrics in the organization. So here's our DORA metric for deployment frequency. Who's familiar with DORA metrics? All right, most of you. Great. Um, so this is one of the, you know, four key metrics that tries to understand directionally, you know, the speed of shipping work within an environment,
- 2:12
and we're seeing it steadily increase. Um, it's tapering off a little bit. I think there's reasons for that. Uh, some of the surge that we saw early on was due to people shipping, you know, PRs more frequently and, and things like that. Um, so we're seeing increases in this. This is one small part of the SDLC PDLC, right? This is, like, the creation, uh, of the PR to actually trying to deploy this into production. It doesn't tell us about revert rates. It doesn't tell us about defect ratios.
- 2:42
It doesn't tell us about change failure rate. It just tells us about how much stuff are we now kind of pushing out into production. But, you know, the data is showing us that this number is steadily increasing, okay? Um, the trends vary a little bit across regions. Um, we are seeing North America just sort of like trending up. We've seen Europe actually pull back a little bit just in the last quarter. Um, I think, you know, there's, there's a number of reasons for that. I mean, obviously the way work is done in these regions are differently. The way that we can
- 3:12
spend money on this technology and think about token spend is different. The regulation is different. Um, but if we just look at the through line in general, you know, we still see improvements, uh, in, in this one metric of deploy frequency even across various regions. We do just see some variance in those regions. Here's one that's kind of interesting though. Uh, the perceived rate of delivery. How much faster do engineers feel in terms of what they're delivering?
- 3:42
It's gone up a little bit. This number shows an increase of about four and a half-ish percent. Um, but given that that's a full year and given all the investment and cost of new spend on AI, it's interesting that this perceived rate has not stayed-- ha-has not kind of crept up a little bit more, especially when you look at some studies like obviously infamous flawed METR study, if you're familiar with that study. Uh, back in February, they actually released a follow-up to that study saying that some of the signals they gathered were sort of
- 4:12
imperfect and they needed to rethink some of that. But what was interesting was that in these sixteen engineers that were part of that study, small study, flawed study, uh, the productivity went down by about nineteen percent, but the perception went up by about twenty percent. So there was like a forty percent spread in perception of productivity versus actual productivity. But here, when we look at this on more of an aggregate, things interestingly stay, uh, sort of flat. All right, what about impact on quality and, and delivery, right? So those are just kind of our speed, our
- 4:42
velocity metrics, but what is this actually doing to our software quality? Um, the first metric we'll look at here is the DORA change failure rate metric, and, uh, let me prepare you, most of you are sitting down. It's very volatile, the impact that we've seen on quality over this period of time. Each line on this graph represents a single company that was part of this particular part of the study, and that line shows whether you've moved up in terms of your change failure rate or moved down in terms of your change failure rate. Y-you wanna go
- 5:12
down. Y-you wanna be on the bottom side of this graph. Uh, but what's interesting is if you look at some of these top lines, you see people increasing as much as two percent, which doesn't sound like much until you realize that the industry benchmark is about four percent. So that means shipping potentially fifty percent more defects than we were shipping before. So you really just wanna understand what side of this graph you're on. You wanna start measuring this stuff, right? Um, this pattern exists without AI, by the way. This has a lot to do... It's not purely causal with AI. A lot of this has to do with, like, your
- 5:42
release pipeline, your automated testing, everything that's surrounding the way that you're releasing the code. The pattern remains the same. The amplitude has changed as a result of AI. We always see this type of shift in volatility, but not usually to these extremes, and that's the AI effect that we're seeing right there. Here's a really interesting one. These are two qualitative quan-- uh, uh, quality metrics, uh, that normally are sort of in sync with one another. Code maintainability, how, uh, how maintainable do developers perceive the code to be in terms
- 6:12
of making changes to it. That number has gone up. Almost 4%, according to our data. By the way, we're looking at about 200,000 engineers in this study to gather this data. The, uh, co- the change confidence though, which traditionally we've seen closely associated with code maintainability. Hey, this code's maintainable, it's modular, it's easy to modify. I feel good about the changes that I'm pushing into production. Nope. Our change confidence number has gone down 6%. This is really an interesting tension because it
- 6:42
means like, okay, agents and assistants are making it easier for me to understand and modify the code that's in front of me, but I trust the outputs less. I'm more afraid now of breaking things than I was a year ago. So it's an interesting psychological effect that we see here. And a lot of this is because PR size is measurably increasing. This is gonna be one of the most important metrics that we look at this year. Hasn't quite doubled, but look at that trajectory, right? This is over the same period of study about a year, and we've seen PRs go from, uh,
- 7:11
around 44 lines on average per PR up to 72. Now, there's num- there's a number of reasons for this. It's not just, you know, the likelihood that the way that these models work is that they're be tr- being trained to output relatively mediocre code. I've been writing code professionally since the late '90s. I remember when it was a great mark of a developer to be able to implement the same use case with as little code as possible. Now it's like, well, no, let's get the thing out and, you know, the code that's being generated is gonna be essentially mediocre because it's law of averages. But on top of that
- 7:41
too, think about if you've got a build pipeline that takes 45 minutes or an hour to complete, and now I've got AI generating functions for me instantly. Am I gonna push four different PRs with four different functions and wait 45 minutes or an hour or whatever that is per build, or I'm just gonna cram all that code into a single PR, right? So there's multiple reasons for this, but it really bears mentioning that, like, every extra line of code is a potential bug, a potential vulnerability. It's more to review. It makes the code less portable. So we need to pay attention to this metric as
- 8:11
well. 'Cause we are also seeing this associated with perception of incremental delivery. I get to work on small incremental changes, and this is actually one of the lowest qualitative developer experience drivers that we've seen impacted over the last year. This number and this sentiment has gone down 10%. And incremental delivery is awful. Uh, is, is as we've learned, also very important for, uh, for the way that we try to deliver in a way that's easy to roll back, uh, and where we can understand and review the changes that are being pushed through.
- 8:42
Okay. Uh, what about demographic differences? Uh, what is different about AI impacts across like, you know, junior devs and stuff like that? Junior engineers are using AI the most. This really should surprise nobody. Anytime we have sort of a new leap, again, writing code professionally since the late '90s, I'm a bit of an OG, I'm on this journey too. But there's less to unlearn when you're coming into the industry, right? In many, many cases, you've already been working with this kind of stuff, like right outta school, and
- 9:12
so you're more likely to implement it, uh, when you actually start working. Now, we've seen some other trends there too. Uh, we have interesting ways of looking per use case at efficiency, uh, which I'll get into in the second talk later today. Um, one of those is looking at agent experience. Like literally we're, we're asking agents about their experience working with humans. We're trying to figure out what use cases, uh, are they working on and how many tokens did they spend. And definitely we also see a trend where the more junior
- 9:41
developers for the same use case will spend more tokens than a more senior developer on the same use case, and that's just a learning curve, right? Again, that should surprise nobody. But in terms of who's using it the most, we do see gen- uh, juniors using it the most. However, despite that additional use, we see staff plus some more senior engineers saving about the same amount of time, right? And, and, and burning less tokens. I, I wanna make sure that that's kinda clear too. So in terms of like raw time savings and the aggregate output of this, it's pretty flat between junior engineers who are using the tech more and
- 10:11
staff engineers who have an easier time spotting hallucinations, understanding the architecture around the changes they're making and things like that. Smaller companies lead in time savings. Sure. You know, the release pipelines are gonna be less complicated. You know, there's less overall complexity in the organization, uh, that's going to, uh, lead to more time savings for a company that already has less friction and has a little bit more agility to it than a, a larger company. What about measurement? How should we
- 10:41
even, even be thinking about measuring this stuff? I mean, it bears mentioning that, that measuring developer productivity and measuring developer experience was a challenge before AI, and we, we never completed that conversation before kind of throwing accelerant on this, right? So this is still a difficult thing, and I'd argue even more difficult to do now because it's confounded by several different aspects of, of AI. But we're all gonna have to answer this question this year, right? Spent 10 million or way more, if you come to my talk later today on
- 11:11
tokens, uh, where's our 10X productivity? When we're thinking about this stuff, it, it's important to bear in mind that we don't throw away what we're measuring, right? Our, our foundational developer experience and developer productivity metrics are still what matter the most, right? We wanna understand how AI is impacting these trusted metrics that we've already been looking at. Like, we wanna maybe look at cohorts of users. We maybe wanna take data from the API telemetry and things that come out of the tools and use that to
- 11:41
understand who's using what and where. But we would only wanna use those cohorts and look at them comparatively across our foundational metrics. What is this actually doing to quality? What is this actually doing to speed? What is this doing to impact to the organization and value generation? And so this is where our AI measurement framework comes from. I won't spend too much time on this. We have a big white paper available. Uh, my colleague Anton, who's standing in the back of the room in the, the white shirt, will be able to hand out some hard copies of reports that we have after the talk that go into the way that
- 12:11
we derive this methodology and how to use it and think about it. Um, it focuses on three key dimensions of measurements, so looking at utilization, our daily active users, weekly active users, that sort of thing. Uh, it looks at impact, so which metric should we be, you know, looking for moving the needle based on utilization? Uh, as our utilization goes up, what do we hope to see in terms of impact to the business? And then cost. It's a fair joke that we're 15 years after the last major hype cycle, and we're still trying to figure out cloud costs, but this stuff is getting
- 12:41
pretty expensive. We're gonna have to measure it. Um, so you can think about this as a bit of a maturity curve as well. Most people start with utilization on the left, just figuring out what's happening in the organization with the tech, who's using it, what are they using, how often, and for what use cases. And then how do we then cross-reference those to the impact metrics, the value generation, what we really wanna understand to know whether these investments are actually working, which is what we really wanna figure out. This is just an example of how we might correlate such a thing. So you could look at two separate cohorts of users, and you could compare
- 13:11
them across metrics like PR cycle time or PR size or pushback in review, right? So lots of... uh, the way to think about this, again, is just right now separate cohorts and then look at how that plays out across these foundational productivity metrics that we've learned to trust. We also need to be able, be able to understand our platform's readiness. I like to say that in 2024 we gave everybody a coding assistant, in 2025 we started building agents. Now we're realizing that our infrastructure wasn't ready for any of this. A lot of the vendors that are
- 13:41
here are, are, are selling solutions that help you with this problem. Um, I think if you can measure your platform's AI readiness, you're gonna have a much better handle on how efficient you're gonna be spending tokens and providing context to these agents. So these are things like clear and accurate well-structured documentation, data structures with straightforward relations, manageable modular code, reliable local CI, and non-flaky test suites. If, by the way, any of this sounds familiar to you, it's because we used to just call this good developer experience, right? But it turns out that what's good for
- 14:10
humans is also good for agents, so we may finally paradoxically be making those investments that we should've been making over the last couple of decades. We also wanna understand the effectiveness of the way that we work with AI, and this is one way of doing that. What you're actually seeing here is feedback, qualitative feedback coming back from agents telling us where they ran into issues with steering with the human, the c- the context that was provided for them, the f- the feedback cycles that they had to go through. So we're actually getting that data from the agents to try to figure out how effective our
- 14:40
teams are at working with AI, and we can break this down into use cases, which is very important too. We wanna understand which use cases are giving us the most efficient token spend, and again, moving the needle for value creation and productivity within the business. Finally, the fifth trend, integrating across the SDLC and PDLC. Why would we wanna do this? Because code generation was never the bottleneck in the first place, right? Even if engineers are getting, like, 100% accurate instant code coming from the models, which they are not, you would
- 15:10
still only be attacking anywhere from maybe 14 to 16% of the overall value stream. So we need to think beyond that, uh, especially because right now the productivity gains that we a- are actually seeing are more modest than expected. Despite additional PR throughput, uh, increases, our median increase in the study that we did from November to, to 2024 to February this year only found about a median 7.7% increase in this velocity metric, a 13% average. But even our top performers were in the 70%
- 15:40
range. Nobody hit 2X. Nobody hit 5X. Nobody hit 10X. And that's because our time savings that come out of AI are still being outweighed by other non-AI factors within our organization, right? We might be saving a lot of time with AI, and that's great, but if we're still eclipsing those time savings because of meeting heavy days and context switching and other sources of interruption, the cumulative effects of the dev environment and friction in the way that we create software around the code, then w- we're not attacking the bottleneck. And as Eli Goldratt from
- 16:10
the theory of constraints and the goal and, uh, inspi- inspiration for the Phoenix Project would tell us that an hour saved on something that isn't the bottleneck is worthless. And so we need to find the bottleneck and fix the bottleneck, and that's what a lot of organizations are doing. Morgan Stanley has been really public about their creation of the DevGen AI agent which is an agent for interpreting legacy code like, uh, mainframe natural and COBOL, and I hate to admit Perl. I've written a lot of Perl in my career, but it's legacy now. The agent actually creates PRDs to hand to
- 16:40
engineers to get rid of that reverse engineering step, and it's saving Morgan Stanley about 300,000 hours a year right now. Zapier's one of my favorite stories in AI right now. They have a whole agent ecosystem that they're using to deal with a lot of administrative tasks and overhead, like daily stand-ups and things like that. They've reduced their stand-ups by using a- agent summaries and things like that in this ecosystem that they've built to two times a week as opposed to five times a week. They're successfully onboarding engineers within about a two-week period, which is really
- 17:10
fast. Our industry benchmark for that is over a month usually. But my favorite part about this story is that they're seeing about a 15% additional value creation per engineer, so they are hiring more than they have in the history of their entire company 'cause they realize the truth here, that they're getting more capacity per single engineer. They're getting a better return on investment per single engineer. And hiring more engineers will improve their competitive edge. This is the right attitude. This is a throughput story. This is an increased innovative capacity story. This is not
- 17:40
a headcount replacement story. Fair, uh, they're automating about 3,000 of their code reviews a week triggered off of a pull request. And, um, they look at superficial stuff. They still need humans in the loop, but the initial superficial, uh, uh, review that comes from the agent is all part of the system of record. It stays in the PR comments. And so the next engineer to actually perform the review will have an idea of what the agent's already looked at, so it can save some time. Finally, Spotify, if anybody remembers the Spotify model from DevOps, they were sort of the North Star
- 18:10
for DevOps. Uh, they built an agent for SREs that will gather remediation steps from runbooks and information and context about the incident and put all that together into SRE communication channels so that when an incident occurs, the SRE has, uh, immediate context and doesn't have to go through multiple minutes of discovery, uh, to figure out how to solve the issue. So if you wanna dive deeper into this data, our Q2 report is coming out in just a few weeks. Uh, but the Q1 report still has a lot of tidbits from this session in it.
- 18:40
And if you subscribe to our newsletter or you come back to the website, by mid-ish end of July we'll have our Q2 report. And if you wanna see a preview of r- that report, that's the talk that I'm giving later this afternoon. Thanks, everybody, for your time. I appreciate it.