← All AI Engineer talks

AI Engineer Code 2025

Leadership in AI-Assisted Engineering

Justin Reock· DX18:11

About this talk

Justin Reock of DX explains how engineering leaders can move beyond AI adoption mandates and perceived productivity toward evidence-based measurement of software delivery, code quality, and business impact. Drawing on METR, DORA, DX data, Project Aristotle, and SWE-bench, he recommends clear policies, experimentation time, psychological safety, transparent metrics, and applying AI across SDLC bottlenecks rather than focusing narrowly on code generation.

Chapters

  1. 0:12Executive AI strategy and conflicting productivity evidence
  2. 2:02DX metrics, delivery quality, and ineffective adoption mandates
  3. 4:33AI policies, SDLC integration, and psychological safety
  4. 7:11Developer experience, system-level productivity, and model settings
  5. 14:59Theory of Constraints and measuring AI adoption outcomes

Talk transcript

  1. 0:12

    [upbeat electronic music] Thanks for joining me in one of the later day sessions. Looks like we, we, we kept a lot of people here. This is a nice full room. Great to see it.

  2. 0:26

    We're gonna go through a lot of content in a short amount of time, so I'm gonna get right into it. If you wanna get deeper into any of this stuff, we have published this, uh, AI strategy playbook for senior executives, and, uh, a lot of the content that I'm gonna go through, I'm not gonna have time to

  3. 0:39

    get quite as deep, but this is just a nice PDF copy that you can come and refer to later. If you miss this QR code, don't worry, I'll show it again, uh, at the end.

  4. 0:47

    So what is the, uh, current impact of GenAI?

  5. 0:51

    Nobody knows, right? We've got Google on the one hand telling us that everyone's 10% more productive. That's interesting. Now they're Google, they were already pretty productive to begin with.

  6. 0:59

    But we have this sort of now infamous Meter, METR study, which has some flaws in the way that study was put together, that showed actually a 19% decrease in productivity using coding assistance.

  7. 1:10

    So there's a lot of volatility, a lot of variability. Uh, what was really interesting about this study, even though I, I mentioned there were some flaws, um, but every engineer that took part in this study felt more productive, but then the data actually bore out that they were less productive.

  8. 1:26

    Kinda interesting, right? We've got this induced flow, uh, that makes us feel really good about what we're doing. So we need to address this. DORA has put out some really good research on this too, but this is based on industry averages.

  9. 1:37

    This is impact based on what do we look at when we see a large sample and an average of how certain factors are being impacted by, in this case, 25% increase in AI adoption.

  10. 1:48

    We see these modest but positive leaning indicators, 7.5% increase in documentation quality and, uh, increase in code quality by about 3.4%. At least that's not leaning in the other direction, right?

  11. 2:02

    And when we started digging through some of DX's data, we have... You know, we're the developer productivity measurement company. We have lots of aggregate data that we can look at with this.

  12. 2:10

    We found the same thing when we looked at averages. We see about a 2.6% increase in overall, uh, change confidence, which is a, a percentage of people who answered positively that they feel confident in the changes that they're putting into production.

  13. 2:24

    Uh, similar positive leaning average when we looked at code maintainability, another qualitative metric, a .11% reduction in change failure rate, uh, which when you think about the industry benchmark being 4%, it's not insignificant.

  14. 2:38

    But this is not the full story because this is what we saw when we broke the same studies down per company. Every company here is a... Every, every bar represents a company, right?

  15. 2:49

    We have some that are seeing 20% increases in change confidence while others are seeing 20% decreases. We're seeing extreme volatility, which is why these averages look so innocuous, but they're belying the greater story of variability.

  16. 3:04

    See the same thing with code maintainability, the same thing with change failure rate. So this is a 2% increase in change failure rate up here at the top. Again, with an industry benchmark of 4%, that means shipping as much as 50% more defects than we were shipping before, right?

  17. 3:21

    We wanna make sure we're on the lower end of this, but how? Like, what should we be doing? Well, we found some patterns here. We see that some organizations are seeing positive impacts to KPIs, but others are struggling with adoption and even seeing some of these negative impacts.

  18. 3:35

    Top-down mandates are not working, right? Driving towards, "Oh, we must have 100% adoption of AI." Great. I will update my ReadMe file every morning, and I will be compliant, right?

  19. 3:44

    We're not actually moving the needle anywhere when we do that. We also find that lack of education and enablement, uh, has a, a big impact on sort of negatively impacting this.

  20. 3:54

    Some organizations just turn on the tech and expect it to just start working and everybody to know the best ways to use it. Uh, and a difficulty measuring the impact or even knowing what we should be measuring, like what metrics should, should we be looking at?

  21. 4:07

    You know, does utilization really tell us much about the full story of GenAI impact? This is another graph from DORA. Uh, this is a, a Bayesian, uh, posterior distribution, which is an interesting way of representing data.

  22. 4:20

    Basically, you want your mass to be on the yellow side of this line, uh, the, the r- uh, the right side of this line for the audience, yeah. And you want a sharp peak, which is telling you that we're pretty confident that this initiative will have this impact.

  23. 4:33

    And if we look at some of the top line initiatives here, these are things like clear AI policies. All right? We wanna make sure we have that. We want time to learn, not just giving people materials, but actually giving them space to experiment, right?

  24. 4:46

    Um, and so these types of factors are the ones that seem to be moving the needle the most. So we're gonna go over some quick tips on how we can do all of these things, and again, the guide will go deeper into this.

  25. 4:57

    We wanna integrate across the SDLC. All right? For most organizations, writing code has never been the bottleneck, right? We can inc- uh, we can increase productivity a bit by helping with code completion, but our, our biggest bottlenecks are elsewhere within the SDLC.

  26. 5:12

    There's a lot more to creating software than just writing code. We wanna unblock usage. We can't just say, "Well, we're worried about data exfiltration, so we can't try this thing."

  27. 5:19

    Like, no, get creative about it. We've got really good infrastructure out there now like Bedrock and Fireworks AI that can let us run powerful models in safe spaces. We have to have open discussions about these metrics.

  28. 5:32

    We need to evangelize the wins, and we need to let our engineers know why we're gathering metrics and data. What is it that we're trying to improve? We have to reduce the fear of AI, right?

  29. 5:42

    We have to make sure that people understand that this is not a technology that is ready to replace engineers. This is a te- a technology that's really good at augmenting engineers and increasing the throughput of our business.

  30. 5:55

    We have to establish better compliance and trust, and we need to tie this stuff to employee success. These are new skill sets. AI's not coming for your job, but somebody really good at AI might take your job.

  31. 6:07

    And so as leaders, we have the opportunity to help our employees become more successful with this technology.

  32. 6:12

    So how do we reduce the fear? Well, first of all, why do we need to do this? Well, there's a lot of good reasons, but I love to point to Google's Project Aristotle.

  33. 6:20

    This was a 2012 study where Google wanted to figure out what are the characteristics of highly performant teams. Uh, they thought that the recipe was just gonna be what Google had, this combination of high performers, experienced managers, and basically unlimited resources, and they were dead wrong.

  34. 6:36

    Overwhelmingly, the biggest indicator of productivity was psychological safety, okay? And so that very much applies now. We also have data, like this is Swebench. I'm sure a lot of you have seen this.

  35. 6:47

    And there are some impressive benchmarks, in that the agents can do, like, a third of the things they're asked to do without any human intervention. That means that they're not able to do two-thirds of them, right? [laughs]

  36. 6:58

    Again, we are augmenting. We're not replacing. We're not ready. We may never be ready. So we need to be very transparent with what we're doing. We need to set very clear intents why, you know, are we, uh, using this, to, to augment, not to replace.

  37. 7:11

    We need to be proactive in the way that we communicate that and not just wait for people to get upset and possibly scared. We need to say, "No, we are here to help you, to give you a better developer experience, and to increase the throughput of the business."

  38. 7:24

    And again, we have to have these discussions about metrics. Now, what metrics? What should we be looking at? Well, DX again, developer experience and productivity measurement company. Um, there are two sort of classes of metrics that we can be looking at, really two levers that matter here, and that's speed and quality, right?

  39. 7:42

    We want to increase PR throughput. We want to increase our velocity, but not by just creating a bunch of slop that's gonna give us a bunch of tech debt later that we're gonna have to deal with.

  40. 7:51

    Then we'll just kick the bottleneck down the road if we do that, right? So we wanna be looking at things like change failure rate, our overall perception of quality, change confidence, maintainability.

  41. 8:01

    And we have three types of metrics that we can be looking at here. We have our telemetry metrics. These are the things coming out of the API, and they're good for some stuff, but they're not always accurate, right?

  42. 8:12

    We know, like, accept versus suggest was kind of, like, all the rage until we realized that engineers need to click accept in the IDE in order for the API to know about it.

  43. 8:22

    Even if they do click accept, who's to say they didn't just go back and rewrite every line that was suggested, right? So that's providing us some context, but we also need to do some experience sampling.

  44. 8:30

    We need to, like, for instance, add a new field to a PR form that says, "I used AI to generate this PR," or "I enjoyed using AI to generate this PR," and get some data that way.

  45. 8:42

    And then self-reported data or survey data. We are big on surveys, but let me underscore, we're big on effective surveys. Ninety percent-plus participation rates engineered against questions that treat developer experience as a systems problem, not a people problem, because that's what it is.

  46. 8:59

    W. Edwards Deming, "Ninety to ninety-five percent of the productivity output of an organization is determined by the system and not the worker." Okay? So foundational developer experience and developer productivity metrics still matter the most, right?

  47. 9:13

    Our AI metrics, like utilization and things, are telling us what's happening with the tech, but these core metrics that we've been able to trust are telling us whether these initiatives are actually working, right?

  48. 9:24

    Are we actually moving the needle and having the outcomes that we wanna see? So top companies are looking at different things, right? We are seeing, like, adoption metrics coming out of Microsoft.

  49. 9:33

    They've also got this great metric called a bad developer day. I'm not gonna go into it, but there's a really good white paper that shows, like, all the different telemetry that they can look at to determine what makes a bad developer day.

  50. 9:44

    Dropbox is looking at similar stuff, adoption, like weekly active users, daily active users, that sort of thing, but also looking at quality metrics like change failure rate. And Booking is looking at similar stuff as well.

  51. 9:55

    And so we built a framework around this. We were first to market with what we call our DX AI Measurement Framework, and this is very much inspired by things like DORA, SPACE framework, DevX, just like our core four metric set, which you can ask me about later.

  52. 10:09

    Uh, and we take these metrics and we, uh, normalize them into these three dimensions of utilization, impact, and cost. And you can kinda think about this as a maturity curve too.

  53. 10:21

    A lot of people start just figuring out, okay, what's happening? Who's using the tech? What's the percentage of pull requests that we're getting that are AI-assisted, maybe through experience sampling?

  54. 10:30

    How many tasks are being assigned to agents? But then we can mature that perspective a little bit, and we can correlate that utilization to impact. What is this actually doing to velocity?

  55. 10:40

    What is this actually doing to quality? And this is when we start getting more mature in our picture of our impact. And then finally, cost. Although I like to joke that we're fifteen years past the last hype cycle, which was cloud, and we still have new companies spinning up that are teaching us how to understand and optimize

  56. 10:54

    our cloud costs, so we will see if we get there. Although I also hear horror stories about people burning through two thousand tokens a-- two thousand dollars worth of tokens a day, so we probably do need to hit that as well.

  57. 11:05

    What about compliance and trust? What can we do to ensure that the output, uh, that, that's being generated is something that can be trusted by our engineers? We have a lot of levers to pull here, but one of the ones that I like to talk about is setting up a feedback loop for our system prompts.

  58. 11:21

    So these could be called system prompts, cursor rules, agent markdown. Pretty much all of the mainstream solutions have something like this, where you can go and provide a set of rules, uh, to control how these models behave.

  59. 11:34

    Uh, and I won't get too much into the technical details here. We have an example where, like, the, uh, models have been providing outdated Spring Boot, uh, stuff. We want Spring Boot 3.

  60. 11:43

    It's, it's been sending us Spring Boot 2 stuff. The big takeaway here is to have the feedback loop. Have a gatekeeper, right? Have somebody or a group in the organization that can receive this feedback, that understand how to maintain and continuously improve these system prompts, right?

  61. 11:58

    And that way, we're always maintaining the way that these assistants or models or agents affect the whole business. It also pays to understand the way that, uh, temperature works, especially when we're building agents, right?

  62. 12:10

    We do have some control over the determinism and non-determinism of these models. Uh, again, like when a model is predicting a next token, it doesn't just have like one token, it has a matrix of tokens, and those are associated with a certain probability of that being like the right token.

  63. 12:25

    And so we have this setting called temperature, which is heat, which is entropy, which is randomness, that can control the amount of randomness involved in actually picking that token.

  64. 12:33

    This is sometimes called increasing the creativity of the model, and it's a number between zero and one. For those reasons I just mentioned, don't use zero or don't use one.

  65. 12:41

    Weird things will happen, but you want some decimal in between zero and one. When we have a lower temperature like we're seeing here, 0.0001, we give it the same task twice, and it gives us the exact same output character for character.

  66. 12:55

    When we set that temperature higher, this is an example of 0.9, I'm asking the agent to create a gradient for me, a simple task. It's giving me two relatively valid solutions.

  67. 13:06

    I did ask it for a JavaScript method, and this is the only one that's giving me a JavaScript method, but the point is, they are wildly different approaches to the same problem when I've increased the creativity of that model.

  68. 13:16

    So we need to think about like use case-wise, where should we have more creativity and where should we have more determinism, and temperature is another setting that we have that can help control this.

  69. 13:26

    You can experiment with all this using like Docker Model Runner, Ollama, LM Studio, that sort of thing. How can we tie this to better employee success? We have to provide both education and adequate time to learn.

  70. 13:38

    So we put together a study where we sampled a bunch of, uh, developers that were saving at least an hour a day, uh, uh, excuse me, an hour a week, and we asked them to stack rank their top five most valuable use cases, and we built a guide around that, a guide that effectively goes through code examples,

  71. 13:56

    prompting examples, uh, of what we determined using this sort of data approach, where we should get more reflexive about our best practice and about, uh, the use cases that we're becoming reflexive in, in, in our use of AI.

  72. 14:09

    And so that's what this guide was about. And, uh, we've had this become required reading in certain engineering groups and, uh, proud of that, and this is another way that we can help educate, but we need to give time.

  73. 14:18

    Uh, we don't have time to go through all of this. I do think it's interesting that the number one use case for this was stack trace analysis, right? So not a generative use case, actually more of an interpretive use case.

  74. 14:28

    Uh, and we see some other ones here that are not too surprising, and there's examples of each of these. What about unblocking usage? How can we make sure that we can creatively ensure that engineers can take the most advantage of this?

  75. 14:39

    Well, leverage self-hosted and private models. That's getting easier and easier to do. Partner with compliance on day one, right? Make sure that what you're doing is in line with your organization's compliance.

  76. 14:51

    You may find that you're making a lot of assumptions about things that you don't think you can do that you can actually do, right? And then think creatively around various barriers.

  77. 14:59

    Finally, how can we integrate across the SDLC? What should we think about doing there? You know, and I'm a big Eli Goldratt Theory of Constraints fan, probably have some others in the audience.

  78. 15:09

    An hour saved on something that isn't the bottleneck is worthless. And when we look at data across, in this case, almost 140,000 engineers, we find that there are definitely good, like annualized time savings with AI that are being eclipsed by sources of context switching and interruption, meeting heavy days.

  79. 15:28

    These other things that it's like, yeah, we can save time here, but we're losing so much more time over there. So find the bottleneck, fix the bottleneck, right? Morgan Stanley's been very public about their, uh, building this thing called DevGen AI that looks at a bunch of legacy code, COBOL, mainframe natural, I hate to admit Perl 'cause

  80. 15:46

    I'm an old school Perl developer, [laughs] uh, but apparently that's legacy now too, and basically creating specs, uh, for developers that can just be handed to developers to start modernizing the code without having to do all that reverse engineering, right?

  81. 15:59

    And they're saving about 300,000 hours annually right now doing this. There's a Wall Street Journal g- uh, Journal article about this, Business Insider article about it. Uh, they're very public about that.

  82. 16:09

    Zapier. Zapier should be the example for everyone. They have a whole series of bots and agents that are doing things like assisting with onboarding. They can now make engineers effective in two weeks.

  83. 16:21

    Industry benchmark on the good side is like a month. On the medium side is like 90 days. And, uh, because they're able to increase the effectiveness of the engineers that they're h- uh, that they're bringing into the organization, they realized that they should be hiring more, right?

  84. 16:38

    As opposed to trying to maintain status quo by like cutting headcount and trying to make individual engineers more productive, they said, "No, we could get more value out of a single engineer.

  85. 16:47

    We should be hiring faster than ever." And they are, and it's really increasing their competitive edge. I think that's the right attitude. Spotify's been helping out their SREs by pulling together context when incidents, uh, are detected, and then taking things like run book steps and, and other areas of context and documentation and pushing them directly into SRE

  86. 17:08

    channels so that those critical minutes of trying to get to the bottom of what's actually happening and what we should we do, do to resolve the incident, uh, they just eliminated that time, right?

  87. 17:18

    It's, it's significantly increased their MTTR. So let's get creative about areas in the SDLC that are our actual bottlenecks. All right, next steps. Uh, distribute this guide, uh, as a reference for integrating AI into the development workflows that you have.

  88. 17:32

    Uh, determine a method for measuring and evaluating GenAI impact. It's really important to make sure that we're not on the bad sides of those graphs that I showed you earlier.

  89. 17:42

    And then track and measure AI adoption and, and see how that correlates to overall impact metrics and iterate on best practices and use cases, and here's a guide again.

  90. 17:51

    Thank you so much. [upbeat music]