← All AI Engineer talks

AI Engineer World's Fair 2024

The ROI of AI: Why You Need Eval Frameworks

About this talk

Sourcegraph CTO and co-founder Beyang Liu explains how engineering leaders can evaluate the return on AI coding assistants beyond intuition. Using Cody and enterprise customer examples, he describes the importance of code search and relevant codebase context, discusses rigorous evaluation approaches and engineering KPIs, and examines developer-productivity measurement alongside lessons from The Mythical Man-Month.

Chapters

  1. 0:00Introduction and the challenge of measuring AI ROI
  2. 1:38Sourcegraph, Cody, and codebase-aware context retrieval
  3. 3:50Developer value and rigorous evaluation frameworks
  4. 16:32Engineering KPIs and productivity measurement
  5. 22:52The Mythical Man-Month and closing audience exchange

Talk transcript

  1. 0:00

    [upbeat music] I want to introduce our first speaker, uh, Beyang Liu, um, CTO of Sourcegraph, and I won't spoil the topic of the talk, but it's gonna be a good one.

  2. 0:19

    Thank you very much.

  3. 0:20

    Awesome. [clapping] Thank you, Peter. How's everyone doing this morning?

  4. 0:25

    Good. Yeah? Everyone, uh, awake, bright and early. Thanks for coming out. Almost awake. Awesome. Um, so before I, uh, dive into the talk here, I just wanted to get a sense of, you know, who we all have in the room.

  5. 0:36

    So, um, you know, who here is, like, a head of engineering or a VP of engineering?

  6. 0:42

    Okay, a good number of you. Who here is just a, a, you know, IC dev, interested in kind of like evaluating how things are going? Okay. And then w- what do you-- what, what do the rest of you do?

  7. 0:54

    Just shout out, uh, your roles. Anyone? Anyone? Middle management. Middle management. Okay, cool. [laughing] And who here has a really, you know, quantitative, very precise, thought-through, uh, evaluation framework for measuring the ROI of AI tools?

  8. 1:10

    Okay, one hand in the back. What company are you from, sir?

  9. 1:16

    Broadsignals. Broadsignals. Broadsignals. Okay, cool. So we got one person in the back. And then who here is sort of like, "We kind of are evaluating it, but it's really kinda like vibes at this point"?

  10. 1:26

    Anyone? Anyone brave enough to... Okay, cool. So you're in the right place. Now, who am I? Why am I qualified to talk on this topic? So I'm the CTO and co-founder of a company called Sourcegraph.

  11. 1:38

    Uh, we're a developer tools company. Uh, if you haven't heard of us, we have two products. One is a code search engine. So we started the company because we were developers ourselves, my co-fou- co-founder and I, and we were really tired of the slog of diving through large, complex codebases and trying to make sense of what was

  12. 1:53

    happening in them. The other product that we have is an AI coding assistant. Uh, so what this, uh, product does is it's essentially...

  13. 2:04

    Oh, sorry. I should mention that we have great adoption among really great companies. So we started this company ten years ago, uh, to solve this problem of tackling, uh, understanding code in large codebases, and today we're very fortunate to have, uh, you know, customers that range from early stage startups all the way through to the Fortune 500

  14. 2:22

    and even some government agencies. So what's the tie-in to AI ROI? So about two years ago, we released a new product called Cody, which is, uh, an AI coding assistant that ties into our code search engine.

  15. 2:36

    So you can kind of think of it as like a Perplexity for code, whereas, you know, your kinda vanilla, run-of-the-mill AI coding assistant, uh, only uses very local context, only has access to kinda like the open files in your editor.

  16. 2:49

    We spent the past ten years building this great code search engine that, uh, is really good at surfacing relevant code snippets from across your codebase. And it just so turns out that, like, that's a great superpower for AI to have, right?

  17. 3:01

    Um, you know, for, for those of us that have started using Perplexity, we can kinda see the appeal. And a big piece of the puzzle is not just the language model itself, but the ability to fetch and rank relevant context from, you know, a whole universe of data and information.

  18. 3:16

    And so we had to solve this problem, uh, in order to sell Cody into the likes of, uh, 1Password, uh, Palo Alto Networks, and, uh, Leidos, which is a big government contractor.

  19. 3:28

    If you flew in here from, uh, somewhere else, you probably entered through one of their security machines. We sold Cody to all these organizations, and so each one of these organizations has kind of like a different way of measuring ROI.

  20. 3:39

    They have different frameworks that they apply. And so we had to answer that question in, uh, a multitude of ways, and that's what I'm here to talk about.

  21. 3:50

    Okay, so I wanna start out with how I would describe the value prop of AI to someone who's actually using it, to a developer.

  22. 3:58

    So coding without AI, I think, you know, we've all felt this before if you've ever written a line of code. Uh, you start out by asking yourself, like, "Oh, this task, it's straightforward.

  23. 4:08

    It should be easy. Let me just go build this feature." I should mention, I poached these slides, uh, from, um, a, a director of engineering at Palo Alto Networks.

  24. 4:18

    He's actually giving another talk, uh, at this conference, Gunjan Patel. I thought he did an excellent job of describing the value prop that he was solving for as a director of engineering when they were purchasing, uh, a coding AI system.

  25. 4:29

    So this is the way he described it. You think it should be straightforward, but then there's all these side quests that you end up going on as a developer.

  26. 4:37

    It's like, "Uh-oh." Like, "I gotta go install a bunch of dependencies," or maybe, you know, there is this, uh, UI component that I have to go and figure out, this framework that I need to learn.

  27. 4:46

    So the gaps appear, and then without AI, bridging the gaps, uh, takes both time and focus. You kinda have to, you know, spin off a, a side process or go on a little mini quest, uh, and then, you know, thirty minutes and two cups of coffee later you're like, "Okay, I got it.

  28. 5:02

    I filled the gap, but what was I doing again?" And so AI helps bridge those gaps. It helps solve this problem. It helps, uh, developers really stay in flow and stay, uh, kind of like cognizant of the high level of what they're trying to accomplish.

  29. 5:18

    And so what this means is that more and more developers can actually do the thing that, that we wanna do, which is build an amazing feature and deliver an amazing experience, um, instead of giving up.

  30. 5:26

    Now, the question is: How do we measure this in a way that we can, uh, demonstrate the business impact of what this does to the rest of the organization?

  31. 5:36

    And so the answer is beans. Okay, who here drinks coffee?

  32. 5:43

    Okay, cool. Do you know the difference between a really great, uh, coffee bean and, you know, your kinda run-of-the-mill Folgers or, you know, the thing that you buy at the supermarket? [clears throat]

  33. 5:54

    Okay. Turns out we're all in the bean business. We think we're in the software business, but We're really selling beans, in a way. So in every company, there's someone, uh, let's call him Bob, who grows the beans, essentially the developer.

  34. 6:09

    Uh, there's another person, let's call, uh, him Pat, who sells the beans. That's your kinda CRO or sales lead. And then you have Alice on the side, who's kinda your CFO, uh, or, or CEO.

  35. 6:20

    Uh, and Alice has gotta count the beans. At the end of the day, you know, uh, [laughs]

  36. 6:24

    not to diss the finance people if there are any in the room, but, like, it's all about counting the beans and seeing how they add up. Now, I don't know if any of you have been paying attention, but in the past two years, the bean business has been revolutionized.

  37. 6:36

    This thing called AI has appeared. And so what does the bean business look like now?

  38. 6:42

    Well, Bob grows the beans with AI, and Pat is selling beans, but with AI. And then Alice is on the side being like, "Well, I'm counting all the beans, and where's the ROI?"

  39. 6:53

    And so this is the answer that, you know, B- basically Bob has to answer. Pat has kinda got it easy because Pat's job is just selling the beans. That's a much more quantifiable, uh, task.

  40. 7:04

    Um, those of us that are involved in software engineering and product development, it's a bit harder to measure. So there's tension in the bean shop, you know? Uh, Alice is asking Bob, you know, "How many more beans are we growing now with AI, Bob?"

  41. 7:16

    And Bob's like, "Well, it's complicated, Alice. Not all beans are the same. You know, there's some good beans, and there's some very bad beans. We're making our beans better."

  42. 7:23

    And then Alice is like, "Well, okay, the bean AI tool costs money, and we gotta measure its impact somehow."

  43. 7:32

    Anyone feel that tension? Anyone have this kinda, like, conversation? We, we've talked to a lot of heads of engineering who ha- see this tension very real with other parts of the org, specifically between finance and engineering.

  44. 7:45

    And I think the core of the problem is that measuring AI ROI for functions where the work is not directly quantifiable through a number is what I like to call NP-hard.

  45. 7:56

    So how many people are familiar with the term NP-hard here? Okay. Cool. We're all pretty technical. So NP-hard basically means if you have a tough challenge, uh, if you, if you have a problem, and you can basically reduce it to a class of very s- uh, difficult problems, it probably means your problem is not solvable.

  46. 8:14

    And measuring AI R- ROI reduces to measuring developer productivity or the productivity of whatever class of knowledge worker, uh, that you're managing. Uh, and so that implies if you can measure the AI ROI precisely, you can also measure developer productivity.

  47. 8:28

    And who here knows how to measure developer productivity?

  48. 8:33

    It's kind of an open question, right? So, uh, using the logic of your standard reduction proof, uh, this problem is intractable. So that's the end of my talk. I'm just here to tell you that this problem is intractable, and we should give up, right?

  49. 8:47

    Well, in the real world, we often find tractable solutions, uh, to intractable problems. And so, uh, what the meat of this talk is, is really sharing a set of evaluation frameworks that we've presented to different customers.

  50. 8:59

    Uh, not all of these are used by, uh, you know, any given customer. Um, but I wanted to give kind of like a, a sampling of the conversations that we've had.

  51. 9:07

    Uh, and hopefully there'll be some time for Q&A at the end where we can kind of talk through this and, and, uh, see what other people are doing.

  52. 9:15

    Okay, so framework number one, uh, is the famous roles eliminated, uh, framework. So this question gets asked a lot these days, especially, you know, on social media. Like, AI is here to take your job.

  53. 9:29

    So how does this framework work? Well, in the classic framing, you buy the tool, the, the labor-saving tool. You observe ... You know, you're in the bean business or the, the widget business.

  54. 9:37

    You observe that this tool yields an X percent increase, uh, in your capacity to build widgets. And then you can cut your workforce to meet the demands for whatever widgets you're selling.

  55. 9:48

    Now, in practice, we have not encountered this framework at all in the realm of software development. Um, we do see it more prevalent in other kinda business units, you know, things that are more viewed as, like, cost centers, like, uh, support and things like that, um, especially like, uh, consumer-facing customer support.

  56. 10:07

    But for software engineering, uh, for whatever reason, we haven't encountered this yet in any of our customers. And we think that the reason here is that, number one, you know, if you view your org as a widget factory, uh, then you're gonna prioritize outputting widgets.

  57. 10:22

    But the, the thing is that very few engineering leaders, effective engineering leaders these days view themselves as widget builders. You know, the widgets are kind of an abstraction, uh, that don't apply to the craft of software engineering.

  58. 10:33

    The other observation here is that the widgets that we're building, which is software at the end of the day, great user experiences, they're not really supply limited. So if you have, like, an extensive backlog, the question is, you know, if we made your engineers 20% more productive, would you go and, you know, do more, uh, 20% more

  59. 10:50

    of your backlog? Or you just cut down 20% of your workforce and say, like, "You know, it's fine. We don't need to get to the backlog." And for 99% of the companies out there that are building software, the answer is no.

  60. 11:01

    We wanna build a better user experience. These issues in our backlog are very important. We just can't get to them. So framework number one is kind of like, uh, talked about frequently, but in practice, we haven't really seen it as an evaluation criteria.

  61. 11:16

    Framework number two is what I like to call A/B testing, uh, velocity. So how this works is you basically segment off, uh, your organization into two groups, uh, the test group and the control group.

  62. 11:29

    And then, uh, you say group one gets the tool and group two does not. And then you go through your standard planning process. Most people, uh, as part of that planning process, what you do is you go and estimate the time that it will take to resolve, uh, certain issues.

  63. 11:44

    So, you know, how long is this feature gonna take to build? How, how long is it gonna, uh, take to work through these bug backlogs? And then because you've divided the groups into two now, uh, you have some notion of, like, you know, how well you're executing against your, your timeline.

  64. 11:57

    So you basically run this A/B test. Um, so Palo Alto Networks, one of our customers, uh, ran something similar to this. Uh, and the conclusion they drew was the, the approximate timelines got accelerated 20 to 30%, uh, using Cody.

  65. 12:12

    And so this is a very kind of, like, rigorous scientific framework. Um, we see it come up now and then, uh, especially when, when companies are of a certain size and, and they, they're very thoughtful about this question.

  66. 12:24

    Um, the criticisms about this framework are no two teams are exactly the same, right? Like, if you lop off your development org, you have, you know, your dev infrastructure on this side, maybe backend, and then you have front-end teams on this side.

  67. 12:35

    It's hard sometimes to compare these, th- these different teams to each other because, uh, software development is very different, uh, in different parts of your organization. There's also confounding factors.

  68. 12:46

    You know, maybe Team X, you know, had, a, an important leader or contributor depart. Uh, maybe Team Y, uh, you know, uh, suffered a bout of COVID that, you know, uh, blew through the team or, or things like that.

  69. 12:59

    So you have to account for these things when, when make your evaluation. And this framework is also high-cost and effort. You basically have to do the subdivision. You give one group access to the tool, and then you have to run it for an extended period of time in order to gain enough confidence.

  70. 13:12

    But provided you have the resources and the time to estimate it, we think this is a pretty good framework for honestly testing the ef- efficacy of, of an AI tool.

  71. 13:24

    Okay, framework number three. Um, I call this time saved as a function of engagement. So if you have a productivity tool,

  72. 13:32

    um, using the product should make people more productive, right? So if you have a code search engine, uh, the more code searches that people do, uh, that, that probably saves them time.

  73. 13:42

    If it didn't save them time, there would be no reason why they would go to the search engine. And so in this framework, what you do is you basically go look at your product metrics, and you break down all the different time-saving actions, uh, you identify them, and then you kinda tag them with an approximate estimate of

  74. 13:58

    how much time is saved in each action. And if you wanna be conservative, you can lower bound it. You know, like, you could say, like, "A code search probably saves me two minutes."

  75. 14:07

    That's maybe, like, a, a lower bound because there's, there's definitely, like, searches where, oh my gosh, like, it saved me, like, half a day or, uh, maybe, like, a whole week of work because it prevented me from going down, uh, an unproductive rabbit hole.

  76. 14:20

    But you can lower bound it and just say, like, "Okay, we're gonna get a lower bound estimate on the total amount of time saved." And then, uh, you go and ask your vendor, "Hey, can you build some analytics and share them with me?"

  77. 14:30

    So this is something that we built for Cody, uh, you know, very fine-grained analytics to show, uh, the admins and the leaders in the org exactly what actions are being, uh, taken.

  78. 14:42

    You know, how many explicit invocations, you know, how many chats, how many questions about the codebase, how many inline code generation actions, and those all map to a certain amount of time saved.

  79. 14:52

    There's one caveat here, which is in, uh, products where you have, like, an implicit trigger, like an autocomplete, um, you can't rely purely on engagement because it's not the human opting in to engage the product, uh, each time.

  80. 15:05

    It's sort of, like, implicitly shown to you. Um, and so there you tend to go with, uh, more of an acceptance rate criteria. You don't wanna do just raw engagements because then the product could just push a bunch of, like, low quality, uh, completions to you, and that would not be a good, uh, meas- metric of, of

  81. 15:20

    time saved. So we have a lot of customers that do this. They appreciate it. One of the nice things about this is that because we're lower bounding, it makes the, the, the value of the software very clear.

  82. 15:31

    Um, so a lot of, uh, developed tools, us included, I think, like, Cody is, like, nine dollars per month, and Sourcegraph is a little bit more than that, but it's like if you back out the math of how much an hour of a developer's time is worth, you know, typically it's around, like, a hundred to two hundred

  83. 15:45

    dollars. It's like if you save, you know, a couple minutes, uh, a month, this kinda pays for itself in terms of productivity. Um, the criticisms of this framework is, of course, it's a lower bound, so you're not, uh, fully assessing the value.

  84. 15:59

    If you go back to that picture I showed earlier of, you know, the dev journey where you're kind of bridging the gaps, uh, I think a big value of AI is actually completing tasks that hitherto or, uh, beforehand were just not completed because people got fed up, uh, or they got lazy, or they just had other things

  85. 16:15

    to do. So this doesn't capture that. Um, it doesn't account for the second order effects of the velocity boost. Um, and it's good for kinda, like, day to day, like, "Hey, is this speeding up the team?"

  86. 16:26

    But it doesn't capture some of the impact on key initiatives of the company.

  87. 16:32

    So that leads to the fourth evaluation framework. So a lot of our customers track certain KPIs that they th- they think are correlated with engineering quality and business impact.

  88. 16:42

    So lines of code generated. Does anyone here think lines of code generated is a good metric of developer productivity?

  89. 16:50

    Okay, I think i- in twenty twenty-four, we can all say that it's not. We have seen this resurfaced in the context of measuring ROI of AI because, uh, it's almost like we've forgotten all the lessons that we learned about human developer productivity, and with AI, uh, generated tools, now it's like, oh, like, you generated, like, you know,

  90. 17:08

    hundreds of lines of code for, uh, uh, a developer in a day. And so we, we noted that we were actually losing some deals on lines of code generated, and when we actually went and looked at the product experience, we were like...

  91. 17:19

    At first, we were like, "Hey, you know, maybe we should... There's a product improvement here that we should be making because people aren't accepting as many lines generated by Cody."

  92. 17:26

    But when we dug into this, we were like, oh, like, you know, the competitor's product, it's just more aggressively kinda, like, triggering and, and that's not, like, the sort of business that we wanna be.

  93. 17:35

    So more and more, we're kind of pushing our customers to not tie to generic metrics or, like, high-level metrics, but identify certain KPIs that attend to changes that you wanna make in your organization.

  94. 17:48

    So with Leidos, um, our, our big kinda government contracting, uh, customer, they identified a set of, uh, zones or actions that they felt were really important. Uh, they wanted to reduce the amount of time spent answering questions, spent bugging teammates, and m- uh, spend more developer time in these areas that they identified as value add.

  95. 18:09

    Uh, three things mainly, building features, writing unit tests, and reviewing code. And so that's what we tracked for, for their evaluation period.

  96. 18:19

    A fifth framework is impact on key initiatives. So this is the kinda like map your product to OKRs, uh, framework. Uh, and so there are a couple of companies where they're in the midst of a big code migration, like they're trying to migrate from, you know, Cobalt to Java or, uh, maybe, you know, from, uh, React to

  97. 18:37

    Svelte or the, or the latest JS framework. And these are kinda like top-level goals that the VP of engineering, um, really cares about. And so if you have a product that accelerates progress, uh, towards this, then the ROI is really just what's the value of bringing that forward by, you know, X number of months or, in some

  98. 18:55

    cases, X number of years, or making it possible at all.

  99. 18:59

    How do you measure how much you pulled it forward? Uh, that's a good question. So the question was how do you measure how much you pulled it forward? It is really a kind of judgment call with the engineering leader at that point.

  100. 19:10

    Um, by the time they have this conversation with us, they've typically already started it or have had a few of these under their belt, and they have to have an idea of the pain.

  101. 19:19

    Uh, and then they can assess kinda the shape of the product and, and the things that we do and estimate how much quicker it'll be done. We also have case studies demonstrating, like, hey, this thing that used to take, you know, a year or longer, we squished it into the span of, you know, a couple months.

  102. 19:35

    Okay. And then the last framework is what I'll call survey. So this sounds like the least rigorous framework, but I think it's still highly valuable. Basically, it's just run a pilot, uh, have your developers use it.

  103. 19:45

    You can compare against another tool, and then at the end of it, just ask your developers, uh, you know, which one was best. Um, more and more nowadays, we don't see this in an unbounded fashion.

  104. 19:56

    Like, you know, in the kinda like twenty twenty-one ZRTP period, uh, people were just like, you know, "Whatever makes the developers happy, let's just go buy that." Nowadays, we see it in more of a bounded fashion, which is, um, you know, a lot of orgs, uh, you know, uh, allocate some part of their budget toward investing in

  105. 20:14

    developer productivity and to develop productivity tools. So that chart over there shows kind of the range of orgs surveyed in, uh, a, a survey run by the Pragmatic Engineer, Engineer newsletter, and it typically ranges somewhere between five and twenty five percent.

  106. 20:28

    And so within that budget allocation, you have a certain amount of budget to allocate to tools. Um, then you basically say, s- uh, subject to that constraint, let me go ask my developers what tools they want the most.

  107. 20:42

    Okay. So I'm basically out of time, but hopefully this gave you a kind of like a sampling of, of the flavors of different frameworks that are involved. This is something that we've had to work through through a lot of customers, ranging from very small startups to very large, uh, companies in the Fortune five hundred.

  108. 20:56

    I just wanna say, you know, be skeptical of anyone saying P equals NP, of saying they, they have, uh, a precise way to measure AI ROI or developer productivity.

  109. 21:04

    No framework is perfect. The most important thing I think is define clear success criteria. Um, that's something that both you, your internal teams, and stakeholders will appreciate, and also the vendor because they know what success looks like.

  110. 21:17

    And then the last kinda final note here is productivity tools are often bottoms up, but we've actually found that top-down mandates with AI can sometimes help because developers can be a little bit of a skeptical crowd.

  111. 21:28

    But if you believe firmly that this is where the future is going and that people have to update their skill set to make productive use of LMS and AI, uh, to code more productively, we've actually seen success in this case, where a CEO basically says, "We're adopting Cody.

  112. 21:42

    We're adopting code AI. Go figure out how to use this. It's not gonna be perfect, but it's something that is gonna play out over the next decade."

  113. 21:51

    And then, sorry, one final thought. Um, as we move towards more automation, um, I like to think of two pictures in mind. So one picture is the graph on the right, which is a kind of like a landscape of, uh, code AI tools.

  114. 22:04

    So on the left-hand side, you have the kinda like inline completions, you know, very basic, completing the next tokens that you're typing. And then on the far right, you have kinda like the fully automated offline agents.

  115. 22:15

    And we're trying to make all these solutions more reliable, right? 'Cause like, uh, generative AI is sort of inherently unreliable. We're trying to make it more general and, and, uh, you know, uh, ma- make it productive in, in more languages and in more scenarios.

  116. 22:28

    And we actually think that the, the, the path there is to go from the left-hand side to the right-hand side, you know, not ju- jumping straight to the full automation because that's a very difficult problem.

  117. 22:39

    So we think the next phase of evolution for us is going from kinda like these inline code completion, uh, scenarios to more what we call online agents that live in your editor but can still react to human feedback and, uh, guidance.

  118. 22:52

    And then the second picture is, you know, the Mythical Man-Month. I think a lot of people are familiar with this classic work. It talks about the classic fallacies with respect to, uh, developer productivity that a lot of companies make.

  119. 23:04

    Um, one question I would pose to all of you is, have the fundamentals really changed? You know, as we have more, quote, unquote, "AI developers," uh, should we measure them by the same evaluation criteria?

  120. 23:15

    And I would ch-- the, the, the challenge question I would pose to all of you is, you know, I think the lessons from the Mythical Man-Month is you prefer to have one very smart engineer who's highly productive over ten, maybe even a hundred mediocre developers who are kind of productive but need a lot of guidance.

  121. 23:32

    And so as AI automation increases, the, the question is, do you want a hundred mediocre AI developers, or do you want a hundred X lever for, uh, the human developers, the really smart people who are gonna craft the user experience?

  122. 23:47

    All right, that's it for me. Um, yeah, if you wanna check out Cody, [audience applauding] that's the URL. And, uh, that's my contact info if you wanna reach out later. Do we have time for questions or...

  123. 23:58

    Um, we have time for perhaps one or two questions.

  124. 24:00

    One or two questions?

  125. 24:00

    I will run around with the microphone-

  126. 24:02

    Okay

  127. 24:02

    ... if anyone, just so we can catch it on audio. Cool.

  128. 24:09

    Thanks. Thanks for Cody. You, you are the best guys. I tried them all. I'm staying with you. [laughs]

  129. 24:14

    Oh, thanks.

  130. 24:14

    Uh, you've got really good insight what's happening now with software development. So what is your intuition, uh, are we going to get away from coding, and the code would become a boilerplate, and we move to meta programming, or it will be still code, uh, as a main, uh, s- output of the senior engineer?

  131. 24:35

    Yeah, that's a really good question. The, the usage patterns that we're observing now is more and more code is being written through natural language prompts. Like, I had an experience just on Friday, actually, where, like, Cody wrote eig- eighty percent of the code because I just kept asking it to write different functions that I wanted.

  132. 24:51

    Uh, and that was nice because I was thinking at, like, the function level rather than the kinda like line-by-line code level. Yeah. And it was nice 'cause it allowed me to stay in flow.

  133. 24:59

    But at the same time, like, the output was still code. And I, I really do think that, like, we're never gonna see a full replacement of code by natural language because it...

  134. 25:07

    Code is nice because you can describe what you want very precisely, and that precision is important as a source of truth for what the software actually does.

  135. 25:15

    Okay. I think that's probably it now.

  136. 25:17

    Okay.

  137. 25:17

    But thank you very much.

  138. 25:18

    Thank you.

  139. 25:18

    Thank you. [audience applauding] [upbeat music]