← All AI Engineer talks

AI Engineer Code 2025

Long Tasks and Experienced Open Source Dev Productivity

Joel Becker· METR1:15:52

About this talk

METR researcher Joel Becker leads an interactive workshop on how AI compute growth and task-completion horizons relate to real-world software-development productivity. Participants examine agent-adoption J-curves, unreliable self-reported productivity, Cursor familiarity, and the challenges of evaluating experienced developers on natural tasks in mature open-source repositories. The discussion also considers AI Village as an example of agents attempting loosely specified, real-world goals.

Chapters

  1. 0:00Compute growth, task horizons, and physical scaling constraints
  2. 6:45Developer-productivity J-curves and self-reported speedups
  3. 14:10Cursor familiarity and developer-study interpretation
  4. 37:36Natural GitHub tasks and open-source repository quality
  5. 55:58AI Village and evaluating agents on fuzzy real-world goals

Talk transcript

  1. 0:00

    [upbeat music] Here's the very simple argument.

  2. 0:23

    If you look at the top notion of compute over time, um, you know, this could be like R&D, um, spending on compute. This could be experimental compute. It could be training compute, but, you know, whatever, um, that some particular lab is, is using.

  3. 0:36

    Just like this, no surprise. If you have another chart of like, um, you know, log time horizon, let's say this, this, uh, meter measure from the, um, this figure that many of you will have seen on Twitter over time, it looks like that.

  4. 0:53

    Um, uh, you know, let, let's say that this was like not merely a coincidence, but these things were causally proportional in the sense that if, uh, if compute growth were to half, then time horizon growth were to half.

  5. 1:08

    So, you know, for, for the, for the sake of argument, let's say that, you know, starting from twenty-eight or so, um, the compute curve begins to bend like that, where this would be no growth and this would be the original growth, something, something like half.

  6. 1:22

    Then if, you know, if they were causally related, and in particular, they were causally proportional to one another, then you'd expect this to go like that. And then for some milestone that you care about, let's say here we've got, uh, one mi-- work, one worked horizon up there, one month. [background chatter]

  7. 1:43

    Then the s-delay implied in AI capabilities is potentially enormous. Um, now, like why, you know, lo-lots of people have stipulated that there might be some slowdown in compute growth.

  8. 1:53

    I, I'm not an expert in, in those forecasts, but I think, I think the prior reasons do seem like somewhat strong to me. One is like physical constraints that we might, we might hit, power constraints as mentioned, or, um, there are various other ones that, that CBOC have a report on that they say they consider all of

  9. 2:08

    which seem to not spike through twenty-thirty, but, you know, potentially, potentially could bite sometime after twenty-thirty. Um, I, I think the more likely one is just like dollars is a constraint.

  10. 2:18

    Like you can't, um, you know, large tech companies can only spend so much at a certain point. Like large nation states can only s-spend so much. Like you can't, um, uh, I guess there are some scenarios in which you, you, you can, you can continue going, but that seems to, to kind of naturally imply this slowing down.

  11. 2:35

    And then the, you know, additional point that this, this paper is trying to make is that under a

  12. 2:41

    con-- a very contestable but standard assumption from economics, um, you should in fact expect these, these two to be causally proportional. Um, I think in particular, you should expect them to be causally proportional, um, to the extent that or for the period that, uh, software-only singularity is not possible, and that's a whole another discussion for me to

  13. 3:02

    talk about. Um, but at least in this kind of somewhat business as usual,

  14. 3:08

    um, um, uh, scenario or sort of until that scenario no longer applies, um, I, I think this is, this is maybe a reasonable model and does imply some slowing of AI capabilities in the, in the near future.

  15. 3:20

    I, I have no plan for this session whatsoever. [laughs]

  16. 3:24

    That also assumes, that also assumes that we don't have a technological advance that dramatically improves capabilities relative to compute. So like, like an, like an unpredictable technological advance, right?

  17. 3:33

    Yeah, yeah. I mean, all, all predictions, you know, assume no unpredictable. [laughs] Um, yeah, and like, um, uh, you know, time horizon or, or like in general in, in AI kind of straight lines on, on log linear plots, um, have, have been a, I think, you know, a very highly underrated, um, forecasting tool.

  18. 3:53

    They've done extremely well over now many orders of magnitude. You know, I, I, I think it's reasonable to have the default expectation that the, um, log linear lines continue through like approximately the same number of orders of magnitude, except maybe if there's, you know, some significant break in the inputs.

  19. 4:08

    Yeah, o-of course, o-on the upside, there could be, um, the-there could be something quite dramatic. Software-only singularity is the first thing that comes to my mind, but, um, uh, but, you know, a-another transformer-style moment seems like another, another candidate, naturally.

  20. 4:22

    Of course. Also, one of the problems with, with testing this will be that, like, I think most of the tasks that you have available- [laughs] -available to test will, will eclipse the maximum possible amount of time that those tasks could take at some point in the evaluation set.

  21. 4:36

    Yeah. So I think, um, you know, there are some ways around this that we're working on. I'd be, I'd be excited to talk about that. They, they all feel pretty early.

  22. 4:43

    Um, uh, but, uh, yeah, you know, I, I think it's, I think it's right that, um, if, uh, if time horizons are doubling, you know, eventually you, you know, the, um, the, the doubling time is such that you can't possibly make long enough tasks [laughs] in, in the, in the relevant period.

  23. 4:59

    Yeah. It's possible also that, like, we actually hit a place where time horizon is no longer a useful measure because actually you now want time-- now you want total time to decrease.

  24. 5:07

    Like, like, like, like what you want is you want the same results at a lower time.

  25. 5:10

    Oh, um, uh, o-one-

  26. 5:14

    Like you want higher reliability at a lower time, time horizon.

  27. 5:16

    One thing to say about time horizon, um, is there's like two notions of time here. Like a, like a human time axis thing. There's, there's like calendars time axis.

  28. 5:25

    The, the time that the model's working for, I think you should like kind of approximate to zero. Um, it's, it's not actually zero. They are, they are taking actions, but they, they largely, um, do their successful work pretty, pretty early on to the extent they're going to be successful on tasks.

  29. 5:40

    Um, so, so my, my guess would be that it will continue to be the case that there's not sort of so much extra juice on that margin of, of making the models complete tasks more quickly.

  30. 5:51

    Although, uh, reliability very much so, obviously. Um.

  31. 5:56

    So most of it's like the human, like, like the iteration loop. Most of the time is spent in like the human machine iteration loop.

  32. 6:02

    Um, the humans are working without AIs, and the AIs are working without humans. So the- for the humans, I guess it's all humans-

  33. 6:09

    That's what I thought when you said like the infrastructure, like was just like-

  34. 6:12

    Yeah. Yeah. Yeah. Yeah. Yeah. Cool. Any questions on meta? I, I can go through, um, uh, some, like, upcoming things that we're, that we're excited about if, if people are excited about those things.

  35. 6:26

    Yeah. I, I, I did have one question-

  36. 6:27

    Yeah.

  37. 6:28

    -about the, uh, perceived one, like, like the time perception-

  38. 6:31

    Yeah.

  39. 6:31

    -one of the-

  40. 6:31

    Open source. Yeah, yeah. Yep.

  41. 6:35

    One, one thing I thought-- And you, you brought it up a little bit in the paper, which is, uh, you know, whether or not familiarity is a confounding factor.

  42. 6:43

    Um, although one of the things-

  43. 6:44

    With, with tools you're thinking.

  44. 6:45

    Yeah, tool familiarity is a confounding factor. And, and of course also, like you also brought up that like tool capability has dramatically changed. But, uh, there was an interesting presentation from Meta at the Developer Community Engineering Summit this year.

  45. 6:56

    Mm-hmm. Mm-hmm.

  46. 6:57

    And they had done a-- They have probably the best infrastructure for quantitative measurement of, like, developer experience in the world of any company. And they're able to tell you basically how long it actually takes to make, uh, make a PR, basically.

  47. 7:10

    They call it the asset meta, but like how much actual effort, like human time effort it took to make a PR. And what they saw was they saw a J-curve when they gave people agents.

  48. 7:20

    And that J-curve was, I don't remember how long it was, like three months or six months. And so one of the things that I also wonder is like if it would be interesting if, if, if there's a cutoff of how much familiarity the person has.

  49. 7:31

    Like, have they been using this as their full-time daily driver for a period of months? Uh, and if there's like a, a interesting cutoff that occurs once they're like a certain level of familiarity occurs.

  50. 7:42

    Yeah. I, I, I'm totally, I'm totally on board with like, not just in this case, but i-in many economically relevant outside of software engineering cases. You know, J-curve like explanations being, being a real thing.

  51. 7:53

    I'm like, yeah. Um, uh, you know, developers, not just developers, um, uh, experiment with tools. You know, you tend to be slower the first time that you're experimenting with tools.

  52. 8:02

    Um, uh, but, you know, if, if you're doing this so that you, you have some investment benefits, you know, later on, you might be, might be more proficient at the tools or in the case of AI, um, maybe you just sort of expect the models will get better.

  53. 8:15

    And so even if you don't become more proficient, it will be like the kind of thing that you want to do. You know, th-those explanations broadly, um, make sense to me.

  54. 8:23

    Um, um, I can give you some reasons why I'm skeptical. [laughs]

  55. 8:27

    Um, I think-- S-So one thing to say is, you know, we-we're, um, um... What are some things to say? Um,

  56. 8:36

    uh, as background, you know, we're continuing with this, with this work, and we'll, and we'll, and we'll see. Um, uh, you know, another thing to say is just like quantitatively, you know, difference between this and this, very large. [laughs]

  57. 8:49

    Yes.

  58. 8:49

    Um, uh, and like how much, how much is J-curve explaining? I think it's not explaining that much.

  59. 8:53

    Well, I'm just find that become like-

  60. 8:54

    Yeah.

  61. 8:54

    We see this over and over actually in software engineering studies, that the one question you can't ask people in a survey is how long did a task take-

  62. 9:02

    Mm-hmm.

  63. 9:03

    -from start to finish the time.

  64. 9:03

    Mm-hmm. Mm-hmm.

  65. 9:03

    Like, you can ask people how much more productive did you feel, and they will give you an accurate response that correlates-

  66. 9:08

    Yeah.

  67. 9:08

    -with your quantitative feedback.

  68. 9:09

    Yeah. Yeah. Yeah.

  69. 9:10

    You ask anybody the amount of time that something takes, they are almost always wrong. So that I was like-- Like, like when I shared this with my colleagues, I was like, "Okay, I'm not surprised about that at all."

  70. 9:19

    But what is interesting is how much is the slowdown aspect. That was what was interesting.

  71. 9:24

    Yeah. Yeah. Yeah. That, um, uh, yeah, point well taken. That, that, that makes a lot of sense. I do, I do, um, uh-- So I think we, des-despite this, were interested in time estimates because, um, you know, we're, we're interested in providing-

  72. 9:39

    Yeah. I mean, in terms of the perceptual, like, yeah, I do think that's relevant too also because, like, the perceptual aspect is also the hype aspect.

  73. 9:45

    Mm-hmm. Mm-hmm. Mm-hmm.

  74. 9:46

    Right? Like, so developers will tell you that they were faster when they weren't-

  75. 9:49

    Mm-hmm.

  76. 9:49

    -and I think that is worth knowing. Yeah.

  77. 9:51

    Mm-hmm. And, you know, to, to the extent that we're interested in, um, uh, measuring the, you know, possibility timing nature of, um, of capabilities explosions or sort of R&D being automated.

  78. 10:03

    One commonly proposed measure to do this is just like ask developers or researchers how much they're being sped up, and for exactly the reasons you're pointing out, I, I don't put a lot of faith in those, um, in those, in those estimates.

  79. 10:13

    So n-nice to, nice, nice to see it like this. Yeah. Some, some more, some more J-curve things. So I--

  80. 10:20

    So the, so the forecasters who, who are not predicting time to complete, right? They, they are, they are just predicting this, this effect size. The n-non-developers, the expert forecasters.

  81. 10:30

    They are told the degree of experience these developers have, and some of the forecasters are, um, in thinking about how this population might be different to other populations, pointing out various facts about the study.

  82. 10:41

    Like, they're more experienced. I expect experienced people to get less, uh, uh, to get less speed up or, you know, the repositories are larger. I think AIs are less capable of working on large repositories.

  83. 10:51

    I expect less speed up. They never, never mention, um, familiarity with tools. My, my sense is that, um, yeah, they, they share the, the sense that I had ahead of time, which was like most of the action is in understanding what AI's-- the kind of things that AI is good at or bad at in the first place.

  84. 11:10

    And all of these developers have experience with LLMs in their core development workflow. It's just Cursor that they're, they're quite-- That three-quarters of them are, are totally unfamiliar with at the start of the study.

  85. 11:19

    Um, so I, I just-- I, I wasn't seeing much, much margin. Um, yeah, I don't know. I, I, I think it, I think it is, I think it is an open, open question.

  86. 11:28

    I, I also, you know, we watched so many hours of screen recordings of these developers working, and, um, I just do not see-- Um, I think they're like prompting very reasonably, you know, in some cases worse than me and my colleagues, in some cases better.

  87. 11:42

    Um, I, I'm not seeing these, like, advanced workflows that they're not accessing.

  88. 11:45

    Yeah. And my experience is, is not that far off from this, is that there are times when, like, I am dramatically slowed down.

  89. 11:51

    Yeah.

  90. 11:51

    And there are times when I am accelerated.

  91. 11:53

    Yep.

  92. 11:54

    Uh, and although as my familiarity with the tool increases-

  93. 11:57

    Yeah.

  94. 11:57

    -I definitely-

  95. 11:58

    Do you want to speed up?

  96. 11:59

    -don't improve a lot because I learn over time-

  97. 12:01

    Yeah.

  98. 12:01

    -what I can tell it to do and what I can't tell it to do.

  99. 12:04

    Yeah.

  100. 12:05

    In addition to, like, just getting better with it, like understanding like, okay, now I need to plan, now blah, blah, blah, blah.

  101. 12:09

    Yeah.

  102. 12:09

    But I-- But that's why-- So the thing I've considered-

  103. 12:12

    That's like before you make a, like, high-level architectural decision that, you know, ten conversations, uh, ten conversations, uh, turns down is gonna blow up in your face. You, like, really try and think about it.

  104. 12:22

    Yeah, yeah, yeah. Exactly. And then-- and also, like, scope it down to, like, a smaller problem. Like I, at first I would try problems that were too large, and like it can't handle that.

  105. 12:32

    Yeah.

  106. 12:32

    But just-- I mean, just for the future, if you ever do... I mean, I think it's obviously really hard with the, with the sample, with the sixteen-person sample size.

  107. 12:38

    But, you know-

  108. 12:40

    Well, now it's changing. [laughs]

  109. 12:41

    That's great. Great. 'Cause, 'cause in the future, what I, I think-

  110. 12:45

    Yeah

  111. 12:45

    ... having a cutoff, like trying to figure out if there is a cutoff of familiarity where the number changes would be interesting to see if that meta result generalizes, like, outside of Meta.

  112. 12:54

    Um, we are, we are on it. I think, um, the AIs have been getting better during this period, which is gonna compound a lot of, a lot of what's going on, obviously.

  113. 13:02

    But, uh, yeah, yeah. Interesting.

  114. 13:04

    The thing is the projects themselves are very optimized for people coming up onto new projects and figuring out how to... You know, they're, they're already-- The ones that struggle to be organized well for humans to come on board and be able to navigate them quickly don't survive very long in the open source ecosystem.

  115. 13:20

    Yeah.

  116. 13:20

    And these are fairly mature open source projects.

  117. 13:23

    Mm.

  118. 13:23

    They're a little bit different from like in enterprise settings where things survive 'cause they make money even if they're a pain to develop on. Right? So the, the context is a bit different.

  119. 13:33

    Mm. These are the, the repos run. Yeah.

  120. 13:38

    Yeah, that, that is a really interesting point 'cause like, uh, actually some of the repos that I was helped the most with were ones that I was completely unfamiliar with and which had no decent documentation of any kind, and where, like, I, I had to come in on this legacy code base that has existed for years and,

  121. 13:53

    like, make a change. And, uh, and, like, the developer who owned it was, like, only partially available to answer questions to me. And so in that case, like, Cloud Code was a huge help.

  122. 14:01

    Yeah. Legacy code bases don't exist 'cause they work well. It's because they make money.

  123. 14:05

    Yeah. [laughs]

  124. 14:08

    Interesting point. Yeah.

  125. 14:10

    The question I had was, um, sort of like, did all the developers have the same level of AI, like, uh, familiarity with Cursor? Or was, was there some variance?

  126. 14:18

    And was that, uh, like, is there a plot of, like, each of their, each of their familiarity?

  127. 14:23

    There's always a plot. [laughs]

  128. 14:28

    There's always a plot.

  129. 14:30

    There's always a plot.

  130. 14:30

    Can you just like, like you kind of like dig into, like, the question of is there, is there a J curve?

  131. 14:34

    Yeah. So, so here's some, here's some evidence. Um, so okay. The, the-- You know, I can show you some plots. I think the, the, the sample size is just small enough that, like, you shouldn't really believe any of the...

  132. 14:45

    I mean, the-- I, I think the plots aren't gonna show much, but then I, I don't wanna say that's like strong evidence this is not something that's going on.

  133. 14:51

    I just think the evidence is kind of weak. The thing that really convinced me is like I watched the videos. [laughs] I watched the videos, and I'm working. And, you know, often they're better at using Cursor than me, and I'm like, "Well, you know, I'm, I'm working on this project using Cursor." [laughs]

  134. 15:03

    Um, but, but here, here are some graphs. So, um, so this is by whether they have various types of, um, uh, AI experience coming into the study and, you know, basically you see no movement in, in point estimates.

  135. 15:16

    People for whom Cursor was a primary IDE before, um, yeah, not, not a huge amount of difference versus people for whom it was not. Um, then the next one is, you know, you might think may-maybe you have a view that, you know, some J curve cutoff comes after this point.

  136. 15:32

    But still, you know, within the, within the study, there's some variation in how experienced people are with AI because they have multiple issues. You know, the-- A-after the first AI issue, they're slightly more, uh, exposed than after second AI issue.

  137. 15:45

    So you might try sort of excluding those data points over time and, and, and seeing, and seeing what pops up. And, you know, they don't, they don't seem to get better at using AI over time.

  138. 15:52

    Although I think there's probably a stats sig issue with that.

  139. 15:54

    You think there's probably what, sorry?

  140. 15:55

    There's probably a stats sig issue with that, that plot right there. Like, those bars are very, very wide.

  141. 16:02

    Oh, I mean, I think, yeah. None-- I, I think, like, all of the, um, plots outside of the main plots, all of these subset things you should, like, not put a lot of stock in.

  142. 16:11

    Yeah.

  143. 16:11

    Um, yeah. I, I, I, I totally, I totally agree. Um, okay. And then lots, lots has been made-- So, so th-this graph is the reason we put it in unclear evidence 'cause we're like, "Ah, things point in different directions."

  144. 16:22

    Um, a lot's been made of, of this plot suggesting, you know, something, something J-shaped. In particular that, you know, at the end, once people have more experience, um, uh, they, they do experience some, some speed up.

  145. 16:33

    Um, here are some issues. You know, first, like the other plots don't. I, I think that's important to, to, to include. Uh, second, these hours are coded very conservatively.

  146. 16:42

    So for instance, someone in the thirty to fifty hours bucket is, um, uh, had Cursor as their primary IDE in twenty twenty-four. They had recorded themselves on their time tracking software as having spent a hundred and forty hours using Cursor.

  147. 16:57

    They conservatively estimated that they'd spent fifty hours using Cursor, and so they-

  148. 17:01

    Oh, okay

  149. 17:01

    ... ended up in our thirty to fifty hours bin. This is someone whose, whose primary IDE was, was, was Cursor last year. Um, and, and, you know, people have been commenting about this.

  150. 17:10

    They've been using Cursor for less than a week. I think that's not a, not a very fair assessment. If you, if you were to move that developer over from the, uh, penultimate bar into the...

  151. 17:18

    Again, you shouldn't believe this because of statistics issues. But, um, if you were to move the, uh, that, that developer from the, um, penultimate, um, effect size estimate to the, to the last one, then you'd see some balancing out where you get back to essentially zero in, in the last bucket.

  152. 17:34

    Uh, yeah. Again, so, so, like, not definitive. I think J curve explanations, you know, still like very on the table. Um-

  153. 17:40

    Is, is it not likely, though, that the fifty-hour group also is similarly underestimating their, their time they've spent using Cursor and that actually if you just had a bur-- longer scale, that you would still see a J?

  154. 17:51

    Oh, that, that is an interesting point. Um, um-

  155. 17:57

    That seems plausible to me. Um, and then, and then I guess I want to... I- I'm not sure it's underestimate because we're using this, like, very conservative-

  156. 18:04

    Yeah, totally. Yeah.

  157. 18:06

    T- totally. Um, yeah. I, yeah, I think that seems plausible to me. And then, um, for this not to be any strong evidence, I'd retreat back to I think you shouldn't really believe in any of these-

  158. 18:15

    Yeah, yeah. No

  159. 18:15

    ... charts. But-

  160. 18:16

    I think the biggest thing is it's small sample size, and there's also a lot of bias in the data set effectively, right? Like, it's a certain kind of data set.

  161. 18:24

    It's open source-

  162. 18:24

    You mean, like, the kinds of, the kinds of developers?

  163. 18:26

    Oh, it's, yeah, open source developers, and also working on open source projects that are pretty mature.

  164. 18:30

    Yeah.

  165. 18:30

    You know, those, those two things are... If you're working with open source developers-

  166. 18:35

    Yeah

  167. 18:35

    ... on projects that are pretty mature, this is probably reasonably indicative maybe, but the sample size is pretty small. But outside of that, it gets a little harder.

  168. 18:43

    Yeah. In talks about this, I'm like, um, uh, I think, yeah, this group is really weird.

  169. 18:49

    Yeah.

  170. 18:49

    It's really interesting. It's, like, interesting for the same reason-

  171. 18:51

    Right

  172. 18:51

    ... it's weird, right? Um, uh, yeah, we, we were interested in, you know, again, studying, um, uh, possible effects on, of AI for R&D speed up or, or, or automation.

  173. 19:01

    Um, there, if any types of developers are not being greatly sped up, it implies the whole thing isn't, isn't being sped up. So, so it is kind of curious to see even-

  174. 19:11

    Yeah

  175. 19:11

    ... even, like, particular weird populations. You might imagine in, like, large, you know, sort of production inference code bases maybe have a bit more of the shape than scrappy experiment scripts.

  176. 19:21

    Yeah. Yeah.

  177. 19:23

    Um, but yeah. Totally.

  178. 19:24

    No, no, I think, I think it's very interesting. It's just, it's hard to generalize. We just don't know.

  179. 19:29

    Yeah. Yeah.

  180. 19:30

    Although-

  181. 19:30

    We are doing this large study, and I think, uh, you know, I th- I think unfortunately, after the lo- large study, which includes more greenfield projects, and, um, I think it's still gonna be hard to surprise. [laughs]

  182. 19:38

    It's the same. Yeah.

  183. 19:39

    Um, for, for not so similar reasons. Yeah.

  184. 19:42

    Although I don't feel like your results are particularly contradictory with any actual independent research that's been conducted. The only research that I've seen that I would say is contradictory to yours is research that has been funded by model shops or agent shops. [laughs]

  185. 19:56

    Um, what can I say about that? I, I do, I do think that most of the research that's, that's put out, um, uh, is associated with, uh, large tech companies.

  186. 20:09

    Um, and, uh, and I, I think there are other methodological concerns that I, yeah, at least as well.

  187. 20:15

    I, I have methodological concerns with that as well. I know people who work at some of those places who have methodological concerns with the work that was output. So...

  188. 20:23

    I mean, I, you know, I think, I think there are, there are concerns as well. So, yeah.

  189. 20:25

    Sure. Sure. But I, I actually feel... But, like, I, I remember somebody sent me your paper, and when I saw the headline, I was like, "No way." [laughs]

  190. 20:33

    Well, me too. [laughs]

  191. 20:35

    Right. And I, I was like, "That sounds like BS."

  192. 20:37

    Yeah, yeah.

  193. 20:37

    I read the paper and I was like, "Oh, this doesn't suck at all." [laughs] Like, uh, yeah.

  194. 20:43

    A little bit.

  195. 20:43

    No, no. [laughs] I feel like at least your high-level conclusion both is intuitive, like, from a person who's read a lot of software engineering research, and also is well justified.

  196. 20:54

    I, like, I think people... I have had people argue with me about the 16 developer thing, but I don't think that actually matters in that particular case because I think they're actually a fairly good control set more or less, right, for an experiment because they, they remove a lot of validity concerns by being experts.

  197. 21:11

    So yeah, they... it's true that they don't represent certain, like, um, like the broad aspects of developers, but they also remove a lot of variance in what you would expect from the population.

  198. 21:20

    And they, and they allow you to have, like, a sort of an epistemological function of like, "Hey, let's isolate that factor away, and then let's, let's see what happens with that."

  199. 21:31

    And that's what-- I like that. And then I thought the way that the study was conducted was completely sufficient to draw a conclusion, a high-level conclusion, but a draw.

  200. 21:38

    Thank you very much. Um, here, here's a, here's a curiosity. So, so we did-- We haven't published this because of the organizational reasons that I won't go into. But, um- [laughs]

  201. 21:48

    ... we did, um, we did conduct this, um, uh... You know, pe- people would throw sort of their various explanations for, for, for what's going on here, you know, many of which have lots of merit, some of which more, more skeptical of.

  202. 21:59

    Um, you know, a natural one is brownfield versus greenfield projects. Um, so we ran this, um, kind of enormous hackathon where we randomized half of teams to, um, use AI versus not, kind of, you know, maximally greenfield or something.

  203. 22:12

    Um, and, uh, and then we'd have a bunch of judges score them. Um, you know, m- many judge scores per project or something to try and even out that noise.

  204. 22:20

    And we'll see, you know, i- is it the case that, like, the bottom, uh, fifty percents are all the AI disallow group and the top, uh, um, the top are all the, um, AI allow groups or something like that.

  205. 22:30

    Now, u- unfortunately, it was sort of even, even smaller. That's, like, part of the reason we're not publishing it is I think the evidence is, is, is really quite weak.

  206. 22:37

    The, the degree of overlap is enormous. Like, the, the point estimate that we, um... I'm a bit nervous about saying this because, you know, it hasn't gone through the kind of review processes that something like this goes through.

  207. 22:47

    So, so, um, maybe I've messed something up. But, um, uh, I think the point estimate is something like four percentage points higher on a, on a... Sorry, four percentile points higher, um, if AI is allowed versus if it's not, after, after controlling for everything else.

  208. 23:03

    That is, like, you know, extremely noisy, and you shouldn't draw any conclusions. But, um, but seemingly maybe kind of, um, small effects, hackathons allowing AI. Um, yeah. Yeah.

  209. 23:17

    So the question I have, I, I guess this is, like, related to the study, but also related to oth- other research that you guys have done. Um, so have you found a similar pattern?

  210. 23:24

    Or I guess first, have you, um, explored, um, like, the effect of AI in other domains, specifically software engineering? Um, and if so, have you also found this kind of surprising result that maybe as much of a speed up?

  211. 23:40

    Um-

  212. 23:40

    Or maybe I-

  213. 23:40

    Um, no, no, no, no. I mean, no, new directions.

  214. 23:44

    Oh.

  215. 23:44

    Um, stuff that we have not done. Um- Uh, yeah, I, yeah, w- you know, we- we're interested in understanding, um, uh, possibility of accelerating R&D. Um, you know, coding is not the only kind of thing that happens at major AI companies.

  216. 24:00

    Much more conceptual work happens. Um, uh, you know, I'd be, I'd be very excited about, um, um, you know, working with math PhD students or very different types of software developers or, um, or, you know, running, running these kind of studies inside of, um, major AI companies or, or large tech companies or, or something like that.

  217. 24:19

    I think, um, w- we are very interested in, you know, not necessarily directly, but some somewhat close analogy to, um, to the large AI company case. So to the extent that something really deviates from that, um, probably less interested.

  218. 24:34

    Interesting. So yeah, so I guess it sounds like, uh, you're interested in measuring capabilities for like other, like, like math research, um, and, uh, some other, like some other, uh, research.

  219. 24:45

    Yeah. I, I'd say I'm interested in like what the hell is going on in AI [laughs]. And, um, you know, how, how am I gonna learn the most about what the hell is going on in AI?

  220. 24:54

    Um, you know, I, I think something, something a bit more conceptual, some- something where, you know, fewer humans are currently working on it, so it's less appearing in training data, um, will help me better sort of triangulate the truth about what's going on in AI.

  221. 25:07

    Um, even if I don't care about math research in particular, um, it'll still sort of draw helpful qualitative lessons is, is kind of the sense I have.

  222. 25:15

    Yeah. I mean, if I was gonna pick the areas that I think it's most successful in or like the areas where I would expect it to be more successful, but where I think it is being less successful, I would pick probably data science-

  223. 25:25

    Hmm

  224. 25:26

    ... as an interesting one. Like, how does data science-- how do, how much are data scientists helped by AI today?

  225. 25:31

    Say, say more about where you expect it to be less successful.

  226. 25:33

    Um, so in a, in a real... So let me give you an example.

  227. 25:37

    Yeah.

  228. 25:38

    I used to work at LinkedIn.

  229. 25:39

    Hmm.

  230. 25:39

    And at LinkedIn there are five thousand tables with the name impressions in the na- in the table, right? So if an analyst wants to understand how many impressions happened on a page, where the hell do they go?

  231. 25:48

    Yeah.

  232. 25:48

    We can't figure that out.

  233. 25:50

    Yeah.

  234. 25:50

    Like, today, there is no existing AI system that we have that can be hooked into like corporate environment like that and processed through... I mean, there's trillions of rows in those tables.

  235. 26:00

    So like, how, like, like, so what a data scientist needs to do is they need to be like, "I need to like, you know, analyze a bunch of data and come to a conclusion."

  236. 26:08

    Right? Uh, and I, I hear lots of like thoughts about building systems. You know, people talk, talked about ML for SQL. The models are much better at writing SQL than they used to be.

  237. 26:18

    But I believe that the state of underlying data is so bad

  238. 26:24

    that the, the, the actual data scientist are gonna get way less value out of the-

  239. 26:29

    Hmm

  240. 26:29

    ... out of the AI than-

  241. 26:30

    Hmm

  242. 26:30

    ... software engineers are always making for.

  243. 26:32

    Hmm. That is, that is-

  244. 26:33

    Interesting

  245. 26:34

    ... that's very curious. I, um... So o- one, one view that some, some more bearish people have looking, looking at the future of AI is, is, um, you know, so much, there's so much tacit knowledge around, there's so much knowledge that sort of, um, embedded inside of companies that you're, you're not going to pick up from, you

  246. 26:50

    know, these like RL training environment startups or, or something, something, something. You know, maybe it, it's not sort of the state of nature that there needs to be many specialized AIs.

  247. 26:59

    Indeed, like much of the lesson of the past few years is that one big general AI seems to, seems to be more performant. But you know, at some point in the future when data is like locked up inside of companies, um, uh, you know, we will have more of this, um, uh, proliferation of, of many more specialist

  248. 27:13

    models as I have, you know, GPTN fine-tuned on, on LinkedIn data in particular, something, something, something, something. I have one reaction that's kind of like that.

  249. 27:20

    Yeah, I don't know. I, I-

  250. 27:22

    I, I do have a disbelief-like reaction. I'm like, ah, science, you know. [laughs]

  251. 27:27

    No. But also like, so, but also like, so contradictory facts.

  252. 27:31

    Yeah.

  253. 27:31

    So the problem, problem is the, all these data sets contain contradictory facts. Like the name of the field will be, uh, like, uh, you know, date started or like title...

  254. 27:43

    It'll be time started, right? And then it will contain only a date, except for it will only contain the date up until like November of last year, and then after that it will contain only the month.

  255. 27:52

    But then after that it will contain maybe the, the seconds that the thing finished. And in order to actually successfully query the data set, you, the data an- you, the data analyst or the data scientist have to know what those cutoff dates were.

  256. 28:02

    Hmm.

  257. 28:02

    Which is not written anywhere.

  258. 28:04

    Hmm.

  259. 28:05

    Although what you could do theoretically is import a bunch of the SQL that other analysts have written to try to figure out like the, how they triangulated these things and work backwards from those reports.

  260. 28:13

    Hmm.

  261. 28:13

    But today, so I think today, for example-

  262. 28:15

    Yeah, the people... Sorry, I've just like, I haven't worked at a large company. People, [laughs] people don't fix this at the source? [laughs]

  263. 28:22

    No, no. So-

  264. 28:23

    I feel like the lesson I learn over and over again at this scale-

  265. 28:26

    At the end of the day

  266. 28:27

    ... data specs really matter.

  267. 28:28

    Really, really matter.

  268. 28:30

    So-

  269. 28:30

    No, I, I, I've also been working in data analysis and, and research, developer research and so on.

  270. 28:34

    Yeah. Yeah.

  271. 28:35

    And so, yeah, so the, like the problem is like their job is like produce this report for this executive, right? Not go make infrastructure to produce this report for this person.

  272. 28:45

    Yeah. Ugh, but I'm like if I... Okay. [laughs] [laughs]

  273. 28:50

    I, I'm with you. I live that dream every day. Yeah.

  274. 28:52

    Right.

  275. 28:52

    No, you just have-- you, you end up having to, right? Is, is-

  276. 28:56

    Yeah

  277. 28:56

    ... you have to build out infrastructure for it. That has to be part of the job description. And, and the other part is you have to fix the problem at the source.

  278. 29:02

    Like you really... I, I, I still remember having a conversation where, where someone said, "It's too difficult to fix it at the source because there's too much complexity of all the systems that-

  279. 29:12

    Hmm

  280. 29:13

    ... end on the source." And I said, "Okay, wait a minute." [laughs] "You're saying it's too complicated to solve at the source downstream somehow a problem that is too big for the en- entire organization to solve?"

  281. 29:23

    Yeah.

  282. 29:23

    "It's easier to solve there? Come on."

  283. 29:25

    Yeah.

  284. 29:25

    "Like that doesn't make any sense."

  285. 29:26

    I just think there's so much potential here, and I have not seen a lot of studies done on like how people who are at, working in that data space are experiencing AI.

  286. 29:36

    And what's fascinating about that is real ML is mostly data work. Like, like ML, especially outside of LLMs, ML outside of LLMs, the majority of ML engineers spend most of their time doing like feature curation-

  287. 29:48

    Rather that they spend actual direct model training.

  288. 29:50

    Mm-hmm.

  289. 29:50

    And like trying to clean up bad data for feature curation. So like theoretically, the potential even for the improvement of ML by enabling ML to be a better data scientist is huge, and I, I suspect that if you...

  290. 30:02

    W- my hypothesis is, if you went into this space, you would discover it is great at telling me how to write SQL, uh, or how to, like, write Pandas and, or Polars or whatever you're using.

  291. 30:13

    It is okay at doing very trivial things, and it fails at all complex tasks.

  292. 30:18

    Mm.

  293. 30:19

    Like, fails completely at all complex tasks.

  294. 30:20

    Mm.

  295. 30:20

    I don't even know... I haven't even seen a benchmark on it.

  296. 30:23

    Hmm. Hmm. C- can you give me an example of a, of a, a complex task?

  297. 30:27

    Sure. Uh, uh, let's say a, a complex task is

  298. 30:34

    determine the time between, uh, give me the P90 of time between deployments for all deployments that happened to Capital One.

  299. 30:42

    It struggled at that?

  300. 30:44

    That... Yeah, that, that... It does seem surprising to me.

  301. 30:46

    That seems surprising, right?

  302. 30:47

    Yeah.

  303. 30:47

    Uh, so, uh-

  304. 30:49

    And like, it... You know, if it has sort of reasonable context about where it would find this-

  305. 30:52

    So it found the data, right? Sure. Sure. Makes sense. And, uh, and, and then, so okay. So fine. So, so give me that number, and then also, uh, make sure that you can break that down, you know, by team hierarchy.

  306. 31:03

    Sure.

  307. 31:03

    So like you give me that in a table so I can break it down by team hierarchy. Uh, where is the team hierarchy data?

  308. 31:10

    Like, uh, how... And oh, h- h- here's a funny thing. Uh, what PRs were in those? So how do I know... Uh, how, how would I... How do I actually determine what the time deployment started and ended was?

  309. 31:22

    'Cause it turns out that's not clear in the base telemetry, and you have to, like, know magic to figure out when the t- when the deployment started and ended.

  310. 31:28

    Um, uh, oh, and also tell me, you know, for my ability to analyze it, tell me how many PRs were in each of those deployments, and which PRs went into each of those deployments.

  311. 31:37

    Well, guess what? The deployment system only rec- this isn't being recorded, right? The, the, the, the-

  312. 31:41

    I think it is being recorded.

  313. 31:42

    Is it? Okay.

  314. 31:42

    Yes, but before you. [laughs]

  315. 31:45

    There we go. Um, so then, you know, imagine the deployment system doesn't contain sufficient information about that data, right? Uh,

  316. 31:53

    then, like, like, where do I get that data? Well, that data doesn't exist in any other system. So what I... Well, maybe I have to go, like... I have to go to GitHub, and I have to call the GitHub API, and, like, the chance of the LLM or an agent figuring that out today is pretty minimal.

  317. 32:10

    Hmm. Yeah, I do still, you know, r- relative to my colleagues, I'm, I'm, I'm pretty bearish on AI progress. I, I, I do still have some reaction that's like, ah, like,

  318. 32:21

    can't you spend a day getting this into a Cursor Rules file? [laughs] You know, like, where, where, where the, um, where the hierarchy exists.

  319. 32:29

    I, I would, I would go... I, I think... That's why I think it's interesting. I think it would be worth studying. I don't... I have not seen any-

  320. 32:35

    Yeah

  321. 32:35

    ... real comprehensive study on the experience that data scientists have. Um-

  322. 32:39

    Um, if you, if you have any ins to, um, uh, to, to us running studies at large tech companies, then I, I am all ears. [laughs]

  323. 32:46

    The only... There, there is a fellow at OpenAI that I was talking to who's one of the speakers who does evals, uh, internal evals, and he has mentioned that he's done some work with data scientists.

  324. 32:56

    So he might know some people who have that data.

  325. 32:59

    Mm.

  326. 33:00

    But it's, it's all been internal between him and, like, Cursor, between him and, like, you know, Anthropic or whatever, right? Um,

  327. 33:08

    uh, yeah, that, and I also think, uh, I, I... One of the ones that I'm curious about too is lawyers.

  328. 33:15

    Curious about, like, more traditional, like, older, like, um, lawyers, doctors, and I think mathematicians are all gonna be interesting to me.

  329. 33:22

    Hmm.

  330. 33:24

    Just because, you know, both lawyers and doctors are so constrained by a legacy history of, like, the constraints around them and how they work.

  331. 33:34

    Um, yeah. Legal, legal issues I'm imagining continue to be a significant barrier.

  332. 33:38

    Yeah. And they're stodgy. Like, I, I, I'm also interested in, like, what's the... how are the stodgy-

  333. 33:45

    So stodginess I feel like is, is a, um, I, I think I'm less bought into as a long-term explanation for econ- I, I... Like, the, the, the legal restrictions, they sort of continue to be the case through time.

  334. 33:55

    The stodginess, I can, like, set up a new law firm that's less stodgy, and then- [laughs] ... it takes the previous law firm. It seems, seems to be-

  335. 34:02

    I agree. I agree. I, I-

  336. 34:03

    ... medium term.

  337. 34:03

    I don't think it's persistent. I just think it's, it's interesting to see... Well, I... One thing that would be interesting to see is, like, if that affects the mental model that they have today.

  338. 34:12

    Like, like, if, if they're... Like, how they've been talked to about it or how their trust in it affects how they use it.

  339. 34:18

    Hmm.

  340. 34:18

    Would be interesting to know, to me. Like, I don't know if it's a worthwhile study. It's more of like one of those things that I wonder about idly. [laughs] You take a lawyer who just got out of college and sort of, you know, has spent a lot more time using ChatGPT, and you take a lawyer who's been in

  341. 34:31

    the business for fifty years and, you know, has, has a, a giant file folder full of Word docs that contain, like, all the briefs that all their, you know, junior associates have written for decades and decades, and he just opens up those briefs and, like, changes a few words in them and then sends them out to the

  342. 34:46

    judge, and he, like, you know, has known those judges for, like, thirty years, forty years, and he knows exactly what they want and, like... You know, is he getting any...

  343. 34:55

    Is he gonna get any value? But is there a value he should get? Is there something that, like... Is there some way that, like, he would be helped by AI?

  344. 35:04

    I certainly know discovery. Discovery in AI is, like, in, in law is, like, a huge, huge problem. And I, I know that, like, there's Harvey. I don't know anything about what success they've had.

  345. 35:17

    I know a lot of people working in that space specifically. Like, that's... It's an ongoing thing, right? There, there's always technology for it, but it's kind of... The adoption of it is a very different thing from-

  346. 35:30

    That's, that's the, that's the thing, right? 'Cause I... One of the first things that I thought of... 'Cause I, I have a little bit of a legal background, and one of the first things that I thought of the first time, like, when ChatGPT 3 came out, I was like, "Oh, this could totally change discovery."

  347. 35:44

    Like, this could... Because discovery is, like, the most painful and most difficult and most expensive. Like, you could have serious social consequences by making discovery less expensive. Like-

  348. 35:53

    Right

  349. 35:54

    ... that is the expensive part of having a lawsuit.

  350. 35:57

    And so, like, you could actually have significant impact on a society if you could make discovery cheap and instantaneous and reliable. Yeah.

  351. 36:09

    I have a question on your graph.

  352. 36:12

    Yeah.

  353. 36:12

    Cursor.

  354. 36:13

    Mm.

  355. 36:14

    I'm not sure you-

  356. 36:14

    Mm, mm.

  357. 36:16

    You missed it up. Keep on going. Few more.

  358. 36:23

    Yeah. Whoa, sorry.

  359. 36:25

    All right.

  360. 36:25

    It's a scatter plot, right?

  361. 36:27

    Um-

  362. 36:28

    It was what? Cursor in fifty hours. And according to-

  363. 36:32

    Oh, I see. Yeah. Yep. Dun, dun, dun.

  364. 36:36

    There.

  365. 36:39

    Uh, I say it's this one.

  366. 36:42

    Yes. That one. So you're saying that people-- the, the developer-- there was no difference. Cursor, we're talking about the ID, that vibe coding, and they use it for fifty-hour problem.

  367. 36:56

    I was very intrigued by that because everyone talks about vibe coding and how Cursor is instrumental. And why did you get to-- how did you get to fifty hours?

  368. 37:07

    Just curious.

  369. 37:08

    Um, so, so this is including time-

  370. 37:11

    Five and fifty hours is-

  371. 37:13

    This is including, uh, time in the experiments, um, that developers have spent in the experiments, plus their, plus their past experience. So for, um, for, for some developers working on some issues as, as part of the experiment, some of them have gotten to more than fifty hours of, um, cursor experience.

  372. 37:29

    Um, uh, and that's, that's just coded up in that, in that bucket at the end.

  373. 37:34

    And was, was it the same task for each group?

  374. 37:36

    Uh, no. These are kind of-- they're, they're natural tasks that pop up on the GitHub repositories. Which, which as mentioned are kind of, um-- I don't wanna-- uh, I'm, I'm a little bit nervous about saying they're weird 'cause it implies they're, um, uh-- I wanna say it's very interesting, and it's very weird.

  375. 37:53

    And, and it's interesting for the same reasons it's weird. These, these are, um, these are repositories in which they have-- these, these are projects in which they have an enormous amount of mental context built up, um, that the, the AIs might not have, um, that they've worked on for, for many, many years, that they can, um...

  376. 38:09

    I'm not sure this is always the case, but, you know, I imagine it in my head that, that they basically know how to execute on the particular task they have before, um, uh, before they even, you know, g-go about attempting it.

  377. 38:21

    Because they're so expert in, in the, in the project.

  378. 38:24

    You mean when positive speed up, is it like, like five percent? Like, what do you mean by-- What's-- How do you quantify the speed up?

  379. 38:31

    Um, so, uh, you might think about, uh, let's, let's go to this one instead. So, um, on the here left-hand side, we have the, um, averages for what the developers say, um, will happen in terms of their time to complete if their issue or their task gets assigned to the AI disallowed or the AI allowed group.

  380. 38:54

    Um, you know, they, they think that if AI is disallowed, it'll take them a bit more time, closer to two hours, and I guess more like an hour and a half or a little bit less if AI is allowed.

  381. 39:04

    Um, but then, you know, we, we randomize this particular task to allow AI or not allow AI, and it turns out, you know, if we randomize to AI allowed, then the, the times are more like a bit above two hours rather than a bit below two hours.

  382. 39:16

    Um, and then you can think of the, uh, change in time estimate as sort of being one divided by the other here. It's not quite that for reasons, reasons I can go into, but it's, you know, it's effectively, um...

  383. 39:27

    What exactly is the transformation? You know, whatever. It's something like AI disallowed over the AI allowed minus one.

  384. 39:35

    So, uh, to, to draw that out and like, um, you know, I might be like, "What's, what's the speed up?" You know, is it like, uh, one point one x that, you know, these, these developers are going one point one times faster when we're actually on a time to complete scale, not a, not a speed scale.

  385. 39:52

    But i-ignoring, ignoring that, um, ignoring that detail, um, you know, is it one point five x? Um, is it zero point five x, that they're actually going sort of twice as slow?

  386. 40:02

    Um, how, how would we get that information? Well, we'd do something like, um, take the AI disallowed times divided by the allowed AI times. You know, if this was, uh, one point one, let's say, times as long as the allowed times, then we'd get to, uh, one point one x speed up.

  387. 40:23

    It's something, something like that that's going on.

  388. 40:29

    And in fact, you know, we find a, we find a slowdown, famously.

  389. 40:34

    I, I just read a fascinating article, last company I can remember. But basically, journalist

  390. 40:41

    was allowed to, uh, using vibe coding, right, uh, do a pull request, meaning there was some feature, and AI was used to assist with building out the requirements.

  391. 40:59

    And she practically, according to the article, just kind of did a little couple of tweaks and then just signed off on it. It was just really fascinating. That was the whole vibe coding thing.

  392. 41:12

    Yeah. I-

  393. 41:13

    She didn't code. Like, that was the whole thing. It was like she didn't have any software development background. That was the whole like... I was just curious

  394. 41:22

    if you've tried to do a study on that.

  395. 41:27

    So I, so I definitely do-

  396. 41:29

    I definitely do to share this out. But, you know, if you've got, like, no idea what's going on-

  397. 41:32

    Yeah

  398. 41:32

    ... then, um, probably, probably these are going to be some, some significant, um, some significant speed up. You know, I, I, I will say, I guess number one, it's not, um, you know, it's not a priori obvious.

  399. 41:44

    Um, you know, in fact, we went out and did this hackathon with, you know, very experienced people and much less experienced people and, and tried to see what happened.

  400. 41:52

    And what we found is, you know, the scores, the judge scores extremely noisy, and I think you shouldn't believe it. But, [laughs] um, you know, the, the, the judge scores were, uh, not that much higher when the AI was allowed versus, versus when it was not.

  401. 42:03

    The, the, the people aren't actually making that much more progress. And then, and then another thing to say is-

  402. 42:08

    I, I think there's going to be more expertise in this, in this room than, than I have. My understanding from, uh, you know, sitting with these open source developers for a while and not, not being a very capable developer myself, um, is, um, is, is that the, the, the quality bar on the repositories in, in this study

  403. 42:26

    is just very high-

  404. 42:27

    It is. It is

  405. 42:27

    ... typically. Um, and so I would be very surprised if a journalist, um, you know, even frankly if, like, a good software engineer without lots of experience on the repository, but, but certainly, you know, someone who wasn't a software engineer was able to get up a clean PR on these repositories first time.

  406. 42:47

    I- in fact, I think that's a lot of the story for what's going on here is that the AIs, you know, they actually kind of do make progress in the right direction some, some good fraction of the time.

  407. 42:56

    But, um, for, you know, for various reasons, sometimes for reasons of correctness, but sometimes for reasons of, like, you know, how they've tried to solve the problem and, you know, whether that's the typical way of solving the problem or, like, how various parts of the project speak to one another.

  408. 43:10

    These, these kind of considerations, you know, they, they haven't properly accounted for that. And so, you know, the humans not only need to spend expensive time verifying, but also, like, clean up, clean up all the stuff.

  409. 43:20

    My, my sense is that someone who didn't have all that experience, like, basically wouldn't know how to do that step, um, and so wouldn't be able to submit a clean PR to these repositories.

  410. 43:29

    You know, that- that said, like, I, relative to these people at least, I suck at software development [laughs] and I, I'm getting up, you know, PRs internally all the time. [laughs]

  411. 43:37

    And I think they're, I think they're worse quality and, um, you know, and they're, and they're getting over time-- they're getting better with time. You know, I do believe that people are coding when they, when they wouldn't be able to code.

  412. 43:47

    They are submitting, you know, PRs at a lower quality standard when they wouldn't be able to do that at all. Um, but getting up, getting up these expert level PRs, I, I do feel kind of skeptical.

  413. 43:56

    And, and that's actually part of what I was getting at is they often get... PRs often get rejected by more novice, uh, folks on these big-- on these bigger quality projects for no other reason other than the developer ergonomics impact of the PR, right?

  414. 44:11

    Mm-hmm.

  415. 44:11

    So the fact that it makes it harder for me to future maintain... 'Cause, 'cause for an open source project, almost all the incentive is biased towards making it easier for me to maintain the project, right?

  416. 44:21

    So every time a PR comes in, if it doesn't make it easier for me to maintain the project, I have a tendency to reject it.

  417. 44:28

    Yeah.

  418. 44:28

    Right? Uh, if it does make it easier to maintain the project, then yay, I'm into it. As a-- That is unlike what you have in a typical business context, right?

  419. 44:37

    Where the most important thing actually is to get something done.

  420. 44:39

    Yeah.

  421. 44:40

    Right? Uh, because you're-- You know, the fact that, that someone's gonna have to spend a lot of time maintaining, it's almost job security, right? But for open source, it's the opposite.

  422. 44:48

    It's actually what causes people to leave projects is when it's difficult to maintain, right? So it is a different bias on what you accept for pull requests.

  423. 44:56

    C- can you remind me the name of the-- name of the [REDACTED:origin] [REDACTED:gender] who maintains the Haskell compiler?

  424. 45:03

    Um-

  425. 45:03

    Simon something?

  426. 45:04

    Yeah, uh, no, I d- I can't remember.

  427. 45:06

    No? Okay.

  428. 45:07

    Simon Name Recall.

  429. 45:08

    So here's, here's, here's one story that might be relevant. Um, you know, a bunch of repositories, uh, in, in the study. They, they all have, you know, broadly these characteristics.

  430. 45:15

    One of them is the Haskell compiler. Famously, on the Haskell compiler, um, there's, like, some chance, I don't know if it's fifty percent or thirty percent or whatever, but there's some chance that if you submit a PR, the...

  431. 45:27

    I'm being recorded.

  432. 45:27

    Simon-

  433. 45:28

    The-

  434. 45:29

    Simon. Simon-

  435. 45:29

    Simon something

  436. 45:30

    ... Marlowe, maybe.

  437. 45:31

    I'm not sure. [laughs] Um, the creator of the Haskell compiler will come into the comments and argue with you for many, many hours, much longer than you spent working on the pull request, um, until, um, until the PR hits exactly your specifications.

  438. 45:45

    Um, combine that fact with the remarkable fact, I think, that the median PR in this study, the time they spend working on the code post-review is zero minutes. That is the m- the median PR is, like, perfect first time around because the professional incentives that these developers are, are like that.

  439. 46:04

    Now, there's a very long tail. On one of them, um, on one of them, I think literally Simon, this [REDACTED:gender], uh, pops up and argues in the comments for many hours, and that, that, that one's a lot longer.

  440. 46:13

    Um, but, um, uh, yeah, they, they are, they are maintaining this extremely high bar.

  441. 46:21

    I'm interested in your other upcoming stuff that you had in your doc.

  442. 46:24

    Yeah, let's do it. Um, so, um, yeah, so, so, you know, so, so one thing I... Mm, what to say? Um, I guess let's, let's go in order. Um, as I, as I think you mentioned, um, you know, if, if, if, uh, capabilities as measured by time horizon keep, keep doubling, it does seem very, very challenging to

  443. 46:45

    keep up with that. Um, i- in the short term, we have a number of directions for, um, for getting on top of that, but, uh, a- and I think that will last, like, through the year.

  444. 46:54

    But through two years, you know, that seems challenging. Um, I think still possible. Through three years, I think still seems possible. You know, it starts, starts to get hotter and hotter.

  445. 47:05

    Um, anyway, in the short term, building these, building these much longer tasks and ways in which we might get around the problem entirely. For instance, um, here's one thing that might be somewhat-

  446. 47:15

    You could also raise the accuracy bar.

  447. 47:17

    Uh, you could raise the accuracy bar. Although, um, you know, we're-- the reason we're interested in this in the first place is we're like, um, you know, is GPT-5 existentially dangerous?

  448. 47:29

    Okay, and the answer is no, I think.

  449. 47:31

    Yeah.

  450. 47:31

    Um, but, like, why-- but, like, why, why do we think the answer's no? Okay, at least, I think there are multiple reasons. But at least we can say, you know, GPT-5 is just, like, not that good at stuff.

  451. 47:39

    Like, you're trying to get it to do, you know, data science on these, on these, um, uh, uh, columns that have very similar names, and it's, like, not clear exactly the logic that's led to the columns.

  452. 47:47

    It's not, it's not doing that sort of thing. I'm like, "You need to do that sort of thing to be ex-" Not exactly that, but, you know, be capable of stuff to be existentially dangerous, and it's not capable of stuff.

  453. 47:56

    Um, so, you know, I see that the, the AIs are failing at these, these hard tasks. I'm like, you know, great, whatever. But if, you know, if it's succeeding ninety percent of the time and not ninety-nine percent of the time on, like, very challenging long tasks, then- You know, I think there are, there are, you know, other

  454. 48:12

    things that we can retreat to, and I'm prepared to retreat to those things. But at least the naive, the reason it's not existentially dangerous is that it's not capable of stuff.

  455. 48:20

    I, I think at least starts to get pushed on as, as, as an explanation. So he- here's another possibility for something we could do. Uh,

  456. 48:29

    whoop. So, um, what if instead of time horizon here, we had, uh, time horizon, uh, whilst,

  457. 48:41

    uh, being watched very closely. And then time. And then, you know, maybe the, the time horizon whilst not being watched very closely is like this, but the time horizon whilst being watched very closely is like this.

  458. 48:59

    That is like if, um, uh, you know, I, I hope that major AI companies have, um, monitoring setups, uh, going where, you know, for example, AIs are looking over transcripts of, of AIs performing work to see if anything suspicious is going on.

  459. 49:17

    And, you know, it's not-- it, it won't be impossible, perhaps, I conjecture, to, to fool those monitors, but, you know, be challenging to, to, to fool those monitors. So, you know, maybe you can do,

  460. 49:29

    you know, forty-hour tasks without-- forty-hour, you know, dangerous capability relevant tasks, um, without any monitoring, but only like two-hour-long tasks if you're, if you're being watched very, very closely.

  461. 49:40

    Being watched, like, if there's another model evaluating the output-

  462. 49:43

    That's the kind of thing I'm talking about. Maybe you escalate to humans and-

  463. 49:45

    With or without safety controls, you mean?

  464. 49:47

    Yeah, yeah, yeah. Roughly, yeah. Um, and, and then, you know, now, now we've bought a load more doublings, [laughs] and we can, we can sort of keep, keep doing the capability extrapolation thing.

  465. 49:56

    That, that's one of the kinds of things I'm thinking about in addition to creating ever longer tasks.

  466. 49:59

    Yeah. I mean, I imagine some of the model shops do have, like, evaluations of capability with and without safety because I'm sure that there, like, there's an argument between their researchers and their safety teams.

  467. 50:10

    Um, yeah, yeah, yeah. Um, yep. Um, um-

  468. 50:14

    Seem like I have seen something about this, but not a lot.

  469. 50:18

    Yeah, yeah. Um, yep. Um, um, yeah. I, I guess I think that,

  470. 50:26

    um, this m- might be sort of like an especially quantitatively important consideration or, um... I, I, I expect that it will r- reduce the effective time horizon by, uh, by like maybe an order of magnitude or two.

  471. 50:40

    Um, yeah, I, I agree that there's a-- there are some important senses in which there's not really a difference, difference in kind.

  472. 50:45

    Yeah. Of course, then I would also worry that, like, publishing that encourages people to, like, focus less on safety or to, like, try to argue against safety because of how it impacts capability.

  473. 50:54

    Yeah. I, yeah, I think there are lots of landmines- [laughs]

  474. 50:58

    -in, in, um, in, in all sorts of safety work, not just, not just in AI.

  475. 51:02

    Oh, of course. Of course.

  476. 51:04

    Um, okay. Next thing. Um, you know, we have this, we have this trend. I, I spoke about this at the beginning, but, you know, we, we have this trend.

  477. 51:12

    Is it going to continue forever? Is this, is this a fact of the universe or does it, you know, somehow depend on inputs or what you think about, um, intelligence explosions or, or something like that?

  478. 51:22

    Um, trying, trying to think about that. Where, where is this line, uh, actually going is, um, is a, is a pretty active area of work. Also, you know, the ways in which, um, this line or, or the, the particular points don't quite correspond to the thing I care about.

  479. 51:37

    So one obvious way is that, um, you know, these, these models are being judged according to, um, uh, you know, I, I think, I think the, um, algorithmic scoring that we use on, on METR tasks is, is, um, is importantly sort of more robust or more covering the relevant concerns.

  480. 51:56

    That might be the case in just sort of SWE-Bench using unit tests, but, but it still sort of-- it still has a lot of the same character. Um, there are, um, you know, considerations like being able to build on this work in future outside of the immediate problem, um, uh, uh, facing you that, that aren't being captured

  481. 52:12

    by, by METR scoring. And maybe if you did capture that, you know, you'd get something a little bit like going from fifty percent success to eighty percent success. You know, you could do hour-long tasks if-- it doesn't matter whether you can build on the work, but, you know, only thirty-minute tasks if it doesn't matter whether you can

  482. 52:26

    build on the work. But bringing, bringing these numbers again to, to something I care about a little bit more. And then, yeah, pro-projecting out both if there are compute slowdowns, um, if, if we are going to enter some regime where, um, uh, AIs are building AIs and, and that leads to some sort of steeping of the curve,

  483. 52:43

    these, these kind of considerations. That's another thing I'm thinking about.

  484. 52:48

    Um, da, da, da, da, da. Oh, and then capabilities measurement from new angles. So here's, um, you know, here's, here's one history of METR that I think is not the accepted history and, um, also probably, um, not a very accurate history.

  485. 53:04

    Certainly not the most accurate history, but, but here's one possible telling.

  486. 53:08

    Um, you know, near the beginning, METR has early access to-- wh- when I wasn't there, and I have sort of no internal knowledge of this. When METR has early access to GPT-4, um, and there are just sort of Q&A datasets going on everywhere or like LSAT datasets or something.

  487. 53:23

    You're like, you know, "Can-- GPT-4, it like seems so smart relative to stuff that, that went before. Can it do stuff?" You know? So you like, you try it out with some tasks.

  488. 53:32

    "Can it, can it do stuff?" And the answer is, you know, it can do some stuff, and it can't do other stuff.

  489. 53:37

    Um, and, um, and people are like, "Oh, that's cool." You know, you've tried this, you've tried this, um, neat new kind of thing, getting models to do stuff instead of, instead of answering questions.

  490. 53:45

    And then, and then later you're like, well, different models. You know, they come out over time. You know, this model comes out in January. This model comes out in February.

  491. 53:51

    Can they do different kinds of stuff if we test them on the same-- if we test them on the same stuff? Then we'll try and think of kind of the most obvious, in some ways, summary statistic of whether they can do stuff.

  492. 54:00

    This like single, single, um, data point or number that reflects whether they can do stuff, the time horizon, plus it over time and see what happens. And you're like, "Oh, that's, that's kind of interesting."

  493. 54:08

    And then you're like, well, what's the next sort of- In some sense, kind of dumbest or, like, most obvious thing you can do. Well, we'll run kind of the most obvious RCT design.

  494. 54:17

    We'll, like, allow AI or not allow AI, and then we'll see, we'll see what happens, and we'll try and... You know, it'll be, it'll be messy. There's lots of, um, th- there are lots of methodological problems that, that people point out, as, as there are with this work.

  495. 54:29

    But they're different kinds of problems. You know, they have different pros and different cons. And maybe with these sort of two different things that give two different answers and have two different sets of pros and cons, we can kind of triangulate the truth from that.

  496. 54:39

    And then now I'm like, well, can, can we pull that rabbit out of the hat one more, one more time? Are there-- or multiple more times. Are, are there other sources of evidence that have, you know, different pros and cons that I, that I won't believe in fully, but they're different pros and cons, and they might give

  497. 54:53

    different answers, and so on and so forth? Um, here are two suggestions for things I'm curious about at the moment. The first is, um, in the wild transcripts. So, you know, agents in Cursor, in Claw Code and, and, and whatever other, other, um, products or services, um, they leave behind traces, um, traces of the diffs that they've,

  498. 55:16

    um, contributed to codes or, or diffs of their actions and their reasoning chains and, and, and so on and so forth. Um, the traces that they leave in the wild are, you know, importantly different from this, where it's more kind of contained and, you know, the tasks sort of neatly packaged and stuff.

  499. 55:32

    This is going to be, you know, like the, like the example with the many different columns that are very confusing. This is gonna be like whatever real crap shows up in the wild, how, how do models handle that?

  500. 55:41

    Um, there are important reasons why you shouldn't believe that kind of information. It's, it's, like, not very experimental. It's, like, hard to know exactly what to make of it.

  501. 55:49

    But it does have these important pros that it's like, it's more real. It's, you know, the data's enormous. Perhaps, um, the, the data on transcripts is enormous. Um, you know, perhaps there's a lot you can learn there.

  502. 55:58

    That's, that's one thing. And then, and then here's another one. There's this, um, there's this group which you guys should check out called, um, Agent Village.

  503. 56:08

    AI Village, sorry. Um, where they, um, they have, um, a, a lot of different models or, or agents kind of living in this village, occasionally talking to humans, trying to accomplish, um, fuzzy goals that are, that are set to them, basically using compute use.

  504. 56:27

    They try and do stuff like, you know, organize this event at the park, or, um, uh, run a few human subjects experiments, or run this merch store, you know, stuff, stuff like that, that's not so clearly specified.

  505. 56:38

    And basically all the time they find that the models fall on their faces and suck. Um, and there are lots of reasons not to believe this evidence.

  506. 56:49

    You know, here are some of the reasons. Number one, um, it is using compute use, and I think compute use is just way worse than CLI-based-- C-c-compute use capabilities are, are considerably worse than CLI-based stuff at the moment, or text-based things in general at the moment.

  507. 57:01

    And maybe you care more about text-based things because that's more relevant to various types of things you care about, and also lots of GUI, GUI-based things, um, can be converted into text things.

  508. 57:11

    Um, it's, um, you know, there's all these different models hanging around in the village. I'm like, why, why are there so many models? Like, why is there a village instead of just, like, some big agent orchestration set up?

  509. 57:20

    I don't re- I don't really understand what's going on there. And, um, anyway, lo- lots of reasons not to believe it. But on the other hand, it is models doing stuff in the world.

  510. 57:30

    It's not benchmark-style tasks. It's like trying to accomplish some goal, and they can't accomplish even sort of, you know, very basic subsets of the goal. And I feel like that's extremely interesting.

  511. 57:40

    And I, I wonder if you could get rid of some of the most obvious cons, you know, make this only text-based. Give them some, um, uh, relevant text-based tools.

  512. 57:48

    Work a bunch on the elicitation to make, to make these models sort of more performant. Get rid of the less performant models in, in the village, so on and so forth.

  513. 57:55

    But then try and get them to do these fuzzy goals. Um, and, you know, just observe, like, where do they mess up? Like, you know, they, they, they, they went about step one, it went great.

  514. 58:06

    But then they sort of, they became incoherent or they, you know, went into a strange psychological basin with one of the other models. Or, you know, they, they weren't able to interact with external services in an appropriate way or, or, or figure out their resource use.

  515. 58:18

    You know, I'd be very interested just kind of qualitatively in what's-- in what goes on when you do that. A- again, keeping in mind that we're interested in, um, the ability of, um, at least...

  516. 58:29

    At, at the moment, I'm most interested in the ability of AIs to, um, automate R&D and, you know, or speaking to why that's not the case at the moment and why that might not be the case in the near future.

  517. 58:37

    Some-something shaped like this seems like it might be, might be kind of-- might curiously point to, to why that's not the case. I'm not sure exactly what's there, but yeah.

  518. 58:47

    And my observation is that they, they are effectively neurodivergent individuals, right? And none of our world was not built for that.

  519. 58:58

    Yeah.

  520. 58:59

    There's-- Everything that we have that are, that are defined for a human to do, they're shaped and sized to humans. Just like, you know, the military, like, you know, how big are packs?

  521. 59:07

    Well, it's based on how much they think a person can reasonably carry, right? And how much we expect someone to handle for their taxes, that's based on what we think a human can do.

  522. 59:15

    And-

  523. 59:15

    Well-

  524. 59:16

    And, and there are-- And if you think about neurodivergent individuals, they struggle with challenges with the way the world's expectations don't align with them. And compared to a neurodivergent individual, these, you know, these intelligences are really, really different, right?

  525. 59:31

    And so all of the rough edges where they don't align with our world, that's why they need an assistant, a human assistant, in order to accomplish anything real in our world.

  526. 59:40

    It's just too hard, uh, for them-

  527. 59:42

    Currently.

  528. 59:44

    What?

  529. 59:44

    Currently.

  530. 59:44

    Yeah. I think someday it could change.

  531. 59:47

    Okay.

  532. 59:47

    Now it's-

  533. 59:48

    Yeah.

  534. 59:48

    They're just hopeless.

  535. 59:49

    Yeah.

  536. 59:49

    Right? They have to get really, really good or the world will have to change. One of those two things.

  537. 59:54

    You know, I, I agree. I, like, so strongly share this sense. But, you know, but if you ask me to really pin down, like, why, why exactly is that the case again, when they're like, you know, beating all the GPQA, G-GPQA experts on these extremely hard science questions and they're, you know, blah, blah, blah.

  538. 1:00:09

    Like, exactly what the hell? Why are they not able to accomplish things in the world?

  539. 1:00:13

    You ever met a neurodivergent individual who wasn't terribly good at some things- [laughs] And completely useless at getting through life?

  540. 1:00:19

    Yeah, yeah. They're all very good at reading books.

  541. 1:00:21

    Yeah. [laughs]

  542. 1:00:23

    There's a lot of those people in the world. I-

  543. 1:00:25

    Yeah.

  544. 1:00:26

    It's not that surprising.

  545. 1:00:27

    Mm.

  546. 1:00:28

    Although my, my only feeling about AI village is it's like, "Well, today is the two hundredth day my car didn't rocket off the Earth in escape velocity and fly to the moon."

  547. 1:00:36

    Mm.

  548. 1:00:37

    Like, that's because you didn't build a rocket yet.

  549. 1:00:42

    Yeah. I mean, I think there, there was a lot of talk a year ago about, you know...

  550. 1:00:47

    M-maybe I'm mischaracterizing, but I thought there was a lot of talk a year ago about computer use capabilities being impressive today.

  551. 1:00:53

    There was. There was a lot of talk about it, and yet I have talked to almost nobody who has used them for any practical-

  552. 1:00:59

    Yeah. [laughs] [laughs] Totally, totally. Um, yeah. But if we, if we move this to text only, and it seems reasonable to complete text only, um, uh, you know, would you still have the rocket concern?

  553. 1:01:11

    No. I wouldn't... I wouldn't really.

  554. 1:01:13

    Yeah. Okay.

  555. 1:01:13

    I do... Well, it depends on what the task was.

  556. 1:01:15

    Sure. Yeah. At least, yeah, the kind of thing that you could-- that a human could do over CLI anyway.

  557. 1:01:22

    So I, I think this, um, this relates to the, uh, Anthropic talk that-

  558. 1:01:26

    Mm.

  559. 1:01:27

    -earlier today, where they talked about how, um, you know, one way to, uh, use, uh,

  560. 1:01:34

    effectively is to give them... If you have a task, like, figure out a way to present the task or transform the task to something that is industry fit, you know, for the model.

  561. 1:01:43

    And I feel like this conversation kind of, you know, ties in on that. Like, um, you know, it interacting with, with Chrome is less in distribution than a COI.

  562. 1:01:52

    So I, I think that could be an interesting area of research is, like, you know, uh, okay, so if you're interested in exploring, like, how well can it perform these really open-ended tasks, like, first, I, I guess creating harnesses and creating an interface that is much more in distribution for them.

  563. 1:02:08

    So that way that's l- you know, less of a, a concern.

  564. 1:02:13

    Yeah. I mean, I, I think also speaks to the point about quote-unquote "neurodivergence models." Um, you know, there's some, uh, it's not so different from management scale or something.

  565. 1:02:22

    Giving, you know, giving appropriately scoped tasks to, to your, to your very talented interns or, uh, very talented neurodivergent interns or something, something like that. I do, I do think that's right.

  566. 1:02:31

    From the... Sorry to be, uh, you know, um, uh, s-sorry, sorry to be so repetitive. From the perspective of capability explosions, um, and automating R&D. You know, I think maybe the models will get extremely good at, um, scoping tasks for themselves such that it's benchmark style or, or, or, or something like that.

  567. 1:02:51

    But, you know, if they can't do that, I'm like, well, there's a lot of things that aren't-- that don't look like benchmarks that crop up in the real world, and you do need to be able to kind of flexibly work with that if you're to, um, do something as complicated as automate a major AI company.

  568. 1:03:04

    Um, um, and, you know, so, so I do, I do think it's, um... Yeah. I think, I think it can both be the case that the AIs are incredibly performant, um, on some particular type of problem, or if you make other types of problems more similar in scope or shape to, to the type of problem that they're

  569. 1:03:22

    best at. And, and also that they, you know, can't flexibly substitute for human workers because that requires, you know, um, yourself setting up the problem in, in, in a way that's appropriate or, or not having those constraints yourself.

  570. 1:03:34

    Yeah. It is interesting, though. Just, just to your point about new capabilities is thinking of all other axis on the graph that you have.

  571. 1:03:40

    Mm.

  572. 1:03:41

    Because I think there's not just-- I, I wonder if there's not just a time horizon issue, but there's a, a task category or a type of work category. Like, like, as your example of compute, like computer use is one of those examples, right?

  573. 1:03:54

    Like, if we think about the capability of computer use versus-- or capability that would require computer use versus the capability that could become-- can be accomplished entirely in text.

  574. 1:04:04

    Yeah. So yeah, sure. But, but, but like a lot of these are like, like almost all these benchmarks are basically text.

  575. 1:04:11

    Um, yes. Yes. Yes. And indeed, you know, the ones, the ones that aren't, the ones that require sort of, um, vision capabilities are, are, are notably lacking behind. Yeah.

  576. 1:04:20

    I, I, um, I'm not sure exactly what to, what to make of this graph. I think one thing I make is that, yeah, one thing I make of it is that, um, uh, you know, there probably is maybe not so much variation in, in sort of slope or doubling time across task dis-task distributions.

  577. 1:04:38

    I think there's only weak evidence for that. But, you know, in, in insects or, you know, the base of where we are now, um, yeah, there, there's, there's possibly a, a great deal of variety, especially on this sort of, um, uh, image-like capabilities versus not to mention.

  578. 1:04:51

    But, but physical ability is even more, you know.

  579. 1:04:55

    Yeah. Right. So there's exactly like... So I mean, you could even go through senses, right? Like, you, you could go through like a tactile... Like, like today, like they would all score zero.

  580. 1:05:05

    Nothing has tactile. So like it can't tell you anything about anything tactile.

  581. 1:05:10

    Um, well, you know, in producing this graph, we, you know, we try and make the models as performant as possible on some held-out set. Uh, so we-

  582. 1:05:17

    I see what you mean.

  583. 1:05:17

    You know, we try and give them some tactile stuff. [laughs] I'm not sure they perform zero. [laughs]

  584. 1:05:22

    Sure. Sure. Sure. I mean, space, we do have some examples-

  585. 1:05:25

    Yeah.

  586. 1:05:25

    -of products.

  587. 1:05:26

    Yeah. Yeah.

  588. 1:05:28

    Like space judgments, spatial judgments, things like that.

  589. 1:05:30

    Yeah.

  590. 1:05:31

    Um, you know, we've, we've obviously seen from figure, fine control and stuff like that with robotic. It's interesting. I, I haven't even-- I don't even know if anybody, maybe somebody has listed out what all of the capabilities that we would expect in the future.

  591. 1:05:47

    Like, if we actually wanted AGI, what is the entire list of cap...?

  592. 1:05:52

    That's a way to start a debate that doesn't end. [laughs]

  593. 1:05:54

    I think it's-

  594. 1:05:55

    Basil Halperin and Arjun Ramani hopefully have a paper on this out in a small number of months.

  595. 1:06:01

    Yeah. And then be interesting to think about where are we at, and do all of the capabilities follow the same-- all the capabilities that we currently measure, do they follow the same, uh, log?

  596. 1:06:12

    Yeah. It, it does seem like a reasonable null hy-hypothesis to, to you as well as me, I think. Not, not, not, not certainty. I mean, who knows? Yeah. Yeah.

  597. 1:06:24

    Um, oh, there was something, there was something I wanted to add there.

  598. 1:06:28

    Um, oh, oh, here, yeah. Here's another thing I'm thinking about, not super in a research capacity, although kind of. Um,

  599. 1:06:36

    um, so, you know, some people like me are sort of skeptical of, of, um, software-only singularity. That is the, the, the idea that you could automate AI research without also automating, um, uh, chip design and maybe also chip production as well.

  600. 1:06:51

    Um, that you'd quickly get bottlenecks by, by computes 'cause there are only... For fixed hardware, there are only so, sort of so many experiments that you can run that, that would be, that would be, um, sufficiently productive to, to, uh, to foom progress upwards.

  601. 1:07:04

    But, you know, even for people like me who are skeptical of that, um, uh, you know, you, you might think that in fact, like, chip production is going to get automated.

  602. 1:07:12

    You know, the robots, like [laughs] they're, they're coming. They can, they can do, they can do the stuff that humans do and then, and then maybe you really do have a fully self-sustaining, um, uh, robots plus AI economy.

  603. 1:07:25

    And so, and, you know, and so you, you, uh, you have some slow trends from, from compute slowing down, but then you have sort of a fooming back up once, once the whole thing is, is, um, is, is in a tight loop.

  604. 1:07:35

    Um, one, one interesting debate that I, uh, heard about recently and would like to think more is, um, uh, you know, I think there's... I- in the public discussion, there's some sense that, you know, why, why are robotics capabilities lagging, um, uh, lagging LLM-like capabilities so much?

  605. 1:07:52

    Well, it's to do with training data or something, something, something like that or, or maybe it's to do with hardware constraints.

  606. 1:07:58

    Yeah.

  607. 1:07:58

    I'm, I'm curious if it's not to do with hardware constraints. Like, what, what, what, what exactly are these hardware constraints? If we put super intelligence inside, hypothetical super intelligence, in- inside of, you know, um, hardware parts that existed today, could it build

  608. 1:08:15

    chip production facilities? And I, I have no idea because I'm su- you know, I'm, I'm beyond, beyond, beyond novice. But it's not obvious to me what the, what the answer is.

  609. 1:08:25

    I think it's, I think it's kind of plausible. I'm not sure you need this, like, um, uh, yeah, I'm not sure you need this, like, very flexible fine motor control in order to do it.

  610. 1:08:33

    Also, I think maybe the fine motor control is there subject to having super intelligence controlling it.

  611. 1:08:38

    I mean, to be fair, like, the key aspects of chip production are done by robots.

  612. 1:08:44

    Um, oh, but, but, but I'm also thinking, like, building the robots and-

  613. 1:08:48

    Yeah

  614. 1:08:48

    ... the whole, you know-

  615. 1:08:49

    And, and that's where I'll tell you, I, I have, I have a friend who spent most of his career doing software development. But during COVID started working on manufacturing things like PAPRs and things like that to help people, and he found out how hard the manufacturing world is and how slow the iteration process is.

  616. 1:09:06

    Mm.

  617. 1:09:07

    And it is really, like, he put it, like, he, he knew it was gonna be worse. He didn't understand that it was, like, next level, like an order of magnitude worse.

  618. 1:09:15

    And I think that probably, like, you know, we, we... From our perspective, people who don't do it, it seems like, "Ugh, how bad can it be," right? And it's...

  619. 1:09:23

    The, the feedback I've had from everybody who actually works in that space is it's way, way different.

  620. 1:09:28

    That's what I've heard as well.

  621. 1:09:29

    Yeah.

  622. 1:09:30

    I've only talked a little bit with, like, people who work in fabs and stuff, but I, I was surprised when I did talk to them of the level of human expertise required-

  623. 1:09:39

    Yeah

  624. 1:09:39

    ... in order to work at the fabs. Like, a lot of those jobs are, like, fairly high-paying actual engineering jobs in order to, like, successfully do this.

  625. 1:09:45

    Also the rate of improvement is actually glacial, right, compared, compared to software, right? I think also because it's cost a billion dollars to build a fab.

  626. 1:09:53

    That's what I'm saying.

  627. 1:09:53

    Like, each iteration is a huge cost of time, money. It, it's brutal.

  628. 1:09:58

    Right.

  629. 1:09:58

    So it's-- I think that's why it's been hard to get it all the way there is just, like, give, give them a couple more centuries, maybe they can get it done [laughs]

  630. 1:10:06

    you know?

  631. 1:10:06

    Is that, is that really your view? Centuries, centuries.

  632. 1:10:09

    I, I do. I do. I, I do think... I, I'm skeptical like you about-

  633. 1:10:12

    Yeah

  634. 1:10:12

    ... like how easy some of these tasks are.

  635. 1:10:15

    Yeah.

  636. 1:10:16

    Um, um, we think they're easy, but in my experience, like, I, I-

  637. 1:10:20

    Oh, that's interesting

  638. 1:10:20

    ... I remember when the self-driving thing came out and when people were, like, pushing that and it was... I actually worked in that space for a while and it was like, I get that we can get really close to it, but getting all the way to something that is acceptable is extremely difficult, right?

  639. 1:10:35

    Uh, uh, and we underestimate how much work is involved in getting that last little bit done. The first 90%, I, I knew we could do it with computers like, you know, 10 years ago pretty much, but getting the last bit that everyone's happy with it-

  640. 1:10:49

    Yeah

  641. 1:10:49

    ... usually a lot of work.

  642. 1:10:50

    I feel this myself. You know, I didn't get a driver's license when I, when I [laughs] got to because, uh, because I expected self-driving cars to come. Um, yeah, I think, I think totally.

  643. 1:10:58

    But it hasn't been that long, you know. And they're, they're expanding to, to, um, to, to the entire Bay Area

  644. 1:11:04

    Yeah, they're, they're-

  645. 1:11:05

    Cities

  646. 1:11:05

    ... they're gonna get there. I don't think it's gonna take a hundred years.

  647. 1:11:07

    Like, is, is the, is the robot economy building the chip productions going to take centuries? [laughs]

  648. 1:11:12

    I don't know. I don't know about... Well, I c- I could see that it might take... It, it's, it's... So part of the trick with self-driving is the economic incentive is moving it along faster, right?

  649. 1:11:22

    And probably the robot building robots kind of thing would also. But, like-

  650. 1:11:27

    Yeah

  651. 1:11:27

    ... you know, where we're at right now is, like, rip rap is kind of as far along as we've got of robots building robots, right? Which is-

  652. 1:11:35

    Oh. Oh, but I, I, I feel like the, you know, I... Is that, is that paying sufficient attention to the charts? GPT-2, 2019. [laughs]

  653. 1:11:43

    Yeah.

  654. 1:11:44

    It's so, it's so recent. Uh, you know, w- I, I, I have some... This is so, this is so-

  655. 1:11:49

    Yeah, yeah. It's true

  656. 1:11:50

    ... um, uh, nonsensical, but I'm like, maybe we're in a sort of GPT-2 moment.

  657. 1:11:55

    Yeah.

  658. 1:11:55

    GPT-2 moment.

  659. 1:11:56

    No, it's a fair point. I, I could be wrong. It's just my guess is it's gonna take a lot longer than we think.

  660. 1:12:01

    Totally.

  661. 1:12:01

    At least to be able to do, like, real mass production-

  662. 1:12:04

    Yeah

  663. 1:12:04

    ... uh, at a scale that, that causes the kind of global impact that you're talking about.

  664. 1:12:10

    Yep.

  665. 1:12:11

    Right? Like, that, that's... I, I think they can already do a great job building one-offs, right? Robots are very good at build- doing one-off builds-

  666. 1:12:18

    Yep

  667. 1:12:19

    ... at a small scale, but-

  668. 1:12:20

    Yep

  669. 1:12:21

    ... it's totally impractical for doing it at a large scale.

  670. 1:12:25

    There is, um, um, um-

  671. 1:12:34

    Mm-hmm.

  672. 1:12:35

    So one, one fact that I think is kind of remarkable

  673. 1:12:39

    is this... Maybe it's this. Is that the rate of... Is it this? Yeah, yeah, yeah. The rate of compute put to robotics models-

  674. 1:12:49

    Mm.

  675. 1:12:50

    -lags behind, um, sorry, is, is, is about the same, but the, but the level's, uh, two orders of magnitude difference. Um, I, I am kind of, um, curious if that's, if that gap closed, um, um, uh, what we'd, what we'd see.

  676. 1:13:07

    It, it does seem like at least sort of more capable robots are in some sense, um, very on the table as something that could be the case ve-very soon if this, if this...

  677. 1:13:17

    No, I'm, I'm not saying all the way. I'm certainly not saying chip production. It just does seem like there's some sort of data hang.

  678. 1:13:23

    Yeah. Yeah.

  679. 1:13:24

    Intuitively.

  680. 1:13:26

    That's interesting.

  681. 1:13:27

    Um, a-a-also, also thinking some sort of, um, um, some, like you don't just need to be scaling data. You can also scale parameters, use the same amount of data, you know, flexible ways to use compute to, to close some gap.

  682. 1:13:42

    Interesting.

  683. 1:13:44

    Yeah.

  684. 1:13:44

    So one and three just give me a very interesting overview of how METR AI is going into fabrication at the very time.

  685. 1:13:51

    And, and what does it say?

  686. 1:13:53

    Um, so it says, so it says there's a lot of areas where right now it's going to help probably pretty dramatically in the near future, and a lot of it's in computational aspects.

  687. 1:14:01

    There's a lot of computational aspects that are extremely expensive, designing like the mask, basically the,

  688. 1:14:08

    the hole that you're using for the laser to get the transistors.

  689. 1:14:12

    Mm.

  690. 1:14:12

    Um, and like calculating that, how to build it and, and ensuring that it conforms to the spec that you've written basically is extremely computationally expensive. Um, and there's a lot of opportunity for AI to help there.

  691. 1:14:26

    Um, and there's also theoretically the possibility for... So like chip-- obviously, uh, chip manufacture is extremely, uh, precise but also fragile, and the opportunity for an AI to detect parameters that are basically out of whack and leading to failure, potential failure in like, uh, imaging a wafer, uh, is could theoretically dramatically improve yield, and yield is a

  692. 1:14:50

    big problem in fab-- in chip manufacture. Like the reason that you get different speeds out of your CPUs is because they actually just have the one line that produces all those CPUs, and some of them come out better and some of them come out worse.

  693. 1:15:02

    And that's why the higher G... That's why the higher gigahertz models are more expensive than the lower gigahertz. Like, like if you have like your NVIDIA, like your home GPUs, your, your fifty-- forty or fifty, fifty or fifty, sixty or fifty, seventy or fifty, eighty, fifty, ninety are all the same chip.

  694. 1:15:17

    Right.

  695. 1:15:17

    That just have different quality.

  696. 1:15:18

    Different, different levels of tolerance, essentially.

  697. 1:15:21

    Yeah. Um, but the problem is that, uh- I'm gonna cut the recording. They're gonna actually cuss out soon, but feel free to continue discussion. Yeah.

  698. 1:15:31

    Oh, cool.

  699. 1:15:32

    You can also hang out here, but I'm just gonna-

  700. 1:15:33

    Sure. [upbeat music]