Mousepower: agents that can’t be measured, can’t be managed. — Maximillian Piras, Yutori

Read the talk

Mousepower: measuring agent value through outcomes and verification

Maximillian Piras explains why agent builders need understandable comparisons, measurable outcomes, and tasks whose results are easier to check than produce.

From a talk by Maximillian Piras

At a glance

Ideas worth remembering

  • Agent value needs a customer-understandable baseline and a trace from token spending to useful outcomes, such as bugs fixed or support requests closed.

  • Faster generation can move the bottleneck to review. A product must help users verify quality before they can determine whether its spending was justified.

  • The proposed matrix favors moderately uncertain execution with clear acceptance criteria. Predictable steps favor scripts; very uncertain steps risk unreliable execution; unclear criteria can make checking require doing the work again.

  • Mousepower is a conceptual guide rather than a defined metric. Repeatable checking may support verifier agents, but the talk presents this direction as a developing idea.

Parallel work makes the cost question unavoidable

Selected presentation frame from Mousepower: agents that can’t be measured, can’t be managed. — Maximillian Piras, Yutori at 130 seconds
Parallel work makes the cost question unavoidable

Maximillian Piras opens with a familiar agent workflow: keep human attention on one activity while agents advance peripheral tasks. His example is preparing better slides while delivering the talk. With a design system and guidance already supplied, an agent can explore typography and layout; several agents can explore different directions in parallel. The attraction is more exploration without moving his own attention away from the immediate task.

This workflow encourages expansion because research and design rarely feel exhausted. There is always another direction to investigate, and background execution makes additional work feel easy to request. The bill changes the question: were those explorations worth their token cost, and could better sequencing have achieved the same value more efficiently? Piras frames the talk around both evaluating that cost and helping customers understand what they received for it.

0:310:33
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:01 · section reference included

Computer use expands access, but customers still need a mental model

Selected presentation frame from Mousepower: agents that can’t be measured, can’t be managed. — Maximillian Piras, Yutori at 169 seconds
Computer use expands access, but customers still need a mental model

At Yutori, Piras works on computer use models: models that operate a computer as a person would. He describes their role when an API or MCP cannot provide the needed information. An agent can interact with the interface to retrieve and manipulate data that would otherwise be inaccessible through those integrations. He explicitly treats this as a less efficient, last-resort approach compared with APIs and MCPs. The benefit is access to a workflow that lacks a direct integration.

He introduces an example of an agent visiting its own company's website and checking its benchmark, then turns to the customer discussions that inform his design work. Customers are excited but often feel they are only scratching the surface. They ask which use cases fit agents and how token costs compare with value. These questions motivate his thesis that agents have a measurement problem: users lack an intuitive basis for choosing work and judging whether delegation paid off.

An enthusiastic technical audience may already feel comfortable managing a fleet of agents, but Piras cautions against generalizing from that experience. Early adopters are unusually motivated to explore the technology, and people who directly or indirectly sell tokens have an additional reason to favor its use. Their habits need not represent customers who have never used an agent. His personal example is his spouse, who still uses ChatGPT by copying and pasting and has resisted his attempts to set up agents. Familiarity and excitement among builders do not establish that the value is legible to everyone else.

2:362:38
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

2:36 · section reference included

Horsepower makes an unfamiliar improvement understandable

Selected presentation frame from Mousepower: agents that can’t be measured, can’t be managed. — Maximillian Piras, Yutori at 374 seconds
Horsepower makes an unfamiliar improvement understandable

Piras develops a historical analogy through his account of James Watt selling steam engines in the 1700s. The application was replacing a horse gin, a power source for a mill. A horse attached to a rotary arm walked in a circle, producing the motion needed to run the machinery. Piras uses a brewery as an illustrative setting, while acknowledging that he does not know its exact production process. The central mechanism is the conversion of an animal's circular movement into mechanical power.

In this account, the adoption barrier was partly conceptual and emotional. People understood horses; a cold, intimidating machine came with claims of improvement that they could not readily evaluate. Watt's response was to study the familiar power source and establish a baseline. Horsepower gave him a way to express the steam engine's efficiency as a multiplier of something customers already understood.

Piras emphasizes that the measure, as he describes it, was neither very scientific nor necessarily accurate. Its contribution was making a potential increase in value understandable enough to encourage a trial. That is a narrower claim than guaranteeing the exact improvement someone would realize in practice. The analogy also recognizes attachment to an existing way of working: efficiency must be communicated in terms that help people weigh a return on investment against their preference for the familiar.

5:575:59
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

5:42 · section reference included

Connect token spending to completed work

Selected presentation frame from Mousepower: agents that can’t be measured, can’t be managed. — Maximillian Piras, Yutori at 676 seconds
Connect token spending to completed work

Piras argues that even technically capable users struggle to calculate agent returns. He describes examples of an annual token budget being consumed in a quarter and token leaderboards encouraging consumption without establishing corresponding value. He calls the resulting cycle overspending and underusing, borrowing the term from Ramp: intense use leads to austerity, users withdraw, and fear of missing out eventually draws them back. A better value measure should help users sustain useful work rather than alternate between enthusiasm and retreat.

He cites a chart posted by Coinbase's CEO in which AI spending reportedly began to diverge from token usage. Piras considers that a useful start, but the example supplies no numerical savings or quantified outcome improvement. Lower spending for a given amount of token activity still leaves the value of the work unresolved.

Tokens are useful for measuring an internal system, but Piras treats them as an output that must be traced to an outcome. His examples are bugs fixed and support requests closed. Those outcomes must then connect to progress on objectives. This creates a chain from spending, through completed work, to the purpose of that work. Counting tokens alone cannot show whether the expense advanced the customer's goals.

9:119:15
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

9:08 · section reference included

Faster execution moves the bottleneck to verification

Selected presentation frame from Mousepower: agents that can’t be measured, can’t be managed. — Maximillian Piras, Yutori at 818 seconds
Faster execution moves the bottleneck to verification

Coding provides Piras's clearest example of why producing more work does not automatically produce more realized value. Agents generate code quickly, but the human review process has not scaled at the same pace. He describes a growing burden of pull requests and cites claims from Anthropic that distinguish progress on coding from the unresolved problem of code review. In his argument, generation speed shifts the bottleneck downstream: reviewers must still determine whether the code is good.

This is also a measurement bottleneck. If users cannot judge quality as quickly as agents produce code, they cannot promptly establish that enough of the output is useful to justify the spending. Piras allows that review may already have been flawed and that agents are exposing an existing weakness. He suggests revisiting the assumptions beneath review rather than assuming that the established process will absorb a much larger volume of generated work.

He does not offer a solution to code review. Instead, he identifies a reason it appears tractable: a shared culture of review provides common assumptions that can converge into a clear rubric. Adapting that rubric could support measurement at computer speed alongside execution at computer speed. The broader product requirement is to give customers a method for verifying output, with measures that fit their mental models, so they can calculate returns and justify continued spending.

11:4011:42
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

11:29 · section reference included

Mousepower is a comparison principle, not a cursor metric

Selected presentation frame from Mousepower: agents that can’t be measured, can’t be managed. — Maximillian Piras, Yutori at 997 seconds
Mousepower is a comparison principle, not a cursor metric

Mousepower is Piras's proposed analogue to horsepower: establish a baseline for how people use computers, then communicate how an agent improves performance on a particular task. He leaves room for improvement along different dimensions rather than assuming one universal measure of efficiency. The comparison should help a customer understand what the agent contributes relative to familiar work.

He illustrates the limits of a literal interpretation with a joke measurement device that he says Claude helped him code. Measuring how fast a cursor moves across a screen might appear to supply a clean human–agent comparison, but he rejects it as a way to measure useful performance. Information work occupies a space with too many dimensions for cursor movement to capture its value. He explicitly says mousepower is an idea rather than a metric: agent builders should supply a rubric for determining whether the agent did good work and whether the tokens were worth buying.

Piras acknowledges that he cannot give a general method for choosing the right measurements for every customer. He instead introduces a task-selection principle inspired by Claude Shannon's ideas about entropy and information. He describes entropy in terms of uncertainty in a probability distribution and mentions cross entropy in connection with agent training. His next move is to apply uncertainty to the task itself, rather than consider only the agent's capabilities.

14:2614:29
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

14:26 · section reference included

Separate uncertainty in the steps from uncertainty in success

Selected presentation frame from Mousepower: agents that can’t be measured, can’t be managed. — Maximillian Piras, Yutori at 1114 seconds
Separate uncertainty in the steps from uncertainty in success

The matrix's horizontal axis represents uncertainty in the steps needed to perform a task. A valuable goal alone does not establish that an agent is a good way to pursue it; the pathway matters too. Piras contrasts booking a flight with painting a masterpiece. A flight purchase has recognizable requirements, including departure and arrival destinations and a seat selection, whether chosen by the person or assigned. Painting a masterpiece has no comparably predictable sequence of steps.

The other axis represents uncertainty in the acceptance criteria. This asks whether there is a clear rubric for grading the result, independently of whether the agent can carry out the work. Considering both axes separates two sources of difficulty: finding a path to completion and recognizing a successful result. Their intersection is intended to guide which tasks agent builders choose.

At the low end of step uncertainty, Piras recommends writing a script: a predictable pathway offers little reason to spend tokens on an agent. At the high end, he warns that unpredictable tasks may fall outside the distribution encountered in pre-training and may provide sparse rewards for reinforcement learning. He presents these as risks, rather than proving that every highly uncertain task is outside a model's abilities. The practical question is whether the task demands useful adaptation or asks the agent to navigate a pathway it cannot reliably model.

16:4216:45
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

16:42 · section reference included

Choose work that is easier to verify than execute

The closing argument adds the cost of acceptance to the cost of execution. When acceptance criteria are highly uncertain, verifying usefulness can become indistinguishable from doing the task. A person may have to perform the work again to decide whether the agent's result was valuable. In that situation, delegation preserves much of the original human burden while adding token spending.

Piras therefore favors tasks with enough uncertainty in execution to justify an agent, but not so much that the pathway becomes unmanageable, together with results that are relatively easy to validate. He describes this as the shape of an NP-style problem: easier to verify than to execute. Here the comparison expresses a practical asymmetry between doing and checking; he does not establish a formal computational classification for the tasks.

A repeatable verification pattern creates a further opportunity: agents can perform the checking as well as the original work. Piras suggests building an agent that verifies another agent's output, so the product includes a way to assess its results. He closes by calling the matrix a thought starter still in development. It offers a direction for choosing tasks and designing verification, without demonstrating a validated measurement method or guaranteeing that an agent verifier will judge correctly.

19:0419:05
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

18:54 · section reference included

Read the complete timestamped transcript
  1. 0:01

    [music]

  2. 0:12

    >> Um well, thanks all for your time.

  3. 0:14

    Really appreciate you dropping by and

  4. 0:16

    it's always a a great honor to speak at

  5. 0:18

    the world's fair. So, I'll do my best to

  6. 0:21

    uh

  7. 0:21

    give you guys some valuable insights and

  8. 0:23

    um

  9. 0:24

    yeah, hopefully make it worth your time.

  10. 0:26

    So, my name's Maximilian Piros and today

  11. 0:28

    I'll be talking about mouse power.

  12. 0:31

    And this is a talk about measuring

  13. 0:33

    agents through mental models.

  14. 0:35

    But before I get into talking about

  15. 0:37

    measuring agents, I'm going to talk

  16. 0:39

    through a bit about how I use them every

  17. 0:40

    day. And it might seem familiar to you,

  18. 0:43

    but just to level set, we'll go through

  19. 0:44

    it. So, uh I tend to background them

  20. 0:46

    like I'm sure a lot of you people are as

  21. 0:48

    well.

  22. 0:49

    Um so, while my active attention is

  23. 0:51

    focusing on one thing, like perhaps

  24. 0:53

    giving this talk to you, I still want to

  25. 0:55

    make some progress on peripheral tasks.

  26. 0:57

    So, I'll keep my attention focused on

  27. 1:00

    giving this talk, while my agents can

  28. 1:03

    help me explore some designs in the

  29. 1:04

    background cuz I think that uh my slides

  30. 1:07

    need a bit of work. So, you know, I've

  31. 1:08

    got my design system already set up.

  32. 1:10

    I've got some guidance uh given to my

  33. 1:12

    agents. So, I'll kick off an agent to

  34. 1:14

    just try to explore some different

  35. 1:16

    directions on the type type two

  36. 1:17

    treatment and the layout. And um you

  37. 1:20

    know, just try to get as many

  38. 1:21

    explorations as possible.

  39. 1:23

    But of course, uh one agent's never

  40. 1:24

    enough. So, I like to kick off a bunch

  41. 1:26

    in parallel. You know, I've got a lot of

  42. 1:28

    slides to get through. So, I need all of

  43. 1:30

    my agents on and exploring it in

  44. 1:32

    different directions and

  45. 1:34

    uh hopefully I can get some interesting

  46. 1:35

    things to make my slides a bit better.

  47. 1:38

    And uh hopefully they can finish the job

  48. 1:40

    pretty soon cuz we're obviously kind of

  49. 1:42

    up against uh the timeline the the uh

  50. 1:44

    deadline here.

  51. 1:46

    So, um

  52. 1:47

    this is generally how I work. I'm sure

  53. 1:48

    it's probably familiar to a lot of you

  54. 1:50

    where we're just trying to kick off

  55. 1:51

    agents for as much as possible in

  56. 1:53

    parallel cuz it always feels like

  57. 1:55

    there's just way more research to do. We

  58. 1:57

    want it to be as thorough as possible.

  59. 1:59

    There's way more design explorations to

  60. 2:01

    do. So, whenever our main focus is on

  61. 2:03

    one thing, why not kick a bunch of

  62. 2:05

    agents off in parallel and just try to

  63. 2:07

    maximize your time. And it's a lot of

  64. 2:10

    fun, of course, until you get the bill.

  65. 2:13

    And then you start to wonder, was it all

  66. 2:15

    worth it? Right? Did

  67. 2:17

    did you vibe code too hard?

  68. 2:20

    Were you token maxing too much? Like,

  69. 2:22

    could you have been more efficient in

  70. 2:24

    how you approached

  71. 2:26

    your sequencing your agents?

  72. 2:28

    And so, this is what I'm going to get

  73. 2:29

    into today. It's It's how do we evaluate

  74. 2:31

    the token cost? And specifically, how do

  75. 2:33

    we help our customers value it?

  76. 2:36

    So, for the past year and a half, I've

  77. 2:38

    had the pleasure of working as a

  78. 2:39

    founding designer at a company called

  79. 2:40

    території and we focus on computer use

  80. 2:43

    models. These are models that learn to

  81. 2:45

    use a computer like a human would and

  82. 2:47

    the use case for them is when you can't

  83. 2:49

    get information from an API or an MCP,

  84. 2:52

    why not just send an agent out to use

  85. 2:54

    the computer like a human would and then

  86. 2:56

    we can extract all types of data and

  87. 2:58

    manipulate it in ways that let us access

  88. 3:01

    all the stuff that wasn't accessible

  89. 3:03

    previously. So, obviously less efficient

  90. 3:04

    than APIs and MCPs, but as a last

  91. 3:07

    resort, just have an agent go use the

  92. 3:09

    computer and try to get the information.

  93. 3:12

    Here's the U території agent using the U

  94. 3:14

    території website.

  95. 3:16

    Um checking out its own benchmark. So,

  96. 3:18

    kind of it's admire itself in a way. So,

  97. 3:21

    yeah, it gets it gets a bit weird. Like,

  98. 3:23

    and a lot of what I do as a founding

  99. 3:25

    designer there is talk to customers, try

  100. 3:27

    to understand how can we make agents as

  101. 3:29

    intuitive as possible, how do we figure

  102. 3:31

    out the mental models they're using to

  103. 3:33

    value the use cases they want to send

  104. 3:35

    out agents for.

  105. 3:36

    And a lot of them do seem pretty

  106. 3:39

    confused so far. A lot of people are

  107. 3:41

    excited about agents, but the phrase

  108. 3:43

    that comes up quite often is that they

  109. 3:45

    feel like they're just scratching the

  110. 3:46

    surface. Um

  111. 3:47

    it seems like it's not quite intuitive

  112. 3:48

    how we can best use them yet. And so, in

  113. 3:51

    a lot of my uh customer discussions,

  114. 3:53

    it's it always comes down to a question

  115. 3:55

    of like, what is the best way to use use

  116. 3:57

    uh to use agents? What are the best use

  117. 3:58

    cases for them? And how do I think about

  118. 4:00

    the trade-offs with regards to token

  119. 4:02

    costs relative to value? So, I think

  120. 4:04

    we're still kind of building this muscle

  121. 4:06

    today. And this leads me to the thesis

  122. 4:08

    of the talk, which is that I think

  123. 4:10

    agents have a measurement problem.

  124. 4:13

    And uh as an example, here's me at work

  125. 4:15

    trying to measure some agents, and one

  126. 4:17

    of my co-workers took this photo and

  127. 4:19

    told me I looked like

  128. 4:20

    I was uh trying to solve the mystery of

  129. 4:22

    Pepe Silvia. So,

  130. 4:23

    uh

  131. 4:24

    as you can see, it's it's

  132. 4:26

    it's not a it's not an easy task to to

  133. 4:28

    measure agents. But, I'm sure some of

  134. 4:30

    you are saying, "Hold on a sec. Like,

  135. 4:32

    what is this guy talking about? I've got

  136. 4:34

    a fleet of agents working for me right

  137. 4:35

    now. We're building our next

  138. 4:37

    million-dollar app as we speak, and I'm

  139. 4:39

    having a a totally fine time uh

  140. 4:41

    measuring my agents." Uh to which I will

  141. 4:43

    agree with you, uh but then I will point

  142. 4:45

    you to the mandatory Upton Sinclair

  143. 4:48

    quote to remind us all that everybody in

  144. 4:50

    this room is very biased, and we're

  145. 4:52

    early adopters, and we're very excited

  146. 4:55

    to explore this new technology, but it

  147. 4:57

    doesn't mean that we represent the

  148. 4:58

    people that ultimately we're going to be

  149. 5:00

    trying to help adopt this technology.

  150. 5:02

    And so, you know, I think it's important

  151. 5:04

    to remind ourselves that in some way or

  152. 5:06

    another, we probably are selling tokens,

  153. 5:08

    whether it's indirectly or directly. And

  154. 5:11

    so, when we think about our own token

  155. 5:12

    usage, uh is it really representative of

  156. 5:14

    all the people out there who have never

  157. 5:15

    touched an agent yet? Uh some people are

  158. 5:18

    still copy and pasting into ChatGPT.

  159. 5:20

    I may be married to one of these people,

  160. 5:22

    and despite how much I tried to get her

  161. 5:24

    to try out agents, she's not let me uh

  162. 5:26

    set her up with it yet. And so, um as a

  163. 5:29

    reminder, uh when we think about helping

  164. 5:31

    people adopt agents, you know, all the

  165. 5:32

    people across the world that we think

  166. 5:34

    could get as much uh excitement and

  167. 5:36

    value as as we do when we run off

  168. 5:37

    parallel agents, let's uh just remember

  169. 5:39

    this quote.

  170. 5:40

    >> [snorts]

  171. 5:41

    >> And um

  172. 5:42

    so, it really boils down to the age-old

  173. 5:44

    problem of a new technology.

  174. 5:47

    And of course, there's tons of history

  175. 5:49

    we can go to to study how people solved

  176. 5:51

    this in the past. We have this really

  177. 5:53

    exciting new thing, but we haven't quite

  178. 5:55

    uh figured out the right ways to

  179. 5:56

    communicate it.

  180. 5:57

    And so for this talk, I'll go back to

  181. 5:59

    the 1700s and we can take some notes

  182. 6:01

    from when James Watt was trying to sell

  183. 6:04

    steam engines.

  184. 6:06

    And at the time he decided that uh a

  185. 6:08

    great use case for his steam engines was

  186. 6:09

    trying to replace a horse gin. And these

  187. 6:13

    are the was the power source of a mill

  188. 6:15

    at the time. So when you're

  189. 6:16

    for uh let's say a brewery and you you

  190. 6:18

    need some power source to to grind your

  191. 6:21

    barley or whatever. I don't know. I'm

  192. 6:22

    not I'm not like a big brewery guy, so I

  193. 6:23

    don't know exactly what how it's made,

  194. 6:25

    but you need a power source and the

  195. 6:27

    power source at the time that was common

  196. 6:29

    was you hooked a horse up to a rotary

  197. 6:30

    arm and the horse walked in a circle and

  198. 6:32

    that's how you generated your power.

  199. 6:34

    Uh and seems crazy today maybe, but um

  200. 6:37

    at the time was commonplace and Watt

  201. 6:38

    thought, you know, it would be much

  202. 6:40

    better than a horse is like a very

  203. 6:41

    efficient machine.

  204. 6:43

    Although he um rightfully acknowledged

  205. 6:45

    that one of the big barriers to adopting

  206. 6:47

    it would be this cognitive dissonance of

  207. 6:50

    trying to tell people who kind of think

  208. 6:52

    in horses, how do you adapt to this to

  209. 6:54

    this uh cold machine that's kind of

  210. 6:56

    intimidating and scary and perhaps uh

  211. 7:00

    somebody's going to say it's going to

  212. 7:00

    solve all your problems, but you you

  213. 7:02

    can't quite see the vision yet. So uh

  214. 7:04

    perhaps that sounds familiar to any of

  215. 7:06

    us working in agents today.

  216. 7:08

    And Watt's solution was that he needed

  217. 7:10

    to understand um the mental model of

  218. 7:12

    these people and specifically to create

  219. 7:13

    a metric that would help him uh give

  220. 7:16

    some baseline of the relative

  221. 7:17

    improvement in efficiency.

  222. 7:19

    And so he literally studied uh horse

  223. 7:22

    gins and tried to get some kind of

  224. 7:25

    armchair measurements of of how is uh

  225. 7:28

    like where the mechanics and the average

  226. 7:29

    um performance of it and eventually came

  227. 7:32

    to a metric called horsepower,

  228. 7:35

    which may sound familiar.

  229. 7:37

    And uh he used this measure to, you

  230. 7:39

    know, this was to quantify the the

  231. 7:41

    general power that the horses were were

  232. 7:43

    um creating at the time and then he

  233. 7:45

    could use that as a basis to show the

  234. 7:46

    multiplier of efficiency that a steam

  235. 7:48

    engine could provide.

  236. 7:49

    And this metric was not very scientific

  237. 7:52

    at the time. It was not necessarily even

  238. 7:54

    accurate, you could say. But the main

  239. 7:57

    thing it did was it communicated an

  240. 7:59

    increase in value and so this led people

  241. 8:01

    who love horses

  242. 8:03

    let them kind of calibrate their

  243. 8:07

    the the efficiency gains that they could

  244. 8:08

    get by by attempting to adopt a steam

  245. 8:11

    engine. So not even necessarily um

  246. 8:14

    what you would get when you use it, but

  247. 8:16

    what would get you over the limit of

  248. 8:17

    trying it out in the first place.

  249. 8:20

    And you know, it's it's it's a pretty

  250. 8:21

    big feat because like although um

  251. 8:24

    he had efficiency on his side with

  252. 8:26

    regards to this metric

  253. 8:27

    you know, let's be honest regardless of

  254. 8:29

    how efficient this was

  255. 8:31

    horses just have

  256. 8:33

    great vibes. So like it's kind of hard

  257. 8:36

    to beat the vibes of horses and so he

  258. 8:37

    knew he had to kind of overcome the

  259. 8:39

    emotion

  260. 8:40

    and actually speak to to something that

  261. 8:42

    gave them an ability to calculate the

  262. 8:44

    ROI.

  263. 8:46

    And oh, sorry. Skipped something.

  264. 8:48

    And so yeah, the the lesson being if

  265. 8:50

    we're not able to give something that is

  266. 8:52

    a tangible ROI for our customers, then

  267. 8:55

    it's very hard for us to communicate

  268. 8:57

    value.

  269. 8:58

    And I think we only need look to our own

  270. 9:00

    industry to see all the examples where

  271. 9:03

    other people in the technology sector

  272. 9:05

    are failing to calculate good ROIs as

  273. 9:07

    well.

  274. 9:08

    And so we might in this room think this

  275. 9:09

    is some sort of solved problem.

  276. 9:11

    But if you look to the other engineers

  277. 9:13

    in the world who are perhaps not as AI

  278. 9:15

    pilled they're theoretically very smart

  279. 9:18

    and should be able to figure out how to

  280. 9:20

    calculate this pretty well, but then you

  281. 9:21

    get these scenarios where people are

  282. 9:23

    blowing through their entire

  283. 9:25

    token budget for a year and they're

  284. 9:27

    blowing through it in in a quarter or

  285. 9:29

    they're like dealing with token

  286. 9:30

    leaderboards and such and so obviously

  287. 9:33

    the incentives haven't quite aligned and

  288. 9:34

    we haven't perhaps got the right measure

  289. 9:36

    of value in terms of the technology

  290. 9:38

    sector itself. And so how then do we end

  291. 9:40

    up scaling past past that and talk to

  292. 9:42

    people who have no idea what we're

  293. 9:44

    talking about, but still try to provide

  294. 9:46

    them um a measure of the like increase

  295. 9:49

    efficiency with agents.

  296. 9:51

    And so right now I think we're kind of

  297. 9:52

    in this doom loop where we're

  298. 9:54

    we're overspending and we're underusing.

  299. 9:56

    Uh this is a term I borrowed from Ramp

  300. 9:58

    um and they have a great blog post on

  301. 9:59

    this. And so it's kind of this vicious

  302. 10:01

    cycle where we're just token maxing and

  303. 10:04

    ourselves into austerity

  304. 10:06

    kind of dropping out of the loop until

  305. 10:07

    we get more FOMO to to get activated

  306. 10:10

    enough to try it again.

  307. 10:12

    And so I think we have to break this

  308. 10:13

    loop and I think the way we do that is

  309. 10:14

    by getting better measures of that will

  310. 10:16

    communicate value.

  311. 10:19

    Some people are obviously on the right

  312. 10:20

    track. There was this chart floating

  313. 10:22

    around on X recently that the Coinbase

  314. 10:24

    Coinbase CEO posted where they hadn't

  315. 10:27

    really started changing the defaults of

  316. 10:29

    what models they will start with and

  317. 10:31

    trying to only save the frontier models

  318. 10:33

    for the hardest tasks and as a result

  319. 10:35

    saw some good

  320. 10:36

    saw AI spend start to diverge from token

  321. 10:39

    usage.

  322. 10:40

    And uh this is a good start. Uh Ramp

  323. 10:42

    also, as I mentioned, has a great blog

  324. 10:43

    post about this. Uh but I think the

  325. 10:44

    problem is still that it's too focused

  326. 10:47

    on tokens.

  327. 10:48

    And tokens are of course uh useful as a

  328. 10:52

    measurement of an internal system, but

  329. 10:54

    at at the end of the day they're just an

  330. 10:56

    output. And so the tokens need to then

  331. 10:59

    be traced very cleanly to an outcome.

  332. 11:01

    So how many

  333. 11:02

    uh bug how many bugs did the tokens uh

  334. 11:05

    how sorry how many um

  335. 11:07

    bugs squashed did the tokens that we

  336. 11:09

    bought um sorry, totally butchered that.

  337. 11:12

    Um how many uh bugs got squashed with

  338. 11:14

    the to with our token spend? How many uh

  339. 11:16

    support requests got closed, etc. So

  340. 11:18

    clean outcomes and then cleanly tying

  341. 11:20

    those to

  342. 11:22

    to progress on our objectives. And so

  343. 11:24

    without uh a very tight measure of ROI,

  344. 11:26

    this becomes very hard to do.

  345. 11:29

    And I think I'll take this further and

  346. 11:31

    um say that it need not even be the the

  347. 11:33

    broader um, technology industry where

  348. 11:35

    it's encountering this problem, but also

  349. 11:38

    many of us in this room perhaps are.

  350. 11:40

    And uh, although we're all probably

  351. 11:42

    enjoying uh, coding with uh, with

  352. 11:44

    various agents and feeling like it it's

  353. 11:46

    it does feel like there's something

  354. 11:47

    there in terms of the increase in

  355. 11:49

    ability and efficiency.

  356. 11:51

    Um, the problem of course is that we're

  357. 11:52

    all kind of dying by a thousand pull

  358. 11:54

    requests. And so, uh, even Anthropic who

  359. 11:57

    has uh, some people on the team have

  360. 11:59

    claimed have solved coding, uh, they

  361. 12:01

    have also admitted that they've not

  362. 12:02

    solved code review. And so, as a result,

  363. 12:05

    um, the the bottleneck is now shifted to

  364. 12:08

    the human review where the efficiency

  365. 12:10

    gains from coding agents aren't quite

  366. 12:11

    are aren't seen yet because we spend

  367. 12:13

    most of the time reviewing the code and

  368. 12:15

    we've not figured out how to scale that

  369. 12:16

    in tandem with uh, the generation of the

  370. 12:18

    code itself.

  371. 12:20

    And so, the bottleneck ends up shifting

  372. 12:22

    to the verification side and thus we

  373. 12:24

    don't have a way to uh, measure value at

  374. 12:27

    scale and uh, to judge quality at at the

  375. 12:29

    same speed. And so, again going back to

  376. 12:31

    the ROI calculations, we generate all

  377. 12:33

    this code, but how do we know uh, we

  378. 12:36

    don't know that enough of it is good to

  379. 12:37

    justify the spend.

  380. 12:40

    And of course, maybe uh, code review was

  381. 12:42

    always flawed, uh, but it's just that

  382. 12:44

    agents are now exposing it for

  383. 12:46

    uh, the are exposing the actual problem.

  384. 12:49

    Um, but I like this quote by Noah Hein

  385. 12:51

    who from a a post about how to solve

  386. 12:53

    code review where he's mentioning

  387. 12:54

    specifically that the assumptions

  388. 12:56

    underneath code review are what's now

  389. 12:57

    being uh, what needs to be revisited.

  390. 12:59

    So, you have to uh, check our priors to

  391. 13:02

    try to figure out a new basis for um,

  392. 13:04

    how we can code review in the age of

  393. 13:06

    agents.

  394. 13:07

    And I'm not going to go into how to

  395. 13:09

    solve code review. I think that's

  396. 13:09

    definitely better uh, a talk that's

  397. 13:11

    better given by somebody else and um, is

  398. 13:14

    totally different subject, but uh, what

  399. 13:16

    I think is important for this talk is

  400. 13:17

    why does code review feel like it is

  401. 13:19

    solvable? And I think that Noah is

  402. 13:20

    hitting on something important here

  403. 13:22

    which is that as a as a culture uh, code

  404. 13:24

    review has a a good uh,

  405. 13:26

    convergence on shared assumptions and

  406. 13:28

    that lets you

  407. 13:30

    um

  408. 13:31

    that lets you measure things at scale

  409. 13:33

    when we can all kind of converge on the

  410. 13:34

    measurement and it becomes somewhat of

  411. 13:38

    clear rubric and so the task at hand now

  412. 13:41

    is we have to adopt we have to adapt

  413. 13:44

    those assumptions for the agentic age.

  414. 13:49

    And so

  415. 13:50

    we can we need to go if we're able to do

  416. 13:52

    that then we can go from execution at

  417. 13:54

    the speed of computer to measurement at

  418. 13:56

    the speed of computer and of course the

  419. 13:58

    measurements need to fit the mental

  420. 13:59

    models of the customers using it.

  421. 14:01

    And

  422. 14:03

    I think the lesson here being that if

  423. 14:05

    you're going to

  424. 14:06

    think of how to build an agent for

  425. 14:08

    something you also have to think about

  426. 14:10

    how do you help the customers build

  427. 14:12

    build or at least create a method for

  428. 14:14

    verifying that the output is good and so

  429. 14:17

    it's not enough to build it we also have

  430. 14:18

    to help them

  431. 14:20

    we also we also also have to help them

  432. 14:21

    get to clear ROI calculations to justify

  433. 14:24

    their spend.

  434. 14:26

    And so this brings me to the idea of

  435. 14:29

    mouse power which could be the

  436. 14:31

    equivalent of horsepower for the agentic

  437. 14:33

    age just as James Watt was able to show

  438. 14:35

    a measure of efficiency relative to the

  439. 14:37

    horses in in the gins in the horse gins

  440. 14:40

    that were the source of power at the

  441. 14:41

    time we perhaps can also figure out how

  442. 14:44

    do we create a baseline of efficiency

  443. 14:46

    for the way we use computers today and

  444. 14:49

    can then demonstrate how much better or

  445. 14:51

    perhaps more performant on certain

  446. 14:53

    vectors an agent could be at that task.

  447. 14:56

    Um and

  448. 14:57

    of course it's not as easy perhaps as

  449. 14:59

    easy a task as he had back then where he

  450. 15:01

    could just study the horse gin cuz it's

  451. 15:03

    not as if we can create some method to

  452. 15:05

    measure cursor movements and like figure

  453. 15:07

    out the delta of how much more efficient

  454. 15:09

    an agent could move them and thus we can

  455. 15:11

    say yeah agents are this much more

  456. 15:13

    performant than humans at these tasks.

  457. 15:15

    Trust me I've I've tried I had Claude

  458. 15:18

    vibe code me this measurement device and

  459. 15:20

    I thought maybe if I can figure out the

  460. 15:22

    movement like the potential movement

  461. 15:24

    across the screen and measure how fast

  462. 15:26

    it went, I could get some clean measure

  463. 15:28

    of mouse power. Uh but of course this is

  464. 15:30

    only joking. Um this is of course um

  465. 15:33

    like a fool's errand because information

  466. 15:35

    space is just way too high dimensional

  467. 15:37

    and so I think mouse power is is never

  468. 15:40

    going to be a metric of course, but it's

  469. 15:41

    more so an idea. Which the idea being if

  470. 15:44

    you're going to sell somebody an agent,

  471. 15:45

    you also have to help them with the with

  472. 15:47

    the rubric of how do we actually verify

  473. 15:50

    that this agent is doing good work and

  474. 15:52

    thus we can uh have a good measure of

  475. 15:54

    saying that these tokens are worth it.

  476. 15:56

    Um so how to do that of course is is

  477. 15:58

    really up to you and I won't be able to

  478. 16:00

    tell you how do you I don't have any

  479. 16:02

    good frameworks for how do you figure

  480. 16:04

    out the right measurements to to help

  481. 16:05

    provide anybody you're building an agent

  482. 16:07

    for. Uh but what I can do is give a

  483. 16:09

    principle uh give an idea that I've been

  484. 16:11

    kicking around which is based um

  485. 16:13

    in information theory. So going back to

  486. 16:16

    Claude Shannon's ideas about measuring

  487. 16:18

    entropy and information.

  488. 16:20

    Uh entropy being uh the uncertainty of a

  489. 16:22

    probability distribution and of course

  490. 16:25

    very much the basis of how we train

  491. 16:27

    agents today.

  492. 16:28

    Things like cross entropy and such uh

  493. 16:30

    being a big factor in determining how

  494. 16:31

    capable an agent is.

  495. 16:33

    Um I think that entropy's an interesting

  496. 16:35

    idea to think through with regards to

  497. 16:36

    not just the performance of an agent,

  498. 16:39

    but also the task that we're setting

  499. 16:40

    them out to to perform on.

  500. 16:42

    And so uh I put together this matrix

  501. 16:45

    which uh it maps on the x-axis axis the

  502. 16:49

    uncertainty in the steps it takes to

  503. 16:50

    perform a task. And so when we're

  504. 16:52

    thinking of building an agent, I think

  505. 16:54

    it's not enough to just think what would

  506. 16:56

    be a valuable task for the agent to do,

  507. 16:58

    but also thinking about how um how much

  508. 17:01

    uncertainty are in the steps to perform

  509. 17:02

    that task itself. So an example would be

  510. 17:05

    uh booking a flight has uh much less

  511. 17:08

    uncertainty than let's say painting a

  512. 17:09

    masterpiece, right? Because you know

  513. 17:11

    there's certain information that has to

  514. 17:13

    be that has to happen in the flight

  515. 17:14

    purchase. There has to be a departing

  516. 17:16

    destination, arriving destination.

  517. 17:18

    There's going to be a seat chosen. It

  518. 17:19

    might be by the person. It might just be

  519. 17:21

    random.

  520. 17:22

    But these things have to happen for that

  521. 17:24

    task to be completed. And on the other

  522. 17:25

    hand, there is the task of like painting

  523. 17:28

    a masterpiece, right? And who knows what

  524. 17:30

    the steps are to that? Maybe you can get

  525. 17:32

    an agent to do it, but it would be very

  526. 17:34

    hard to figure out how we can actually

  527. 17:36

    create a a relatively predictable

  528. 17:39

    pathway to that.

  529. 17:40

    But then on the other axis is the the

  530. 17:43

    uncertainty in the acceptance criteria

  531. 17:45

    itself. So not just can the agent

  532. 17:47

    perform the task, but can we help

  533. 17:49

    somebody actually or is is there

  534. 17:51

    actually a a clean rubric for how it's

  535. 17:53

    graded? And so thinking about ideas on

  536. 17:56

    on these two axes and where they

  537. 17:57

    intersect, perhaps gives us a better

  538. 17:59

    guide for how to build agents and we can

  539. 18:01

    run through a few examples. So if we

  540. 18:05

    look at the at the left side, your right

  541. 18:07

    side.

  542. 18:09

    Yes.

  543. 18:10

    No, you're left as well.

  544. 18:12

    Then

  545. 18:13

    No, last speaker was also a confused

  546. 18:15

    about.

  547. 18:16

    So uh

  548. 18:17

    Yeah, on the left side when uncertainty

  549. 18:20

    in the task steps are low, then it's a

  550. 18:22

    very it's a very predictable outcome or

  551. 18:25

    it's a very predictable pathway to

  552. 18:26

    achieve that goal. And so then, you

  553. 18:28

    know, why would you waste tokens? Just

  554. 18:30

    write a script. On the other side, when

  555. 18:33

    the

  556. 18:34

    the steps to do perform the task are

  557. 18:36

    very high in in uncertainty, then you

  558. 18:39

    you have very unpredictable information.

  559. 18:40

    And so it's probably at risk of being

  560. 18:43

    out of distribution in pre-training and

  561. 18:45

    probably has very sparse rewards for

  562. 18:46

    reinforcement learning. And so perhaps

  563. 18:48

    it's not a a good task for an agent

  564. 18:50

    because it's just much harder to figure

  565. 18:52

    out how to actually model that data.

  566. 18:54

    And so obviously in the middle is is um

  567. 18:58

    is so I'm I think I'm out of time, but

  568. 19:00

    I'm not getting kicked off yet.

  569. 19:02

    So I'll just finish this up quickly. Um

  570. 19:04

    so yeah, in the middle is is probably

  571. 19:05

    the sweet spot, but then on the other

  572. 19:06

    axis,

  573. 19:07

    what's the uncertainty in verifying that

  574. 19:09

    this is actually valuable? So when you

  575. 19:11

    have high uncertainty in the acceptance

  576. 19:13

    criteria, you pretty much are in a spot

  577. 19:14

    where verification is indistinguishable

  578. 19:16

    from execution. So, why would you build

  579. 19:18

    an agent for something that to verify

  580. 19:21

    was useful, a person pretty much has to

  581. 19:22

    do the work again. So, like waste of

  582. 19:25

    tokens obviously.

  583. 19:27

    And then it leaves that that middle area

  584. 19:29

    where you have this interesting

  585. 19:30

    intersection of tasks that are um

  586. 19:33

    they're not too uncertain in that

  587. 19:36

    they or they they have a degree of

  588. 19:38

    uncertainty where they're not great

  589. 19:41

    they're not just a a script or they're

  590. 19:43

    not out of distribution for training,

  591. 19:45

    but they have enough uncertainty to be

  592. 19:46

    interesting, but at the at the same time

  593. 19:49

    they also have a property of being

  594. 19:50

    relatively easy to

  595. 19:52

    validate the

  596. 19:54

    to check the value of them. And so they

  597. 19:56

    become in this place where they kind of

  598. 19:58

    become the shape of an NP-style problem,

  599. 20:00

    which means they're easier to verify

  600. 20:01

    than to execute. And the reason I say

  601. 20:03

    that is because if you can figure out a

  602. 20:05

    pretty repeatable pattern for verifying

  603. 20:07

    their their work, you can actually just

  604. 20:09

    throw agents to that problem as well.

  605. 20:11

    And so of course you don't just build

  606. 20:13

    the agent, you perhaps build the agent

  607. 20:15

    that verifies the work of the agent.

  608. 20:17

    Um and so yeah, this is perhaps this is

  609. 20:21

    a thought starter mostly kind of have

  610. 20:23

    kind of still in the works, so

  611. 20:25

    I'm happy to hear any thoughts on it,

  612. 20:26

    but if with this guidance I hope when

  613. 20:29

    you're building your next agent you can

  614. 20:30

    also figure out how to also build its

  615. 20:32

    mouse power.

  616. 20:34

    And thanks very much.