AI Engineer World's Fair 2026

An AI Research Agent That Runs Your Experiments — Tim Sweeney, Weights & Biases

Read the talk

An AI Research Agent That Runs Your Experiments

ARIA turns a research request into GPU jobs, experiment analysis, and visual reports inside Weights & Biases. A live batch shows how that loop works; the team's own traces and nightly evaluations show how they decide whether the agent itself is improving.

From a talk by Tim Sweeney

At a glance

Ideas worth remembering

  • Keep long-running training outside the agent's main loop: ARIA starts experiments through Launch and polls while GPU jobs execute.

  • Evaluate useful research behavior along separate dimensions: the example task checks correctness, interesting insights, and a six-tool-call limit.

  • Production traces become more useful when human review and live judges turn observed behavior into repeatable tasks and candidate evaluations.

  • Domain context and available tools are practical places to improve an agent before adding elaborate harness or memory machinery.

  • The live batch completed 12 experiments and nearly matched the earlier best result; autonomous execution did not guarantee a better model.

Start with a project that already has 200 experiments

A research workspace can accumulate experiments faster than a person can explain them. Tim Sweeney, Principal Engineer at Weights & Biases by CoreWeave, opens a project with more than 200 training jobs. A scatter plot shows loss declining over time. The next problem is deciding what those runs teach—and what experiment to run next.

Source frame: Start with a project that already has 200 experiments
Source frame: Start with a project that already has 200 experiments

The project uses Karpathy's autoresearch, a small codebase that trains an LLM and provides a manageable setting for repeated changes. ARIA, the AI research and iteration agent, opens as a chat inside the workspace. Project context and images can accompany a request, so the conversation can begin with the experiment data already close at hand.

The useful starting point is a long-running conversation rather than the agent's onstage introduction. In that conversation, ARIA has helped download the code, configure Launch jobs, and set up GPUs. It can then iterate on both the code and its hyperparameters. Those are different kinds of changes: one alters the implementation being trained; the other changes how that implementation trains.

0:150:17
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:13 · section reference included

A chat request becomes a batch of GPU jobs

Sweeney asks ARIA to conduct another batch of experiments and try to find the best model live. He adds encouragement—because, naturally, the models need encouragement. The consequential decision comes next: ARIA avoids a large architecture change, treating it as too risky for this iteration. Sweeney expects it to try hyperparameter modifications instead, and a shell call starts the experimentation loop.

Source frame: A chat request becomes a batch of GPU jobs
Source frame: A chat request becomes a batch of GPU jobs

W&B Launch supplies the connection to compute. A Launch queue lets humans or agents submit long-running experiment jobs to a cluster, including jobs that need GPUs. Sweeney opens terminal output from his Kubernetes cluster to show experiments executing, then returns to ARIA, which is polling while it waits for the work to finish. The chat initiates and monitors the work; the cluster performs the training.

Where does the time-consuming work happen after a research request? The flow below separates the agent's decision and shell execution from queued GPU training. The polling path makes the separation visible: launching a job does not mean its result is immediately available.

How it fits togetherFrom request to running experiments

Conduct another batch and look for a better model.

ARIA starts the experiment loop, Launch connects it to GPU compute, and the agent polls while training proceeds.

4:334:35
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

4:30 · section reference included

Turn accumulated runs into explanations and charts

While training runs, a second request asks for a summary of the project's highest-performing runs. This is useful when a teammate joins a project or returns from time away: the experiment history needs to become an explanation of what worked. A previously prepared pattern-finding example goes further, identifying an emerging model family, batch size as an influential parameter, and a promising architectural recipe. These are leads for further research, rather than a demonstrated guarantee that any one change causes an improvement.

Source frame: Turn accumulated runs into explanations and charts
Source frame: Turn accumulated runs into explanations and charts

The analysis can become a native W&B artifact rather than remaining a chat answer:

  • Reports: A report combines Markdown with embedded plots, charts, and graphics—Sweeney calls it a “markdown file on steroids.” The example includes a project thesis, data panels, and a parameter-importance chart showing relationships among parameters.
  • Workspaces: ARIA is tuned and prompted to construct workspaces and plots using W&B's built-in chart types. That gives the analysis a form colleagues can inspect in the interface they already use.

A check-in shows the summary request querying W&B, applying patches, and writing code, while the training conversation continues polling for results. Analysis and training therefore have their own work to complete. ARIA's value here comes from connecting research data, executable tools, and visualization primitives within the same product.

The iOS announcement extends that workflow to a phone: researchers can steer hyperparameter-tuning jobs while away from their desks. The larger ambition is an end-to-end automated research platform that handles job orchestration, GPU workloads, research lookup, and hypothesis collaboration. Sweeney frames the division of labor around letting ARIA handle mechanics while researchers concentrate on new ideas, architectures, and parameters.

6:026:03
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

5:55 · section reference included

Separate conversation handling from long-running compute

ARIA's backend begins with a familiar sequence: web and iOS clients communicate with an API server, the server writes to a turn database, and a worker harness processes the work. The harness then connects to the capabilities needed to carry out a research request.

Source frame: Separate conversation handling from long-running compute
Source frame: Separate conversation handling from long-running compute

Those connections serve distinct purposes:

  • Sandbox execution: Shell calls and Python data-science work give the agent a way to manipulate code and analyze data.
  • Model inference: An LLM provider supplies the model calls used by the worker; W&B Inference is presented as one option.
  • External training: Launch handles workloads outside the agent's main loop, including training that can take days, with CoreWeave GPUs supplying compute.
  • Observability: Sessions, turns, tool calls, and errors are logged so the team can inspect what the agent actually did. Sweeney says ARIA logs 100% of its traces to W&B Weave.

The trace store also feeds development. Offline work turns observed behavior into tasks, evaluates candidate agents, and records results in a shared dashboard. One loop improves the tests; another forms hypotheses, implements candidates, and studies their evaluation results. The loops are complementary but also adversarial: better tests can expose weaknesses in a candidate that previously looked promising. The selected agent returns to production through a registry.

10:3010:32
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

10:12 · section reference included

Read the shape of a conversation, then inspect its failures

The Weave dashboard gives a bird's-eye view through span volume, conversation volume, and token tracking. Sweeney then opens a conversation feed filtered to internal employees. Its span view makes trace topology visible: different colors and shapes distinguish tool calls, LLM calls, and thinking blocks. Opening a conversation reveals the system prompt, user messages, shell calls, and reasoning blocks behind that shape.

Source frame: Read the shape of a conversation, then inspect its failures
Source frame: Read the shape of a conversation, then inspect its failures

That detailed view is where the research lead, engineer, and product manager add notes and feedback before turning behavioral discoveries into tasks. A contextual Summarize button can also start an ARIA conversation about the item being inspected. In this example, ARIA analyzes one of its own conversations to recommend improvements to ARIA. The agent assists the review, while the team still has the underlying trace available.

Signals add automated triage to live traffic. LLM judges flag user frustration, low-quality responses, and other behaviors, helping the team find clusters to address in a later iteration. One displayed judge identifies frustration because the user expresses dissatisfaction with a loss curve. Its reasoning is inspectable, which gives reviewers a concrete explanation to assess rather than just a flag. The flag identifies dissatisfaction; it does not by itself establish what caused the poor experience.

13:1313:15
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

12:59 · section reference included

Turn observed behavior into nightly tests

A task is a YAML file that acts as a unit test for the agent. The demonstrated prompt asks ARIA to compare two runs that both give good results: what differs, and what can be learned? The task pairs that request with definitions of success. This makes a broad product expectation—give useful research advice—specific enough to evaluate repeatedly.

Source frame: Turn observed behavior into nightly tests
Source frame: Turn observed behavior into nightly tests

The example checks three separate properties:

  • Correctness: An LLM judge evaluates the answer against correctness defined for this particular question.
  • Interesting insights: A second LLM judge asks whether the comparison produces insights worth knowing. A correct answer can still be unhelpful.
  • Expediency: A rule-based judge checks whether the agent produces a result within six tool calls, giving efficiency an explicit test alongside answer quality.

About 200 tasks form a suite that runs nightly, with results tracked in Weave. The displayed candidate scores 73%, compared with the production model's 72%; Sweeney says the team plans to move it forward that Friday. These scores describe performance on the team's suite. The recording does not establish whether the one-percentage-point difference is statistically significant or predicts improvement outside that suite.

How does a production conversation influence the next release? The cycle below shows the intermediate steps: review produces tasks, tasks test candidates, and the results inform a team decision. Production traces provide the starting material; they become useful release evidence only after the team defines the behavior it wants to test.

How it fits togetherThe agent improvement loop

Collect conversations and agent actions.

Live behavior supplies examples; explicit tasks and candidate evaluations turn those examples into release decisions.

16:1016:12
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

16:02 · section reference included

Keep humans in the loop—and let the experiment finish

The production advice follows directly from this development loop:

  • Instrument behavior: Log sessions, turns, tools, and feedback to catch behavioral bugs. An agent can complete its calls without producing the behavior a user needs.
  • Use evaluations for decisions: Researchers and engineers should develop tasks together and treat performance metrics as go/no-go evidence.
  • Keep human judgment: Use the product and review its best and worst traces as a team. LLM judges miss behavioral nuances, so automated scores cannot cover the whole review.
  • Improve context and tools: Before adding elaborate harness or memory machinery, give the agent the business domain, available primitives, and relevant business data. Sweeney's team found substantial opportunities in those simpler additions.
Source frame: Keep humans in the loop—and let the experiment finish
Source frame: Keep humans in the loop—and let the experiment finish

The final check returns to the batch started onstage. A new point has appeared in the workspace, and the batch has run 12 experiments. The reported result is 5.833, close to the earlier best but without beating it. ARIA has carried a request through experiment selection, shell execution, queued GPU training, waiting, and a visible result. The automation completed a research iteration; an improvement remained something the experiments had to earn.

18:5119:01
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

18:43 · section reference included

Resources

From the talk

  • The tracing and evaluation product used in the talk to inspect conversations, run live judge signals, and compare nightly candidate results.

Read the complete timestamped transcript
  1. 0:01

    [music]

  2. 0:13

    Okay. Hello everyone and thank you for

  3. 0:15

    attending this session. My name is Tim

  4. 0:17

    Sweeney, a principal engineer at Weights

  5. 0:20

    and Biases and Coreweave. And for the

  6. 0:22

    next 20 minutes, we're going to talk

  7. 0:23

    about Arya, our new AI research and

  8. 0:26

    iteration agent. Let's go ahead and get

  9. 0:28

    started.

  10. 0:29

    So, uh, first off, just by way of making

  11. 0:32

    some noise, some clapping, uh, who here,

  12. 0:35

    um, identifies as an ML researcher?

  13. 0:37

    You're someone that trains models,

  14. 0:38

    trains the brain?

  15. 0:41

    I heard one. Wow. Okay. Great work.

  16. 0:43

    Great work. Uh, what about who here is

  17. 0:45

    the applied engineer, the namesake of

  18. 0:47

    this conference? Who here actually

  19. 0:48

    builds the bots?

  20. 0:50

    >> Okay, good. Expected much more. And who

  21. 0:53

    here is in AI management? You are

  22. 0:55

    helping fund this compute.

  23. 0:58

    Okay. Okay. Nice. From the back. Lovely.

  24. 1:01

    Um, well, now that I know a little bit

  25. 1:02

    about you, just a little bit about me.

  26. 1:04

    Uh, again, my name is Tim. I have a

  27. 1:06

    masters in machine learning, uh, and

  28. 1:07

    reinforcement learning from Georgia

  29. 1:09

    Tech. So, I've been that, uh, researcher

  30. 1:11

    currently building Weights and Biases

  31. 1:13

    agent, Arya. So, identify as that

  32. 1:15

    applied engineer. And in a previous life

  33. 1:18

    was the PM of Twitter's ML stack. So, I

  34. 1:21

    hope you hopefully can connect with you

  35. 1:22

    middle management as well. [laughter]

  36. 1:24

    >> [gasps]

  37. 1:24

    >> Um today's agenda is kind of broken into

  38. 1:27

    three sections and hopefully each of you

  39. 1:29

    personas walk away with something

  40. 1:30

    valuable. So first we're going to learn

  41. 1:32

    about Arya itself and how it can

  42. 1:34

    supercharge your AI and ML workflows.

  43. 1:37

    We're going to dive into auto research

  44. 1:38

    and see that live in a live demo in just

  45. 1:40

    a moment.

  46. 1:42

    Then we're going to pull back the

  47. 1:43

    curtain and learn how we use weights and

  48. 1:44

    biases and uh coreweave to actually

  49. 1:47

    build Arya because a lot of you in the

  50. 1:49

    audience are building agents yourself

  51. 1:50

    and we believe a lot of these components

  52. 1:52

    can help you in your endeavors. And then

  53. 1:54

    towards the end we'll just take a step

  54. 1:56

    back and identify a few key tips and

  55. 1:58

    tricks for making sure that you're able

  56. 1:59

    to productionize your systems

  57. 2:00

    effectively.

  58. 2:02

    For those of you who might not be

  59. 2:03

    familiar, Weights and Biases is the

  60. 2:05

    world's leading AI development platform.

  61. 2:07

    We've been in business now for nine

  62. 2:09

    years and have happily joined the core

  63. 2:11

    family about a year ago. Uh we have a

  64. 2:13

    number of products in our suite but are

  65. 2:15

    really known for our models training

  66. 2:17

    inference and weave stack which really

  67. 2:19

    helps collect data uh about the AI

  68. 2:21

    development and machine learning

  69. 2:22

    workflows and makes that information

  70. 2:24

    actionable and uh enables users to make

  71. 2:27

    the best decisions about what to do

  72. 2:28

    next.

  73. 2:30

    So without further ado, let's go ahead

  74. 2:32

    and dive into Arya, our agent. Uh we'll

  75. 2:34

    show a demo and then we'll get back to

  76. 2:35

    some slides.

  77. 2:40

    Okay, beautiful.

  78. 2:42

    Let's make this a bit bigger. Holler at

  79. 2:44

    me if you need it to be bigger. So, uh,

  80. 2:46

    what you're looking at here is a weights

  81. 2:48

    and biases workspace. For you, for

  82. 2:50

    anybody that isn't familiar, on the

  83. 2:51

    lefth hand side, I actually see a list

  84. 2:53

    of a bunch of different experiments. In

  85. 2:55

    this particular project, I have over 200

  86. 2:57

    training jobs. And on the right hand

  87. 3:00

    side, I see a scatter plot of, in this

  88. 3:02

    case, declining metrics, which is good.

  89. 3:04

    means our loss is going down over time.

  90. 3:06

    And this view would be very familiar for

  91. 3:08

    anyone that uses our tool. Now, to

  92. 3:10

    ground this, we're actually uh uh using

  93. 3:13

    the Carpathy Auto Research project,

  94. 3:15

    which I'm sure many of you are familiar

  95. 3:16

    with, but if you're not, it's just a

  96. 3:18

    very simple project that trains an LLM,

  97. 3:20

    and it's a great foundation for auto

  98. 3:23

    research type demonstrations because

  99. 3:25

    it's a very simple codebase and allows

  100. 3:27

    us to improve iteratively over time. So,

  101. 3:30

    let's jump back to the project and open

  102. 3:31

    up Arya by clicking this blue button in

  103. 3:33

    the upper right. When I click this

  104. 3:35

    button, I'm uh presented with the

  105. 3:37

    familiar chat interface with, you know,

  106. 3:39

    how can I help you today, a few call to

  107. 3:40

    actions, and you know, I can add

  108. 3:43

    different context in my project or maybe

  109. 3:45

    add images, etc. Um, everyone here is

  110. 3:48

    agent builders, so I don't need to bore

  111. 3:49

    you with the details of what an agent

  112. 3:51

    interface looks like. But let's go ahead

  113. 3:53

    and just, you know, enter in a basic

  114. 3:55

    intro here. Let's say, "Hello, Arya.

  115. 3:58

    You're on stage at AI World's Fair 2026.

  116. 4:00

    Please introduce yourself." So, it's

  117. 4:02

    going to go ahead and chug along and

  118. 4:03

    hopefully emit some sort of nice emoji.

  119. 4:06

    Yay. He I'm Arya. I'm talking to the

  120. 4:08

    audience. Great. But now, let's dive

  121. 4:10

    into the meat of why you came here. So,

  122. 4:13

    I'm going to open up this chat here. And

  123. 4:15

    this is a longunning chat where I've

  124. 4:17

    been running again over 200 experiments

  125. 4:19

    using the auto research loop. Um, it

  126. 4:22

    helped me download the code, set up my

  127. 4:24

    launch jobs, set up my GPUs, and is able

  128. 4:26

    to autonomously iterate on the code

  129. 4:28

    itself and the hyperparameters.

  130. 4:30

    We'll take a look at what it's doing in

  131. 4:31

    a moment, but while we're doing this,

  132. 4:33

    I'm going to kick off a live iteration

  133. 4:35

    right here. So, what I'm going to say is

  134. 4:37

    please conduct another batch of

  135. 4:39

    experiments. You are on stage at the AI

  136. 4:41

    Engineer Worlds Fair 2026 and we're

  137. 4:44

    hoping to find the best model live. I

  138. 4:46

    believe in you. Uh, because we know we

  139. 4:48

    have to encourage our models. Um, so

  140. 4:50

    it's been doing this for a while. What

  141. 4:51

    it what it's doing here is it's saying,

  142. 4:53

    "Okay, great. Um, I don't want to make a

  143. 4:55

    big architecture swing. That feels a

  144. 4:57

    little bit too risky." So, it's probably

  145. 4:58

    going to go for uh some modifications to

  146. 5:01

    the hyperparameters. And then it's

  147. 5:03

    kicking off a shell call here that is

  148. 5:05

    actually um executing that uh executing

  149. 5:08

    that experimentation loop. And we're

  150. 5:10

    going to check in on this periodically

  151. 5:11

    throughout this presentation.

  152. 5:13

    But I want to help explain what's going

  153. 5:15

    on behind the scenes. So behind the

  154. 5:17

    scenes I have set up a weights and

  155. 5:19

    biases launch queue. Launch is our our

  156. 5:21

    product that allows you to connect your

  157. 5:23

    compute clusters and allows humans and

  158. 5:25

    agents to launch longunning

  159. 5:27

    experimentation jobs particularly by

  160. 5:29

    leveraging GPUs.

  161. 5:31

    Here I'm looking at a uh a terminal

  162. 5:33

    output of my Kubernetes cluster where

  163. 5:35

    we're actually seeing live execution of

  164. 5:37

    experiments happening. So this is

  165. 5:39

    happening live right here. This is not a

  166. 5:40

    fake demo. Um great. And if we jump

  167. 5:44

    back, we see that at this point it

  168. 5:46

    started the cues and now it is simply

  169. 5:48

    polling and waiting for our work to be

  170. 5:50

    complete. So we'll jump back to that in

  171. 5:52

    a in a moment. But before but let's dive

  172. 5:55

    into a few other examples. So uh

  173. 5:57

    something else that is interesting you

  174. 5:59

    can do is maybe you might want to ask it

  175. 6:02

    something like please summarize the

  176. 6:03

    highest performing runs in this project.

  177. 6:05

    This use case would be something like

  178. 6:06

    maybe a new user come or a new uh team

  179. 6:08

    member is joining your project and want

  180. 6:10

    to understand the research. Um or maybe

  181. 6:12

    you've uh someone's been doing some work

  182. 6:14

    while you were on PTO and you want to

  183. 6:16

    get caught up. We'll see what this comes

  184. 6:17

    up with in a moment. Some other

  185. 6:19

    pre-anned uh examples are finding

  186. 6:21

    patterns in your project. So here we can

  187. 6:24

    see that I asked it, hey, can you find

  188. 6:26

    some patterns in this research? And we

  189. 6:28

    see that um it identified that a new

  190. 6:30

    family of models emerged as the as the

  191. 6:33

    auto uh auto research was happening. Uh

  192. 6:35

    it identified that batch size seems to

  193. 6:37

    be a really high high uh lever uh

  194. 6:40

    parameter. It identified an

  195. 6:42

    architectural recipe that seemed to be

  196. 6:44

    quite promising and a number of other

  197. 6:46

    insights that would have taken me hours

  198. 6:48

    or days to discover on my own. And Arya

  199. 6:50

    is able to do it right for me directly

  200. 6:52

    in the interface that I already live.

  201. 6:55

    Not only is it able to emit text based

  202. 6:57

    uh textbased outputs, but it also deeply

  203. 7:00

    integrates with a number of weights and

  204. 7:01

    biases visualization utilities. So here

  205. 7:04

    I've actually asked it to emit a weights

  206. 7:06

    and biases report which for those who

  207. 7:08

    aren't familiar is essentially a

  208. 7:09

    markdown file on steroids. It's got uh

  209. 7:12

    embedded embedded plots, charts and and

  210. 7:14

    and graphics. And so here uh you know

  211. 7:16

    it's talked about the thesis of the

  212. 7:18

    project. It's it's emitted a number of

  213. 7:20

    of data panels. And uh I actually think

  214. 7:23

    it's quite interesting. It used um one

  215. 7:25

    of our more esoteric panels, the uh

  216. 7:27

    parameter importance chart to uh tell me

  217. 7:30

    the correlation of various different

  218. 7:32

    parameters within this uh within this

  219. 7:33

    training job.

  220. 7:35

    Uh in addition to uh reports, it's also

  221. 7:38

    great at working with workspaces. So if

  222. 7:40

    you're a weights and biases user, uh you

  223. 7:43

    spend a lot of your time uh designing

  224. 7:45

    and working with workspaces. Well, Arya

  225. 7:48

    is actually customtuned and prompted to

  226. 7:50

    really understand how to build

  227. 7:51

    workspaces, build plots, and complement

  228. 7:54

    that that data analytics with real live

  229. 7:56

    graphics using the built-in proprietary

  230. 7:58

    charts that weights and biases users

  231. 8:00

    know and love. Um, so with that, let's

  232. 8:03

    go ahead and check back on some of our

  233. 8:04

    our prompts. We can see that the please

  234. 8:06

    summarize this project prompt is cooking

  235. 8:08

    away. It's querying weights and biases.

  236. 8:10

    It's applying patches. It's writing its

  237. 8:12

    own code. So, we'll come back and check

  238. 8:13

    on that in a moment. and our longunning

  239. 8:15

    training job is uh still pulling for the

  240. 8:18

    results. We can see that we're cooking

  241. 8:20

    away on our GPUs. So, we're we're frying

  242. 8:22

    some GPUs [music] and doing some data

  243. 8:24

    science all live. And while that's

  244. 8:26

    cooking, let's go ahead and jump back to

  245. 8:27

    the presentation. We'll come back in a

  246. 8:28

    moment.

  247. 8:30

    Uh

  248. 8:33

    oh, no, we're not looking at a

  249. 8:34

    dictionary. We're looking at a PO.

  250. 8:36

    Great. Uh okay, so quick recap here.

  251. 8:38

    What did Arya show? What did we show in

  252. 8:40

    these last five minutes? First, we show

  253. 8:42

    that uh Arya can serve as your data

  254. 8:44

    science companion right inside of

  255. 8:46

    Weights and Biases, helping you discover

  256. 8:48

    insights that you wouldn't you wouldn't

  257. 8:49

    be able to discover as your experiments

  258. 8:51

    and as your team size grows.

  259. 8:55

    Next, we address the problem of

  260. 8:57

    complicated reporting and complicated

  261. 8:58

    plotting. Weights and biases users are

  262. 9:00

    are really want to turn their insights

  263. 9:02

    into visual communication tools. They

  264. 9:04

    want to communicate with their peers and

  265. 9:06

    their colleagues. So Arya's built from

  266. 9:08

    the ground up to understand those

  267. 9:09

    primitives and help co-pilot and drive

  268. 9:12

    right along right alongside in the UI

  269. 9:15

    and announcing now today for the first

  270. 9:17

    time we are releasing Arya on our iOS

  271. 9:19

    device or on our iOS app. So uh uh Arya

  272. 9:23

    released on Monday and our iOS app now

  273. 9:26

    has Arya built in. So if you're

  274. 9:27

    conducting hyperparameter tuning jobs,

  275. 9:29

    if you're training models, or if you're

  276. 9:31

    just researching within the weights and

  277. 9:32

    biases ecosystem, you can go touched

  278. 9:34

    grass at Yerba Buena uh gardens and

  279. 9:37

    steer your uh hyperparameter tuning jobs

  280. 9:39

    all from your mobile device. And what is

  281. 9:41

    this all building up to? This is

  282. 9:43

    building up to an a fully automated

  283. 9:45

    endto-end research platform where we're

  284. 9:47

    not seeking to replace uh RL

  285. 9:49

    researchers, but complement your

  286. 9:50

    workflows. Arya's great at orchestrating

  287. 9:53

    jobs, understanding GPU workloads,

  288. 9:55

    responding to events within the within

  289. 9:57

    the Wandi ecosystem, and listening to

  290. 9:59

    researchers, uh, uh, looking up archive

  291. 10:01

    papers, and collaborating on hypothesis.

  292. 10:04

    So, we can let Arya drive the mechanics

  293. 10:05

    that you don't want to deal with while

  294. 10:07

    you focus on the new ideas, new

  295. 10:08

    architectures, and new parameters that

  296. 10:10

    you wanted to try.

  297. 10:12

    Um, great. So, that's Arya in a

  298. 10:14

    nutshell. We're really hoping that you

  299. 10:16

    give it a shot. And uh we'll jump back

  300. 10:18

    to the auto research at the end and see

  301. 10:20

    if we got a new best record. But before

  302. 10:22

    we do that, let's talk about how we use

  303. 10:24

    weights and biases and coreweave to

  304. 10:26

    actually build Arya. So now speaking to

  305. 10:28

    a lot of the the AI agent builders in

  306. 10:30

    the room, here's a quick architecture on

  307. 10:32

    the lefth hand side. You see that we

  308. 10:34

    have a web client, iOS client that

  309. 10:36

    communicates with our API server that

  310. 10:37

    then dumps data into our turn database

  311. 10:39

    and is worked on by our harness, our our

  312. 10:41

    worker harness. This is sort of

  313. 10:43

    archetypical of probably what most of

  314. 10:45

    you are all building in the room and is

  315. 10:47

    exactly what we have on our back end.

  316. 10:49

    But that harness worker is a magic is a

  317. 10:51

    is a magic box and it connects to a

  318. 10:53

    number of important utilities. First is

  319. 10:55

    a sandbox where it can execute arbitrary

  320. 10:57

    shell calls uh do do Python data science

  321. 11:00

    etc. And we invite you to try coreweave

  322. 11:02

    weights and biases sandbox to fit into

  323. 11:04

    your architecture.

  324. 11:06

    Next up you need an LLM provider of

  325. 11:08

    course and so if you're maybe using GLM

  326. 11:10

    5.2 two or one of your fine-tuned

  327. 11:12

    models. We invite you to use uh weights

  328. 11:14

    and biases inference and connect that to

  329. 11:16

    your worker as well.

  330. 11:18

    If you're like us, you need to run

  331. 11:20

    longunning workloads outside of the main

  332. 11:22

    loop of the agent where you're actually

  333. 11:24

    training for day for sometimes days at a

  334. 11:26

    time. Weights and biases launch can

  335. 11:28

    actually help facilitate that and

  336. 11:29

    coreweave GPUs can help make that

  337. 11:31

    compute even better.

  338. 11:33

    And then lastly, and really most

  339. 11:35

    importantly, we need an observability

  340. 11:36

    layer. It's critical that your agents

  341. 11:38

    are able to log out their what's going

  342. 11:40

    on with their sessions, their turns,

  343. 11:42

    their tool calls, any errors they're hap

  344. 11:43

    that that's happening, etc. Uh we have a

  345. 11:46

    product called Weights and Biases Weave

  346. 11:48

    that we log 100% of our traces to where

  347. 11:50

    us and our team can learn from. And

  348. 11:52

    that's where we move from production to

  349. 11:54

    offline where our team is able to use

  350. 11:56

    Weights and Biases Weave to drive

  351. 11:58

    insights and identify behaviors,

  352. 12:00

    implement tasks with tasks which are

  353. 12:02

    essentially unit tests for your models

  354. 12:04

    and evaluate those models in a loop.

  355. 12:07

    We have a model repository which you

  356. 12:08

    might choose to use weights and biases

  357. 12:10

    artifacts to store your agents or models

  358. 12:12

    and you we emit our evaluation results

  359. 12:15

    to weave where we have a common

  360. 12:16

    dashboard that we can make go no-go

  361. 12:18

    decisions on various prompt changes or

  362. 12:20

    architectural changes that then feeds

  363. 12:23

    into a research loop which we call our

  364. 12:24

    improvement loop where we form

  365. 12:26

    hypotheses implement candidate agents

  366. 12:28

    and analyze the evals. So we have two

  367. 12:30

    sort of complimentary yet adversarial

  368. 12:32

    research loops going on going on offline

  369. 12:35

    feeding data from weights and biases

  370. 12:37

    weave ultimately to identify the best

  371. 12:39

    model so that we can promote that to

  372. 12:41

    production through our registry and

  373. 12:42

    close the data flywheel. So in the next

  374. 12:45

    just uh three seven minutes or so we'll

  375. 12:47

    just talk about uh weights and biases

  376. 12:49

    weave and show how we as a team actually

  377. 12:51

    use weave to facilitate this workflow

  378. 12:53

    and we believe this is something that

  379. 12:55

    you would benefit from as well all of

  380. 12:56

    you agent builders in the room.

  381. 12:59

    Yes, another demo. Great.

  382. 13:03

    Okay. Okay, we have new responses. So,

  383. 13:05

    it's going to be exciting when we open

  384. 13:06

    this up later. See if uh we've got some

  385. 13:09

    better metrics. Um, okay. Let me zoom

  386. 13:11

    out just a little bit here. So, here I'm

  387. 13:13

    looking at the agent dashboard. This is

  388. 13:15

    the live weights and biases agent or

  389. 13:18

    Arya agent dashboard uh built in weave.

  390. 13:21

    Man, that is a lot of uh branded

  391. 13:23

    buzzwords there. This is the dashboard

  392. 13:25

    that you would get if you use our tool.

  393. 13:26

    and uh you have a you know uh span

  394. 13:29

    volume, conversation volume, token

  395. 13:31

    tracking, etc. Think of this as like a

  396. 13:33

    uh a bird's eye view of your agent. For

  397. 13:36

    me, however, I really like this

  398. 13:38

    conversations view, which I do have

  399. 13:39

    pre-loaded in this tab. This

  400. 13:41

    conversations view is a live feed of all

  401. 13:44

    of the conversations that are going

  402. 13:45

    through Arya, but it's filtered down to

  403. 13:47

    just the internal employees. So, it's a

  404. 13:49

    little bit of a of a reduced set here.

  405. 13:51

    Um what I what I love is this middle

  406. 13:53

    spans view which gives me a visual

  407. 13:56

    indicator of the topology of a trace.

  408. 13:58

    Different colors and and shapes indicate

  409. 14:01

    different things that are happening

  410. 14:02

    within the agent. So things like tool

  411. 14:04

    calls, LLM calls, thinking blocks, etc.

  412. 14:06

    which really help me understand again

  413. 14:08

    the shape and topology of that

  414. 14:10

    particular conversation. I can of course

  415. 14:12

    open up one of these conversations and

  416. 14:14

    view our our conversation view where I

  417. 14:17

    can see the system prompt, the user

  418. 14:18

    message, shell calls, reasoning blocks,

  419. 14:21

    etc. This is where my research lead,

  420. 14:23

    myself and my PM go to add notes, add

  421. 14:26

    feedback, add emojis, and talk about and

  422. 14:28

    discover those insights and those

  423. 14:30

    behavioral nuances we spoke about

  424. 14:32

    earlier so that we can turn them into

  425. 14:33

    tasks.

  426. 14:35

    Arya's built in to the weights and

  427. 14:37

    biases system as well. Here you'll see a

  428. 14:39

    summarize button and these are sprinkled

  429. 14:40

    throughout the weights and biases

  430. 14:42

    application. I simply click summarize

  431. 14:44

    and we start a new chat contextualized

  432. 14:47

    to the thing that I'm looking at. So it

  433. 14:49

    it sees this and says give me a brief

  434. 14:51

    summary of this particular conversation.

  435. 14:53

    So if you if you're paying a attention

  436. 14:55

    closely, you'll realize that what we're

  437. 14:57

    doing is using Arya to analyze Arya's

  438. 15:00

    own conversations to then make

  439. 15:01

    recommendations about how to improve

  440. 15:02

    Arya all within the UI.

  441. 15:06

    Um okay, great. While that's cooking

  442. 15:08

    away, I want to show you the last item

  443. 15:10

    uh within the Weave ecosystem here, and

  444. 15:12

    that's signals. We've heard a lot today

  445. 15:14

    about the value of evals and the value

  446. 15:16

    of LLM judges. Weave actually offers an

  447. 15:18

    integrated LLM judge experience. So

  448. 15:21

    here, if I zoom out a little bit, you'll

  449. 15:23

    see that I have a user frustration

  450. 15:25

    signal, a lowquality response signal,

  451. 15:27

    ask user signal, etc. These are LLM

  452. 15:29

    judges that run live against against our

  453. 15:32

    live traffic. And we can see various

  454. 15:34

    different signals like user frustration

  455. 15:35

    moments or lowquality responses. These

  456. 15:38

    help our team identify these clusters of

  457. 15:40

    behavior for us to go fix in next week's

  458. 15:42

    iteration. Let's go ahead and do a live

  459. 15:44

    look and see what it says. Um this says

  460. 15:47

    the user explicitly states that I'm not

  461. 15:49

    satisfied with the loss curve. It looks

  462. 15:51

    bad and it apparently that indicates

  463. 15:53

    frustration. So here we can see an LLM

  464. 15:56

    judges live reasoning for why that

  465. 15:58

    particular flag was uh indicated.

  466. 16:02

    Uh let's see, four minutes left.

  467. 16:03

    Perfect. Um so, uh with that, I've been

  468. 16:06

    using the term task a lot. And so what

  469. 16:08

    we're do, what I've showed so far is is

  470. 16:10

    this live production loop where we are

  471. 16:12

    are are are tracing our our prod logs.

  472. 16:14

    We're looking at them as humans, maybe

  473. 16:16

    even using LLMs to complement that

  474. 16:17

    analysis. And what we end up doing is

  475. 16:20

    transforming those into tasks. Now, this

  476. 16:22

    gets a bit technical here, but our tasks

  477. 16:24

    are all described as YAML files. You can

  478. 16:26

    think of a task as essentially a unit

  479. 16:28

    test for your model. So here we say we

  480. 16:30

    have a an example user prompt that says

  481. 16:33

    check this run and that run. Both of

  482. 16:35

    these are giving good results. What can

  483. 16:37

    we learn from this? What's the

  484. 16:38

    difference? So this is an example of

  485. 16:40

    something we want Arya to be good at for

  486. 16:42

    all of you. And after the uh requisite

  487. 16:45

    metadata we see that we've defined an

  488. 16:47

    LLM judge. So here we've defined what

  489. 16:50

    correctness means in the context of that

  490. 16:52

    question.

  491. 16:53

    And we've then we've defined a second

  492. 16:55

    LLM judge that determines if the

  493. 16:58

    insights are actually interesting.

  494. 17:00

    [laughter] And then we've uh defined a

  495. 17:02

    third rule-based judge that says were

  496. 17:05

    you able to actually generate a result

  497. 17:07

    within just six tool calls meaning it

  498. 17:09

    got there with some degree of

  499. 17:10

    expediency. These are all then clustered

  500. 17:13

    together into we have about like 200 of

  501. 17:15

    these. They're all clustered together

  502. 17:17

    into an eval suite that runs nightly.

  503. 17:19

    And again we use weave to track all

  504. 17:21

    those evals. So here, I know it's a bit

  505. 17:23

    small on this screen, but what you're

  506. 17:25

    looking at is a listing of every night's

  507. 17:27

    eval. This is literally two nights ago,

  508. 17:29

    the evaluation for our candidate model

  509. 17:31

    got 73% on our production or on our eval

  510. 17:34

    suite against the 72% that our prod

  511. 17:37

    model got, which means we're definitely

  512. 17:38

    going to push that forward uh this

  513. 17:40

    Friday. Uh and we can see a kind of a a

  514. 17:42

    performance plot on the right. So these

  515. 17:44

    utilities are what you would get out of

  516. 17:45

    the box if you're uh if you decide to

  517. 17:47

    pick up weave and use this tool. Um,

  518. 17:50

    jumping back to the last conversation we

  519. 17:52

    had where it asked me where we asked,

  520. 17:54

    uh, can you please give a quick summary

  521. 17:56

    of this trace, we see that it actually

  522. 17:58

    analyzed the conversation, understood

  523. 18:00

    what the user was doing, and then

  524. 18:02

    ultimately decided that this was a

  525. 18:04

    pretty strong trace. Um, let's see,

  526. 18:07

    we've got two and a half minutes left,

  527. 18:08

    so let's just quickly recap here. Uh,

  528. 18:11

    first off, uh, what we use weave to do

  529. 18:13

    is a, collect production traffic. Super

  530. 18:15

    critical to collect all of your

  531. 18:16

    production traffic so you can learn and

  532. 18:17

    iterate. Secondly, we use it to generate

  533. 18:20

    insights both as humans as well. We we

  534. 18:22

    do it as humans. We use Arya and we use

  535. 18:24

    LLM judges to identify those behavioral

  536. 18:26

    nuances. We then enrich our tasks. We

  537. 18:30

    implement models and we evaluate using

  538. 18:32

    weights and biases weave as a shared

  539. 18:34

    dashboard where we can make decisions

  540. 18:35

    together as a team that then ultimately

  541. 18:38

    allows us to promote the best model

  542. 18:40

    forward with confidence.

  543. 18:43

    So speaking of confident

  544. 18:44

    productionization, let me speak uh

  545. 18:46

    briefly to the managers in the room. So

  546. 18:48

    a few tips for being successful here.

  547. 18:51

    First is um invest in agent-oriented

  548. 18:53

    observability. Uh I'm a bit biased. I

  549. 18:56

    believe that weights and biases weave is

  550. 18:57

    the uh observability platform of the

  551. 18:59

    future. Uh but pick your favorite

  552. 19:01

    flavor. Whatever it is, log your

  553. 19:03

    sessions, log your turns, log your tools

  554. 19:04

    and feedback. This introduces an ability

  555. 19:07

    to catch a new class of bugs in our

  556. 19:08

    world called behavioral bugs. Not

  557. 19:10

    exceptions, not performance, but

  558. 19:12

    behavioral bugs.

  559. 19:14

    Next up, tasks and evals are the new

  560. 19:16

    world of CI. You've heard a lot about

  561. 19:18

    this. If you are a software engineer,

  562. 19:19

    you've written unit tests your whole

  563. 19:21

    life. You must develop a practice where

  564. 19:23

    your researchers are sitting on the same

  565. 19:24

    scrum team as you developing tasks and

  566. 19:26

    you're viewing the performance metrics

  567. 19:28

    as true go no-go decisions. But in order

  568. 19:31

    to complement that, you must use humans

  569. 19:33

    as a necessary judge. There are

  570. 19:35

    behavioral nuances that LLM will not

  571. 19:37

    catch. You must be using your product

  572. 19:39

    and you must be manually reviewing these

  573. 19:41

    traces as a team at the end of the week

  574. 19:43

    on a board looking at the best and worst

  575. 19:45

    traces to understand how your model is

  576. 19:47

    performing.

  577. 19:49

    And then lastly, um just maybe one one

  578. 19:51

    more tip is to add value through context

  579. 19:53

    and tools. It can be really tempting to

  580. 19:56

    uh try to overengineer the harness and

  581. 19:58

    do a bunch of creative stuff around

  582. 19:59

    memory and things like this. We found

  583. 20:01

    that a a lot of lowhanging fruit can be

  584. 20:03

    ascertained through simply giving your

  585. 20:04

    agent context about your business

  586. 20:06

    domain, the underlying uh primitives

  587. 20:08

    that you have available and your

  588. 20:09

    particular uh business data. Um so with

  589. 20:12

    that, let's go ahead and check in on our

  590. 20:15

    uh our our research agent here and let's

  591. 20:19

    go ahead and toggle our workspace. And

  592. 20:21

    what we should be seeing is yes indeed a

  593. 20:24

    little dot that uh oh okay our previous

  594. 20:27

    dot which was done at lunch was 5.83.

  595. 20:30

    831. This got 5.833. So we were right on

  596. 20:33

    the edge of having a live improvement,

  597. 20:35

    but pretty darn close. Uh so that's what

  598. 20:37

    the uh that's what the model was able to

  599. 20:39

    produce. It actually uh ran uh quite a

  600. 20:42

    few tests here. I see I'm over time, so

  601. 20:44

    I will click close pretty soon. But we

  602. 20:46

    ran 12 different experiments within that

  603. 20:48

    experiment batch and uh we'll be running

  604. 20:50

    more all night. So please try out Arya,

  605. 20:52

    scan the QR codes, check out the docs.

  606. 20:54

    Uh we really love to see what you do

  607. 20:56

    with it and um looking forward to

  608. 20:57

    serving you. Thank you very much.

  609. 21:13

    >> [music]