AI Engineer World's Fair 2026
An AI Research Agent That Runs Your Experiments — Tim Sweeney, Weights & Biases
Read the talk
An AI Research Agent That Runs Your Experiments
ARIA turns a research request into GPU jobs, experiment analysis, and visual reports inside Weights & Biases. A live batch shows how that loop works; the team's own traces and nightly evaluations show how they decide whether the agent itself is improving.
From a talk by Tim Sweeney
At a glance
Ideas worth remembering
Keep long-running training outside the agent's main loop: ARIA starts experiments through Launch and polls while GPU jobs execute.
Evaluate useful research behavior along separate dimensions: the example task checks correctness, interesting insights, and a six-tool-call limit.
Production traces become more useful when human review and live judges turn observed behavior into repeatable tasks and candidate evaluations.
Domain context and available tools are practical places to improve an agent before adding elaborate harness or memory machinery.
The live batch completed 12 experiments and nearly matched the earlier best result; autonomous execution did not guarantee a better model.
Start with a project that already has 200 experiments
A research workspace can accumulate experiments faster than a person can explain them. Tim Sweeney, Principal Engineer at Weights & Biases by CoreWeave, opens a project with more than 200 training jobs. A scatter plot shows loss declining over time. The next problem is deciding what those runs teach—and what experiment to run next.
The project uses Karpathy's autoresearch, a small codebase that trains an LLM and provides a manageable setting for repeated changes. ARIA, the AI research and iteration agent, opens as a chat inside the workspace. Project context and images can accompany a request, so the conversation can begin with the experiment data already close at hand.
The useful starting point is a long-running conversation rather than the agent's onstage introduction. In that conversation, ARIA has helped download the code, configure Launch jobs, and set up GPUs. It can then iterate on both the code and its hyperparameters. Those are different kinds of changes: one alters the implementation being trained; the other changes how that implementation trains.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A chat request becomes a batch of GPU jobs
Sweeney asks ARIA to conduct another batch of experiments and try to find the best model live. He adds encouragement—because, naturally, the models need encouragement. The consequential decision comes next: ARIA avoids a large architecture change, treating it as too risky for this iteration. Sweeney expects it to try hyperparameter modifications instead, and a shell call starts the experimentation loop.
W&B Launch supplies the connection to compute. A Launch queue lets humans or agents submit long-running experiment jobs to a cluster, including jobs that need GPUs. Sweeney opens terminal output from his Kubernetes cluster to show experiments executing, then returns to ARIA, which is polling while it waits for the work to finish. The chat initiates and monitors the work; the cluster performs the training.
Where does the time-consuming work happen after a research request? The flow below separates the agent's decision and shell execution from queued GPU training. The polling path makes the separation visible: launching a job does not mean its result is immediately available.
Conduct another batch and look for a better model.
ARIA starts the experiment loop, Launch connects it to GPU compute, and the agent polls while training proceeds.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Turn accumulated runs into explanations and charts
While training runs, a second request asks for a summary of the project's highest-performing runs. This is useful when a teammate joins a project or returns from time away: the experiment history needs to become an explanation of what worked. A previously prepared pattern-finding example goes further, identifying an emerging model family, batch size as an influential parameter, and a promising architectural recipe. These are leads for further research, rather than a demonstrated guarantee that any one change causes an improvement.
The analysis can become a native W&B artifact rather than remaining a chat answer:
- Reports: A report combines Markdown with embedded plots, charts, and graphics—Sweeney calls it a “markdown file on steroids.” The example includes a project thesis, data panels, and a parameter-importance chart showing relationships among parameters.
- Workspaces: ARIA is tuned and prompted to construct workspaces and plots using W&B's built-in chart types. That gives the analysis a form colleagues can inspect in the interface they already use.
A check-in shows the summary request querying W&B, applying patches, and writing code, while the training conversation continues polling for results. Analysis and training therefore have their own work to complete. ARIA's value here comes from connecting research data, executable tools, and visualization primitives within the same product.
The iOS announcement extends that workflow to a phone: researchers can steer hyperparameter-tuning jobs while away from their desks. The larger ambition is an end-to-end automated research platform that handles job orchestration, GPU workloads, research lookup, and hypothesis collaboration. Sweeney frames the division of labor around letting ARIA handle mechanics while researchers concentrate on new ideas, architectures, and parameters.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Separate conversation handling from long-running compute
ARIA's backend begins with a familiar sequence: web and iOS clients communicate with an API server, the server writes to a turn database, and a worker harness processes the work. The harness then connects to the capabilities needed to carry out a research request.
Those connections serve distinct purposes:
- Sandbox execution: Shell calls and Python data-science work give the agent a way to manipulate code and analyze data.
- Model inference: An LLM provider supplies the model calls used by the worker; W&B Inference is presented as one option.
- External training: Launch handles workloads outside the agent's main loop, including training that can take days, with CoreWeave GPUs supplying compute.
- Observability: Sessions, turns, tool calls, and errors are logged so the team can inspect what the agent actually did. Sweeney says ARIA logs 100% of its traces to W&B Weave.
The trace store also feeds development. Offline work turns observed behavior into tasks, evaluates candidate agents, and records results in a shared dashboard. One loop improves the tests; another forms hypotheses, implements candidates, and studies their evaluation results. The loops are complementary but also adversarial: better tests can expose weaknesses in a candidate that previously looked promising. The selected agent returns to production through a registry.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Read the shape of a conversation, then inspect its failures
The Weave dashboard gives a bird's-eye view through span volume, conversation volume, and token tracking. Sweeney then opens a conversation feed filtered to internal employees. Its span view makes trace topology visible: different colors and shapes distinguish tool calls, LLM calls, and thinking blocks. Opening a conversation reveals the system prompt, user messages, shell calls, and reasoning blocks behind that shape.
That detailed view is where the research lead, engineer, and product manager add notes and feedback before turning behavioral discoveries into tasks. A contextual Summarize button can also start an ARIA conversation about the item being inspected. In this example, ARIA analyzes one of its own conversations to recommend improvements to ARIA. The agent assists the review, while the team still has the underlying trace available.
Signals add automated triage to live traffic. LLM judges flag user frustration, low-quality responses, and other behaviors, helping the team find clusters to address in a later iteration. One displayed judge identifies frustration because the user expresses dissatisfaction with a loss curve. Its reasoning is inspectable, which gives reviewers a concrete explanation to assess rather than just a flag. The flag identifies dissatisfaction; it does not by itself establish what caused the poor experience.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Turn observed behavior into nightly tests
A task is a YAML file that acts as a unit test for the agent. The demonstrated prompt asks ARIA to compare two runs that both give good results: what differs, and what can be learned? The task pairs that request with definitions of success. This makes a broad product expectation—give useful research advice—specific enough to evaluate repeatedly.
The example checks three separate properties:
- Correctness: An LLM judge evaluates the answer against correctness defined for this particular question.
- Interesting insights: A second LLM judge asks whether the comparison produces insights worth knowing. A correct answer can still be unhelpful.
- Expediency: A rule-based judge checks whether the agent produces a result within six tool calls, giving efficiency an explicit test alongside answer quality.
About 200 tasks form a suite that runs nightly, with results tracked in Weave. The displayed candidate scores 73%, compared with the production model's 72%; Sweeney says the team plans to move it forward that Friday. These scores describe performance on the team's suite. The recording does not establish whether the one-percentage-point difference is statistically significant or predicts improvement outside that suite.
How does a production conversation influence the next release? The cycle below shows the intermediate steps: review produces tasks, tasks test candidates, and the results inform a team decision. Production traces provide the starting material; they become useful release evidence only after the team defines the behavior it wants to test.
Collect conversations and agent actions.
Live behavior supplies examples; explicit tasks and candidate evaluations turn those examples into release decisions.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Keep humans in the loop—and let the experiment finish
The production advice follows directly from this development loop:
- Instrument behavior: Log sessions, turns, tools, and feedback to catch behavioral bugs. An agent can complete its calls without producing the behavior a user needs.
- Use evaluations for decisions: Researchers and engineers should develop tasks together and treat performance metrics as go/no-go evidence.
- Keep human judgment: Use the product and review its best and worst traces as a team. LLM judges miss behavioral nuances, so automated scores cannot cover the whole review.
- Improve context and tools: Before adding elaborate harness or memory machinery, give the agent the business domain, available primitives, and relevant business data. Sweeney's team found substantial opportunities in those simpler additions.
The final check returns to the batch started onstage. A new point has appeared in the workspace, and the batch has run 12 experiments. The reported result is 5.833, close to the earlier best but without beating it. ARIA has carried a request through experiment selection, shell execution, queued GPU training, waiting, and a visible result. The automation completed a research iteration; an improvement remained something the experiments had to earn.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
From the talk
The tracing and evaluation product used in the talk to inspect conversations, run live judge signals, and compare nightly candidate results.
Read the complete timestamped transcript
- 0:01
[music]
- 0:13
Okay. Hello everyone and thank you for
- 0:15
attending this session. My name is Tim
- 0:17
Sweeney, a principal engineer at Weights
- 0:20
and Biases and Coreweave. And for the
- 0:22
next 20 minutes, we're going to talk
- 0:23
about Arya, our new AI research and
- 0:26
iteration agent. Let's go ahead and get
- 0:28
started.
- 0:29
So, uh, first off, just by way of making
- 0:32
some noise, some clapping, uh, who here,
- 0:35
um, identifies as an ML researcher?
- 0:37
You're someone that trains models,
- 0:38
trains the brain?
- 0:41
I heard one. Wow. Okay. Great work.
- 0:43
Great work. Uh, what about who here is
- 0:45
the applied engineer, the namesake of
- 0:47
this conference? Who here actually
- 0:48
builds the bots?
- 0:50
>> Okay, good. Expected much more. And who
- 0:53
here is in AI management? You are
- 0:55
helping fund this compute.
- 0:58
Okay. Okay. Nice. From the back. Lovely.
- 1:01
Um, well, now that I know a little bit
- 1:02
about you, just a little bit about me.
- 1:04
Uh, again, my name is Tim. I have a
- 1:06
masters in machine learning, uh, and
- 1:07
reinforcement learning from Georgia
- 1:09
Tech. So, I've been that, uh, researcher
- 1:11
currently building Weights and Biases
- 1:13
agent, Arya. So, identify as that
- 1:15
applied engineer. And in a previous life
- 1:18
was the PM of Twitter's ML stack. So, I
- 1:21
hope you hopefully can connect with you
- 1:22
middle management as well. [laughter]
- 1:24
>> [gasps]
- 1:24
>> Um today's agenda is kind of broken into
- 1:27
three sections and hopefully each of you
- 1:29
personas walk away with something
- 1:30
valuable. So first we're going to learn
- 1:32
about Arya itself and how it can
- 1:34
supercharge your AI and ML workflows.
- 1:37
We're going to dive into auto research
- 1:38
and see that live in a live demo in just
- 1:40
a moment.
- 1:42
Then we're going to pull back the
- 1:43
curtain and learn how we use weights and
- 1:44
biases and uh coreweave to actually
- 1:47
build Arya because a lot of you in the
- 1:49
audience are building agents yourself
- 1:50
and we believe a lot of these components
- 1:52
can help you in your endeavors. And then
- 1:54
towards the end we'll just take a step
- 1:56
back and identify a few key tips and
- 1:58
tricks for making sure that you're able
- 1:59
to productionize your systems
- 2:00
effectively.
- 2:02
For those of you who might not be
- 2:03
familiar, Weights and Biases is the
- 2:05
world's leading AI development platform.
- 2:07
We've been in business now for nine
- 2:09
years and have happily joined the core
- 2:11
family about a year ago. Uh we have a
- 2:13
number of products in our suite but are
- 2:15
really known for our models training
- 2:17
inference and weave stack which really
- 2:19
helps collect data uh about the AI
- 2:21
development and machine learning
- 2:22
workflows and makes that information
- 2:24
actionable and uh enables users to make
- 2:27
the best decisions about what to do
- 2:28
next.
- 2:30
So without further ado, let's go ahead
- 2:32
and dive into Arya, our agent. Uh we'll
- 2:34
show a demo and then we'll get back to
- 2:35
some slides.
- 2:40
Okay, beautiful.
- 2:42
Let's make this a bit bigger. Holler at
- 2:44
me if you need it to be bigger. So, uh,
- 2:46
what you're looking at here is a weights
- 2:48
and biases workspace. For you, for
- 2:50
anybody that isn't familiar, on the
- 2:51
lefth hand side, I actually see a list
- 2:53
of a bunch of different experiments. In
- 2:55
this particular project, I have over 200
- 2:57
training jobs. And on the right hand
- 3:00
side, I see a scatter plot of, in this
- 3:02
case, declining metrics, which is good.
- 3:04
means our loss is going down over time.
- 3:06
And this view would be very familiar for
- 3:08
anyone that uses our tool. Now, to
- 3:10
ground this, we're actually uh uh using
- 3:13
the Carpathy Auto Research project,
- 3:15
which I'm sure many of you are familiar
- 3:16
with, but if you're not, it's just a
- 3:18
very simple project that trains an LLM,
- 3:20
and it's a great foundation for auto
- 3:23
research type demonstrations because
- 3:25
it's a very simple codebase and allows
- 3:27
us to improve iteratively over time. So,
- 3:30
let's jump back to the project and open
- 3:31
up Arya by clicking this blue button in
- 3:33
the upper right. When I click this
- 3:35
button, I'm uh presented with the
- 3:37
familiar chat interface with, you know,
- 3:39
how can I help you today, a few call to
- 3:40
actions, and you know, I can add
- 3:43
different context in my project or maybe
- 3:45
add images, etc. Um, everyone here is
- 3:48
agent builders, so I don't need to bore
- 3:49
you with the details of what an agent
- 3:51
interface looks like. But let's go ahead
- 3:53
and just, you know, enter in a basic
- 3:55
intro here. Let's say, "Hello, Arya.
- 3:58
You're on stage at AI World's Fair 2026.
- 4:00
Please introduce yourself." So, it's
- 4:02
going to go ahead and chug along and
- 4:03
hopefully emit some sort of nice emoji.
- 4:06
Yay. He I'm Arya. I'm talking to the
- 4:08
audience. Great. But now, let's dive
- 4:10
into the meat of why you came here. So,
- 4:13
I'm going to open up this chat here. And
- 4:15
this is a longunning chat where I've
- 4:17
been running again over 200 experiments
- 4:19
using the auto research loop. Um, it
- 4:22
helped me download the code, set up my
- 4:24
launch jobs, set up my GPUs, and is able
- 4:26
to autonomously iterate on the code
- 4:28
itself and the hyperparameters.
- 4:30
We'll take a look at what it's doing in
- 4:31
a moment, but while we're doing this,
- 4:33
I'm going to kick off a live iteration
- 4:35
right here. So, what I'm going to say is
- 4:37
please conduct another batch of
- 4:39
experiments. You are on stage at the AI
- 4:41
Engineer Worlds Fair 2026 and we're
- 4:44
hoping to find the best model live. I
- 4:46
believe in you. Uh, because we know we
- 4:48
have to encourage our models. Um, so
- 4:50
it's been doing this for a while. What
- 4:51
it what it's doing here is it's saying,
- 4:53
"Okay, great. Um, I don't want to make a
- 4:55
big architecture swing. That feels a
- 4:57
little bit too risky." So, it's probably
- 4:58
going to go for uh some modifications to
- 5:01
the hyperparameters. And then it's
- 5:03
kicking off a shell call here that is
- 5:05
actually um executing that uh executing
- 5:08
that experimentation loop. And we're
- 5:10
going to check in on this periodically
- 5:11
throughout this presentation.
- 5:13
But I want to help explain what's going
- 5:15
on behind the scenes. So behind the
- 5:17
scenes I have set up a weights and
- 5:19
biases launch queue. Launch is our our
- 5:21
product that allows you to connect your
- 5:23
compute clusters and allows humans and
- 5:25
agents to launch longunning
- 5:27
experimentation jobs particularly by
- 5:29
leveraging GPUs.
- 5:31
Here I'm looking at a uh a terminal
- 5:33
output of my Kubernetes cluster where
- 5:35
we're actually seeing live execution of
- 5:37
experiments happening. So this is
- 5:39
happening live right here. This is not a
- 5:40
fake demo. Um great. And if we jump
- 5:44
back, we see that at this point it
- 5:46
started the cues and now it is simply
- 5:48
polling and waiting for our work to be
- 5:50
complete. So we'll jump back to that in
- 5:52
a in a moment. But before but let's dive
- 5:55
into a few other examples. So uh
- 5:57
something else that is interesting you
- 5:59
can do is maybe you might want to ask it
- 6:02
something like please summarize the
- 6:03
highest performing runs in this project.
- 6:05
This use case would be something like
- 6:06
maybe a new user come or a new uh team
- 6:08
member is joining your project and want
- 6:10
to understand the research. Um or maybe
- 6:12
you've uh someone's been doing some work
- 6:14
while you were on PTO and you want to
- 6:16
get caught up. We'll see what this comes
- 6:17
up with in a moment. Some other
- 6:19
pre-anned uh examples are finding
- 6:21
patterns in your project. So here we can
- 6:24
see that I asked it, hey, can you find
- 6:26
some patterns in this research? And we
- 6:28
see that um it identified that a new
- 6:30
family of models emerged as the as the
- 6:33
auto uh auto research was happening. Uh
- 6:35
it identified that batch size seems to
- 6:37
be a really high high uh lever uh
- 6:40
parameter. It identified an
- 6:42
architectural recipe that seemed to be
- 6:44
quite promising and a number of other
- 6:46
insights that would have taken me hours
- 6:48
or days to discover on my own. And Arya
- 6:50
is able to do it right for me directly
- 6:52
in the interface that I already live.
- 6:55
Not only is it able to emit text based
- 6:57
uh textbased outputs, but it also deeply
- 7:00
integrates with a number of weights and
- 7:01
biases visualization utilities. So here
- 7:04
I've actually asked it to emit a weights
- 7:06
and biases report which for those who
- 7:08
aren't familiar is essentially a
- 7:09
markdown file on steroids. It's got uh
- 7:12
embedded embedded plots, charts and and
- 7:14
and graphics. And so here uh you know
- 7:16
it's talked about the thesis of the
- 7:18
project. It's it's emitted a number of
- 7:20
of data panels. And uh I actually think
- 7:23
it's quite interesting. It used um one
- 7:25
of our more esoteric panels, the uh
- 7:27
parameter importance chart to uh tell me
- 7:30
the correlation of various different
- 7:32
parameters within this uh within this
- 7:33
training job.
- 7:35
Uh in addition to uh reports, it's also
- 7:38
great at working with workspaces. So if
- 7:40
you're a weights and biases user, uh you
- 7:43
spend a lot of your time uh designing
- 7:45
and working with workspaces. Well, Arya
- 7:48
is actually customtuned and prompted to
- 7:50
really understand how to build
- 7:51
workspaces, build plots, and complement
- 7:54
that that data analytics with real live
- 7:56
graphics using the built-in proprietary
- 7:58
charts that weights and biases users
- 8:00
know and love. Um, so with that, let's
- 8:03
go ahead and check back on some of our
- 8:04
our prompts. We can see that the please
- 8:06
summarize this project prompt is cooking
- 8:08
away. It's querying weights and biases.
- 8:10
It's applying patches. It's writing its
- 8:12
own code. So, we'll come back and check
- 8:13
on that in a moment. and our longunning
- 8:15
training job is uh still pulling for the
- 8:18
results. We can see that we're cooking
- 8:20
away on our GPUs. So, we're we're frying
- 8:22
some GPUs [music] and doing some data
- 8:24
science all live. And while that's
- 8:26
cooking, let's go ahead and jump back to
- 8:27
the presentation. We'll come back in a
- 8:28
moment.
- 8:30
Uh
- 8:33
oh, no, we're not looking at a
- 8:34
dictionary. We're looking at a PO.
- 8:36
Great. Uh okay, so quick recap here.
- 8:38
What did Arya show? What did we show in
- 8:40
these last five minutes? First, we show
- 8:42
that uh Arya can serve as your data
- 8:44
science companion right inside of
- 8:46
Weights and Biases, helping you discover
- 8:48
insights that you wouldn't you wouldn't
- 8:49
be able to discover as your experiments
- 8:51
and as your team size grows.
- 8:55
Next, we address the problem of
- 8:57
complicated reporting and complicated
- 8:58
plotting. Weights and biases users are
- 9:00
are really want to turn their insights
- 9:02
into visual communication tools. They
- 9:04
want to communicate with their peers and
- 9:06
their colleagues. So Arya's built from
- 9:08
the ground up to understand those
- 9:09
primitives and help co-pilot and drive
- 9:12
right along right alongside in the UI
- 9:15
and announcing now today for the first
- 9:17
time we are releasing Arya on our iOS
- 9:19
device or on our iOS app. So uh uh Arya
- 9:23
released on Monday and our iOS app now
- 9:26
has Arya built in. So if you're
- 9:27
conducting hyperparameter tuning jobs,
- 9:29
if you're training models, or if you're
- 9:31
just researching within the weights and
- 9:32
biases ecosystem, you can go touched
- 9:34
grass at Yerba Buena uh gardens and
- 9:37
steer your uh hyperparameter tuning jobs
- 9:39
all from your mobile device. And what is
- 9:41
this all building up to? This is
- 9:43
building up to an a fully automated
- 9:45
endto-end research platform where we're
- 9:47
not seeking to replace uh RL
- 9:49
researchers, but complement your
- 9:50
workflows. Arya's great at orchestrating
- 9:53
jobs, understanding GPU workloads,
- 9:55
responding to events within the within
- 9:57
the Wandi ecosystem, and listening to
- 9:59
researchers, uh, uh, looking up archive
- 10:01
papers, and collaborating on hypothesis.
- 10:04
So, we can let Arya drive the mechanics
- 10:05
that you don't want to deal with while
- 10:07
you focus on the new ideas, new
- 10:08
architectures, and new parameters that
- 10:10
you wanted to try.
- 10:12
Um, great. So, that's Arya in a
- 10:14
nutshell. We're really hoping that you
- 10:16
give it a shot. And uh we'll jump back
- 10:18
to the auto research at the end and see
- 10:20
if we got a new best record. But before
- 10:22
we do that, let's talk about how we use
- 10:24
weights and biases and coreweave to
- 10:26
actually build Arya. So now speaking to
- 10:28
a lot of the the AI agent builders in
- 10:30
the room, here's a quick architecture on
- 10:32
the lefth hand side. You see that we
- 10:34
have a web client, iOS client that
- 10:36
communicates with our API server that
- 10:37
then dumps data into our turn database
- 10:39
and is worked on by our harness, our our
- 10:41
worker harness. This is sort of
- 10:43
archetypical of probably what most of
- 10:45
you are all building in the room and is
- 10:47
exactly what we have on our back end.
- 10:49
But that harness worker is a magic is a
- 10:51
is a magic box and it connects to a
- 10:53
number of important utilities. First is
- 10:55
a sandbox where it can execute arbitrary
- 10:57
shell calls uh do do Python data science
- 11:00
etc. And we invite you to try coreweave
- 11:02
weights and biases sandbox to fit into
- 11:04
your architecture.
- 11:06
Next up you need an LLM provider of
- 11:08
course and so if you're maybe using GLM
- 11:10
5.2 two or one of your fine-tuned
- 11:12
models. We invite you to use uh weights
- 11:14
and biases inference and connect that to
- 11:16
your worker as well.
- 11:18
If you're like us, you need to run
- 11:20
longunning workloads outside of the main
- 11:22
loop of the agent where you're actually
- 11:24
training for day for sometimes days at a
- 11:26
time. Weights and biases launch can
- 11:28
actually help facilitate that and
- 11:29
coreweave GPUs can help make that
- 11:31
compute even better.
- 11:33
And then lastly, and really most
- 11:35
importantly, we need an observability
- 11:36
layer. It's critical that your agents
- 11:38
are able to log out their what's going
- 11:40
on with their sessions, their turns,
- 11:42
their tool calls, any errors they're hap
- 11:43
that that's happening, etc. Uh we have a
- 11:46
product called Weights and Biases Weave
- 11:48
that we log 100% of our traces to where
- 11:50
us and our team can learn from. And
- 11:52
that's where we move from production to
- 11:54
offline where our team is able to use
- 11:56
Weights and Biases Weave to drive
- 11:58
insights and identify behaviors,
- 12:00
implement tasks with tasks which are
- 12:02
essentially unit tests for your models
- 12:04
and evaluate those models in a loop.
- 12:07
We have a model repository which you
- 12:08
might choose to use weights and biases
- 12:10
artifacts to store your agents or models
- 12:12
and you we emit our evaluation results
- 12:15
to weave where we have a common
- 12:16
dashboard that we can make go no-go
- 12:18
decisions on various prompt changes or
- 12:20
architectural changes that then feeds
- 12:23
into a research loop which we call our
- 12:24
improvement loop where we form
- 12:26
hypotheses implement candidate agents
- 12:28
and analyze the evals. So we have two
- 12:30
sort of complimentary yet adversarial
- 12:32
research loops going on going on offline
- 12:35
feeding data from weights and biases
- 12:37
weave ultimately to identify the best
- 12:39
model so that we can promote that to
- 12:41
production through our registry and
- 12:42
close the data flywheel. So in the next
- 12:45
just uh three seven minutes or so we'll
- 12:47
just talk about uh weights and biases
- 12:49
weave and show how we as a team actually
- 12:51
use weave to facilitate this workflow
- 12:53
and we believe this is something that
- 12:55
you would benefit from as well all of
- 12:56
you agent builders in the room.
- 12:59
Yes, another demo. Great.
- 13:03
Okay. Okay, we have new responses. So,
- 13:05
it's going to be exciting when we open
- 13:06
this up later. See if uh we've got some
- 13:09
better metrics. Um, okay. Let me zoom
- 13:11
out just a little bit here. So, here I'm
- 13:13
looking at the agent dashboard. This is
- 13:15
the live weights and biases agent or
- 13:18
Arya agent dashboard uh built in weave.
- 13:21
Man, that is a lot of uh branded
- 13:23
buzzwords there. This is the dashboard
- 13:25
that you would get if you use our tool.
- 13:26
and uh you have a you know uh span
- 13:29
volume, conversation volume, token
- 13:31
tracking, etc. Think of this as like a
- 13:33
uh a bird's eye view of your agent. For
- 13:36
me, however, I really like this
- 13:38
conversations view, which I do have
- 13:39
pre-loaded in this tab. This
- 13:41
conversations view is a live feed of all
- 13:44
of the conversations that are going
- 13:45
through Arya, but it's filtered down to
- 13:47
just the internal employees. So, it's a
- 13:49
little bit of a of a reduced set here.
- 13:51
Um what I what I love is this middle
- 13:53
spans view which gives me a visual
- 13:56
indicator of the topology of a trace.
- 13:58
Different colors and and shapes indicate
- 14:01
different things that are happening
- 14:02
within the agent. So things like tool
- 14:04
calls, LLM calls, thinking blocks, etc.
- 14:06
which really help me understand again
- 14:08
the shape and topology of that
- 14:10
particular conversation. I can of course
- 14:12
open up one of these conversations and
- 14:14
view our our conversation view where I
- 14:17
can see the system prompt, the user
- 14:18
message, shell calls, reasoning blocks,
- 14:21
etc. This is where my research lead,
- 14:23
myself and my PM go to add notes, add
- 14:26
feedback, add emojis, and talk about and
- 14:28
discover those insights and those
- 14:30
behavioral nuances we spoke about
- 14:32
earlier so that we can turn them into
- 14:33
tasks.
- 14:35
Arya's built in to the weights and
- 14:37
biases system as well. Here you'll see a
- 14:39
summarize button and these are sprinkled
- 14:40
throughout the weights and biases
- 14:42
application. I simply click summarize
- 14:44
and we start a new chat contextualized
- 14:47
to the thing that I'm looking at. So it
- 14:49
it sees this and says give me a brief
- 14:51
summary of this particular conversation.
- 14:53
So if you if you're paying a attention
- 14:55
closely, you'll realize that what we're
- 14:57
doing is using Arya to analyze Arya's
- 15:00
own conversations to then make
- 15:01
recommendations about how to improve
- 15:02
Arya all within the UI.
- 15:06
Um okay, great. While that's cooking
- 15:08
away, I want to show you the last item
- 15:10
uh within the Weave ecosystem here, and
- 15:12
that's signals. We've heard a lot today
- 15:14
about the value of evals and the value
- 15:16
of LLM judges. Weave actually offers an
- 15:18
integrated LLM judge experience. So
- 15:21
here, if I zoom out a little bit, you'll
- 15:23
see that I have a user frustration
- 15:25
signal, a lowquality response signal,
- 15:27
ask user signal, etc. These are LLM
- 15:29
judges that run live against against our
- 15:32
live traffic. And we can see various
- 15:34
different signals like user frustration
- 15:35
moments or lowquality responses. These
- 15:38
help our team identify these clusters of
- 15:40
behavior for us to go fix in next week's
- 15:42
iteration. Let's go ahead and do a live
- 15:44
look and see what it says. Um this says
- 15:47
the user explicitly states that I'm not
- 15:49
satisfied with the loss curve. It looks
- 15:51
bad and it apparently that indicates
- 15:53
frustration. So here we can see an LLM
- 15:56
judges live reasoning for why that
- 15:58
particular flag was uh indicated.
- 16:02
Uh let's see, four minutes left.
- 16:03
Perfect. Um so, uh with that, I've been
- 16:06
using the term task a lot. And so what
- 16:08
we're do, what I've showed so far is is
- 16:10
this live production loop where we are
- 16:12
are are are tracing our our prod logs.
- 16:14
We're looking at them as humans, maybe
- 16:16
even using LLMs to complement that
- 16:17
analysis. And what we end up doing is
- 16:20
transforming those into tasks. Now, this
- 16:22
gets a bit technical here, but our tasks
- 16:24
are all described as YAML files. You can
- 16:26
think of a task as essentially a unit
- 16:28
test for your model. So here we say we
- 16:30
have a an example user prompt that says
- 16:33
check this run and that run. Both of
- 16:35
these are giving good results. What can
- 16:37
we learn from this? What's the
- 16:38
difference? So this is an example of
- 16:40
something we want Arya to be good at for
- 16:42
all of you. And after the uh requisite
- 16:45
metadata we see that we've defined an
- 16:47
LLM judge. So here we've defined what
- 16:50
correctness means in the context of that
- 16:52
question.
- 16:53
And we've then we've defined a second
- 16:55
LLM judge that determines if the
- 16:58
insights are actually interesting.
- 17:00
[laughter] And then we've uh defined a
- 17:02
third rule-based judge that says were
- 17:05
you able to actually generate a result
- 17:07
within just six tool calls meaning it
- 17:09
got there with some degree of
- 17:10
expediency. These are all then clustered
- 17:13
together into we have about like 200 of
- 17:15
these. They're all clustered together
- 17:17
into an eval suite that runs nightly.
- 17:19
And again we use weave to track all
- 17:21
those evals. So here, I know it's a bit
- 17:23
small on this screen, but what you're
- 17:25
looking at is a listing of every night's
- 17:27
eval. This is literally two nights ago,
- 17:29
the evaluation for our candidate model
- 17:31
got 73% on our production or on our eval
- 17:34
suite against the 72% that our prod
- 17:37
model got, which means we're definitely
- 17:38
going to push that forward uh this
- 17:40
Friday. Uh and we can see a kind of a a
- 17:42
performance plot on the right. So these
- 17:44
utilities are what you would get out of
- 17:45
the box if you're uh if you decide to
- 17:47
pick up weave and use this tool. Um,
- 17:50
jumping back to the last conversation we
- 17:52
had where it asked me where we asked,
- 17:54
uh, can you please give a quick summary
- 17:56
of this trace, we see that it actually
- 17:58
analyzed the conversation, understood
- 18:00
what the user was doing, and then
- 18:02
ultimately decided that this was a
- 18:04
pretty strong trace. Um, let's see,
- 18:07
we've got two and a half minutes left,
- 18:08
so let's just quickly recap here. Uh,
- 18:11
first off, uh, what we use weave to do
- 18:13
is a, collect production traffic. Super
- 18:15
critical to collect all of your
- 18:16
production traffic so you can learn and
- 18:17
iterate. Secondly, we use it to generate
- 18:20
insights both as humans as well. We we
- 18:22
do it as humans. We use Arya and we use
- 18:24
LLM judges to identify those behavioral
- 18:26
nuances. We then enrich our tasks. We
- 18:30
implement models and we evaluate using
- 18:32
weights and biases weave as a shared
- 18:34
dashboard where we can make decisions
- 18:35
together as a team that then ultimately
- 18:38
allows us to promote the best model
- 18:40
forward with confidence.
- 18:43
So speaking of confident
- 18:44
productionization, let me speak uh
- 18:46
briefly to the managers in the room. So
- 18:48
a few tips for being successful here.
- 18:51
First is um invest in agent-oriented
- 18:53
observability. Uh I'm a bit biased. I
- 18:56
believe that weights and biases weave is
- 18:57
the uh observability platform of the
- 18:59
future. Uh but pick your favorite
- 19:01
flavor. Whatever it is, log your
- 19:03
sessions, log your turns, log your tools
- 19:04
and feedback. This introduces an ability
- 19:07
to catch a new class of bugs in our
- 19:08
world called behavioral bugs. Not
- 19:10
exceptions, not performance, but
- 19:12
behavioral bugs.
- 19:14
Next up, tasks and evals are the new
- 19:16
world of CI. You've heard a lot about
- 19:18
this. If you are a software engineer,
- 19:19
you've written unit tests your whole
- 19:21
life. You must develop a practice where
- 19:23
your researchers are sitting on the same
- 19:24
scrum team as you developing tasks and
- 19:26
you're viewing the performance metrics
- 19:28
as true go no-go decisions. But in order
- 19:31
to complement that, you must use humans
- 19:33
as a necessary judge. There are
- 19:35
behavioral nuances that LLM will not
- 19:37
catch. You must be using your product
- 19:39
and you must be manually reviewing these
- 19:41
traces as a team at the end of the week
- 19:43
on a board looking at the best and worst
- 19:45
traces to understand how your model is
- 19:47
performing.
- 19:49
And then lastly, um just maybe one one
- 19:51
more tip is to add value through context
- 19:53
and tools. It can be really tempting to
- 19:56
uh try to overengineer the harness and
- 19:58
do a bunch of creative stuff around
- 19:59
memory and things like this. We found
- 20:01
that a a lot of lowhanging fruit can be
- 20:03
ascertained through simply giving your
- 20:04
agent context about your business
- 20:06
domain, the underlying uh primitives
- 20:08
that you have available and your
- 20:09
particular uh business data. Um so with
- 20:12
that, let's go ahead and check in on our
- 20:15
uh our our research agent here and let's
- 20:19
go ahead and toggle our workspace. And
- 20:21
what we should be seeing is yes indeed a
- 20:24
little dot that uh oh okay our previous
- 20:27
dot which was done at lunch was 5.83.
- 20:30
831. This got 5.833. So we were right on
- 20:33
the edge of having a live improvement,
- 20:35
but pretty darn close. Uh so that's what
- 20:37
the uh that's what the model was able to
- 20:39
produce. It actually uh ran uh quite a
- 20:42
few tests here. I see I'm over time, so
- 20:44
I will click close pretty soon. But we
- 20:46
ran 12 different experiments within that
- 20:48
experiment batch and uh we'll be running
- 20:50
more all night. So please try out Arya,
- 20:52
scan the QR codes, check out the docs.
- 20:54
Uh we really love to see what you do
- 20:56
with it and um looking forward to
- 20:57
serving you. Thank you very much.
- 21:13
>> [music]