AI Engineer World's Fair 2026
Robot Demos Are Easy. Reliability Is Hard — Jason Ma, Dyna Robotics
Read the talk
Robot Demos Are Easy. Reliability Is Hard
Jason Ma explains how Dyna Robotics combines broad robot training with targeted recovery demonstrations—and why learning what to do after a mistake matters as much as learning the task itself.
From a talk by Jason Ma
At a glance
Ideas worth remembering
Commercial reliability includes recovering from mistakes and meeting customer quality criteria, as well as completing the nominal task.
A video-based reward model helps locate failures through changes in estimated progress. Humans collect recovery demonstrations for those cases, then fine-tune and repeat.
The reported 99.4% result concerns napkin folding over 24 hours. Task mastery and transfer to a new customer site require separate evidence.
Broad pre-training can supply recovery knowledge from other tasks; task-specific post-training and active learning turn that foundation into a useful commercial workflow.
Deployment gives research a useful problem
A robot that can perform a useful physical task once has cleared an important hurdle. A robot that can keep doing it in a customer's workplace faces a different test: changing conditions, imperfect hardware, and mistakes that alter what happens next. Jason Ma, co-founder and CTO of Dyna Robotics, introduces the company's goal as one platform that can perform many economically useful tasks with enough reliability for commercial use.
The research and deployment cycle starts with what the models and hardware can already do. Those capabilities determine which customer workflows Dyna can attempt. Deployment then supplies both training data and a sharper diagnosis: which parts of the model or hardware still fail? Ma reports more than five deployment sites at the time of the talk. Their value goes beyond demonstrating progress; they help choose which of robotics' many research problems deserve attention next.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Broad data supplies knowledge; robot data supplies precision
The basic manipulation pipeline begins with demonstrations. A human controls the robot to perform a task, such as folding a T-shirt, and the resulting data trains a neural network. At execution time, the network receives a representation of the scene—in Ma's introductory example, a camera view of the table—and outputs joint positions or torques that move the robot. Learning the visual task and learning the physical action are connected through those demonstrations.
Scaling that pipeline runs into a data problem: robot demonstrations are scarce. Dyna's pre-training data pyramid combines three sources with different jobs:
- Off-robot data: Human-worn camera recordings and public datasets provide diverse experience without requiring every example to be collected on a robot. Simulation data is described as something the team is considering.
- Robot-task data: Demonstrations across industrial, household, laundromat, and hotel tasks teach the precise actions that off-robot observations alone cannot supply.
- Deployment data: Experience from customer environments helps close the gap between training in a laboratory and operating somewhere else.
Ma reports more than 200,000 hours of data in the training pipeline, across these sources.
The architecture pairs a high-level reasoning model with a low-level world action model. The reasoning component supplies semantic understanding of the situation; the action component produces fine-grained, high-frequency movements. Physical work requires both: understanding what should happen does not by itself supply the precise interaction needed to complete a fold or recover from a bad grasp. The talk gives this division of responsibility without specifying the interface or control frequency.
A useful pre-trained model makes new tasks cheaper to learn. Ma presents two dexterous tool-use examples trained with less than one hour of task-specific data, including tasks absent from pre-training. The office gallery also includes cleaning trash, opening a box, and folding towels and T-shirts. These examples establish the appeal of a general foundation: each new task can build on capabilities already learned. But repeated commercial work raises the next question—how often does the robot fail, and what happens when it does?
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
An impressive success rate can still interrupt the workflow
Ma places deployed generalist manipulation models at roughly 80–90% success. That leaves frequent interruptions in a repetitive job. To see the compounding effect, under an illustrative assumption of independent attempts with constant success probability (p), ten consecutive successes have probability (p^{10}): about 10.7% at 80% success and 34.9% at 90%. The recording's captioned figure of less than 0.1% for ten repetitions is inconsistent with that calculation, so it cannot support the numerical comparison. The underlying concern remains: reliable individual attempts and long uninterrupted runs are different requirements.
Specialized systems offer another path: build a robot and learning pipeline for one operation, such as pick and place. The tradeoff is reuse. A pipeline built around one task cannot necessarily adapt quickly to an arbitrary new one. Dyna's goal is to retain the generalist model's range while reaching the reliability a customer needs for a specific workflow.
Restaurant napkin folding gives that goal a concrete test. Employees fold napkins individually in the back of the restaurant, and customers Dyna spoke with wanted to automate the work. The task is repetitive, commercially useful, and difficult enough to expose weaknesses in manipulation. Narrowing the research to this workflow makes the desired result tangible: keep producing napkins the restaurant will accept.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Napkin folding tests grasping, quality, and recovery
Dyna-1, a generalist robot foundation model fine-tuned for napkin folding, achieved a reported 99.4% success rate over a 24-hour run. Ma also presents four distinct 24-hour office trials. In the time-lapse, daylight arrives around 7 a.m. and changes the lighting while the robot keeps folding. These are speaker-reported results; the supplied material gives neither evaluation sample sizes nor independent verification, so the percentage should remain attached to this napkin-folding evaluation rather than become a reliability claim for every task or site.
The operation has three nominal steps: pick exactly one napkin from the stack, fold it, and place it in a bin. The first step already creates trouble. Parallel-jaw grippers may pull out extra napkins, leaving the robot with a scene that differs from a clean demonstration. The fold also has to meet a customer's quality standard. Ma contrasts an acceptable grade-five fold with an unacceptable grade-three fold, separated by roughly one inch in the position of the second fold seam. Moving cloth into a folded shape is only part of success; the final geometry matters.
Dyna's initial pre-training and post-training recipe reached roughly 80% success on this task. The recurring problem appeared after a mistake: the robot entered an unfamiliar state, got stuck, and could not recover. More successful demonstrations would teach the normal path, but the missing behavior was what to do after leaving that path. Reliability therefore required training on the situations the robot created through its own errors.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Progress scores turn rare mistakes into training targets
Dyna adds a reward model that watches robot video and estimates progress toward completing a task. In the examples, progress moves from zero toward one as the operation finishes. A successful napkin fold produces a roughly rising estimate; mistakes produce dips. During the ninth napkin in one autonomous run, the estimate stops rising monotonically as the robot makes an error. The useful signal is the change in progress, which helps identify where the operation went wrong.
Once the policy succeeds around 90% of the time, having a person watch every attempt is an inefficient way to find the remaining failures. The reward model runs in the background and directs researchers or operators to the video cases that need attention. Humans then collect targeted demonstrations of recovery from those states, and Dyna fine-tunes the policy again. This is a human-in-the-loop active-learning cycle: autonomous operation finds the gaps, progress scoring helps locate them, and human demonstrations supply the missing actions.
Where does the progress signal change the training process? The diagram follows the loop from a napkin-folding attempt back to an improved policy. Its key relationship is the handoff from detecting trouble to collecting recovery data: the reward model helps choose what to teach, while the policy learns how to act. Ma reports that repeated cycles produced a model able to recover from a wider range of errors.
The policy acts while producing video of its attempts.
Video scoring selects difficult cases for human recovery demonstrations; fine-tuning changes the policy used in the next autonomous run.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A bad grasp becomes a recoverable state
Return to the single-napkin pickup. In a recovery highlight, the robot accidentally takes more than one napkin. Instead of remaining stuck with the extra cloth, it separates the napkins and continues making progress. The observable change is from a failed selection to a state in which folding can proceed. Targeted recovery training addresses precisely this gap: the policy needs useful actions for the aftermath of a bad grasp, as well as actions for the clean pickup.
Cloth makes exhaustive recovery coverage impractical. A napkin can deform into an enormous range of configurations, so the team cannot demonstrate every possible error. Ma reports that repeated active learning produced recovery behavior that generalized to new mistakes. The strongest example is a robot pulling over the entire napkin stack, then using pulling and stretching movements to recover and keep working. Continuous operation depends on handling the mess the robot creates, rather than assuming the table always remains in its expected state.
The commercial examples move the same recipe into customer workflows. Dyna deploys napkin-folding robots in restaurant back offices and fine-tunes a model for towel folding at a Sacramento laundromat. In the laundromat example, a worker repeatedly replenishes the bin while the robot folds stacks of towels. The automation handles the folding operation within a workflow that still includes human material handling.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Task mastery and site generalization are separate tests
The restaurant and laundromat examples come with an important qualification: Dyna collected data at the customer sites used for those deployments. Mastering a task in one location does not automatically establish that the model will perform it somewhere unfamiliar. Collecting and adapting at every new site adds work to expansion, so the next objective is to preserve task performance without additional site data.
Ma describes a diverse-task data recipe intended to enable that transfer, then presents T-shirt folding at CoRL 2025 in Korea as the example. The robot arrived at an exhibition booth and began folding different T-shirts without additional site data. It faced attendees, including people who deliberately covered its camera with a shirt, and reportedly continued folding over three days at the conference. This tests environmental transfer and resilience to disruption; it does not establish the same 99.4% figure reported for napkin folding.
A Red Bull partnership extends the examples to opening cans at live events. Ma describes repeated operation at a music festival despite changing background lighting, but the demonstration video did not play properly during the presentation. This example rests on his account. Together, the deployments point toward the desired product experience: bring a robot to a site and have it start useful work, while retaining the ability to perform a chosen commercial task well.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Enterprise deployment comes before a household product
The audience questions clarify what Dyna is building now and what remains an ambition:
- Developer tools: Dyna is building its own hardware stack and is not currently working on a developer kit. A kit may become part of the roadmap later.
- Education: Teaching children to train robots is an interesting future direction for Ma, but Dyna has not explored that use case.
- Consumer robots: The long-term goal includes a model and hardware platform usable by anyone, including households. The current go-to-market focus is enterprise customers.
Enterprise customers let Dyna deploy and iterate at the site. A consumer product would need more polish, tolerate fewer mistakes, and address privacy and safety concerns that the company is not taking on at this stage. Ma treats commercial environments as a useful testing ground for the immediate bottleneck: getting the AI and hardware to work well together. The sequencing follows the research cycle introduced at the start—choose a workable deployment, learn from it, and improve the system.
Voice commands could fit above the manipulation stack. Ma sketches an approach in which onboard speech-to-text converts a person's speech into text, then the reasoning and world models interpret the instruction and produce actions. He offers this as a possible integration rather than a demonstrated voice system. Recognizing the command would still leave the harder capability problem: Dyna's models and other robot foundation models are not ready to execute arbitrary instructions. A robot that anyone can teach or steer through language remains the goal; low-level manipulation is still a limiting step.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Recovery draws on experience beyond the chosen task
The final question asks how intelligence should be divided between general pre-training, task-specific training, and adaptation at a deployment site. Ma adds a consequential explanation for the recovery demonstrations: some of the behavior came from interpolating recovery experience gathered on other tasks. A model exposed to thousands of tasks can draw on physical knowledge that a model specialized from the outset would miss.
That gives the three training stages distinct purposes. General pre-training builds semantic and physical understanding, including experience with recovery. Post-training adapts that foundation to the chosen job. Active learning then concentrates demonstrations on the mistakes that remain during operation. For napkin folding, the resulting policy needs both a precise routine and enough transferable physical knowledge to find a way forward when the routine goes wrong. Dyna's intended reusable product is this recipe across tasks, rather than a separate learning pipeline for each one.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Read the complete timestamped transcript
- 0:01
[music]
- 0:12
>> Thanks for the kind introduction. So,
- 0:13
I'm Jason, co-founder of Diana.
- 0:16
Today, I'll talk about how we're
- 0:18
developing high-performance and very
- 0:20
robust generalist robotics policies.
- 0:23
Yeah.
- 0:24
So, let's jump into it. So, today's talk
- 0:27
will focus on, you know, how we're
- 0:28
bringing robots into commercial grade
- 0:30
and what we are doing to make these
- 0:32
models, like I said, very
- 0:33
high-performance and robust. So, just a
- 0:35
brief introduction to our company. So,
- 0:37
our mission is to build very robust
- 0:40
foundation model autonomy in the real
- 0:41
world. So, we want to train models and
- 0:43
these days also hardware to have a
- 0:46
single platform that can do many
- 0:48
economically useful physical tasks in
- 0:50
the real world, like the ones we're
- 0:52
showing here. So, the company the
- 0:54
company was founded in September 2024.
- 0:56
We are a series A company, have raised
- 0:58
about $120 million, and we
- 1:03
And our thesis is that to actually bring
- 1:05
robots into the real world, being at
- 1:08
commercial grade doing useful tasks, the
- 1:10
company needs to combine doing frontier
- 1:12
research with a lot of commercial
- 1:14
deployments, so we can build what we
- 1:16
consider a research and deployment
- 1:18
flywheel.
- 1:19
Uh which is that the research and
- 1:21
hardware we do in-house, the R&D informs
- 1:24
the kind of tasks, the kind of workflows
- 1:26
in the real world that we can
- 1:27
commercialize. So, we actively try to
- 1:29
deploy. So, these days we have more than
- 1:31
five deployment sites doing a bunch of
- 1:33
different tasks, which I will talk about
- 1:35
in a bit. And then by building product
- 1:38
and by deploying robots, it does several
- 1:40
things. One is that it can help us
- 1:42
gather high-quality deployment data, and
- 1:45
it can also tell us what our models and
- 1:47
what our hardware is not good at yet.
- 1:48
So, it helps us sharpen our research
- 1:50
focus to figure out what is the right
- 1:52
problem to work on in robotics. Because
- 1:55
if you're familiar with the robotics
- 1:57
field, there are too many problems.
- 1:58
There are too many different fields you
- 2:00
can spend your effort on. And by
- 2:02
deploying and by building a product, we
- 2:04
know exactly the right kind of problems
- 2:05
that we need to focus to actually make a
- 2:08
robotics not just a demo or videos you
- 2:10
see on YouTube, but rather in the real
- 2:12
world impacting millions of people's
- 2:14
life.
- 2:16
So, before I jump into what we
- 2:18
specifically work on at covariant, just
- 2:19
a very quick introduction on the kind of
- 2:22
models we're training for robotic
- 2:24
manipulation. So, at a high level, uh we
- 2:27
do data collection where you know,
- 2:28
there's a lot of data being collected.
- 2:30
So, here's a video of me uh manipulating
- 2:33
a robotic hardware to, you know, fold a
- 2:35
t-shirt. And once you collect enough of
- 2:37
this kind of data, you can put all of
- 2:40
them in a large neural network. And a
- 2:42
neural network essentially takes in a,
- 2:45
you know, representation of the world,
- 2:47
in this case just camera feed of what
- 2:49
the, you know, a table looks like, and
- 2:51
the outputs robot's, you know, joint
- 2:53
positions or torque to actually control
- 2:55
the robot to do the task that you have
- 2:57
collected data on.
- 2:59
Right?
- 3:00
So, you know, this is the high level of
- 3:03
how you robot models in the real world
- 3:06
function. So, how are we actually
- 3:08
scaling this up to actually create
- 3:10
models that's generalizable and it can
- 3:12
do a lot of different tasks. So, at
- 3:14
covariant we focus on what we call the
- 3:16
pre-training data pyramid for the real
- 3:17
world, where we gather a lot of diverse
- 3:20
off-robot data, just because robotics
- 3:22
data is very scarce. So, we have data
- 3:24
captured from humans wearing cameras, uh
- 3:27
from public data sets, and these days
- 3:29
we're also considering some simulation
- 3:31
data. But if you only have off-robot
- 3:33
data, it's not actually enough to get
- 3:35
robots to do very precise actions. So,
- 3:38
on the robot themselves, we also collect
- 3:40
diverse tasks collected in many
- 3:41
different scenarios, on many different
- 3:44
types of tasks, like industrial tasks,
- 3:46
household, uh in the laundromat, and uh
- 3:48
you know, hotels. And then finally,
- 3:50
because we're also deploying robots, we
- 3:52
can collect very high-quality deployment
- 3:55
data, which helps the model to close
- 3:57
train and test distribution gap. Because
- 4:00
when you're developing robots, you know,
- 4:02
most of the time your robots are in your
- 4:03
facility, in your laboratories. But if
- 4:05
you're deploying robots, it's in a very
- 4:07
different environment. So, we found that
- 4:09
by combining these three data sources,
- 4:11
we can actually train large-scale
- 4:12
foundation models that work very well in
- 4:14
the real world. And so far, we have more
- 4:16
than 200,000 hours of data in our
- 4:18
training pipeline.
- 4:20
And the model architecture roughly
- 4:22
follows a high-level reasoning model
- 4:25
with a low-level world action model that
- 4:27
can actually output dexterous actions at
- 4:30
fine-grained high frequency, right? So,
- 4:33
this is the kind of architecture we have
- 4:35
because in the real world, if you think
- 4:36
about a robot doing physical tasks, it
- 4:39
needs to have a semantic understanding
- 4:41
of the world, but also needs to
- 4:43
understand physical interaction at a
- 4:45
fine-grained level to be actually able
- 4:47
to, you know, recover from mistakes and
- 4:49
do very precise actions to complete the
- 4:51
task.
- 4:52
So, once you combine,
- 4:54
you know, this architecture with a lot
- 4:56
of data, what happens is that the model
- 4:58
can be rapidly fine-tuned to do a bunch
- 5:00
of different tasks in the real world.
- 5:03
Yeah, so here are just a gallery of the
- 5:05
kind of benchmark tasks we're doing in
- 5:07
the office. So, it spans from, you know,
- 5:09
like cleaning trash, opening a box, to,
- 5:12
you know, things like folding towels,
- 5:14
folding t-shirts, and also just a bunch
- 5:16
of other tasks our researchers have
- 5:17
thought about.
- 5:19
And what's really interesting is that
- 5:20
with a good pre-trained model, even
- 5:23
without any,
- 5:24
you know, for some of these tasks I'm
- 5:26
showing you here, they're not in the
- 5:28
pre-training data at all. But with a
- 5:30
good pre-training, you can actually
- 5:31
post-train the model to do these kind of
- 5:33
highly dexterous tool use tasks with
- 5:36
very little amount of data. So, both of
- 5:38
the videos you're seeing here have only
- 5:40
been trained on with less than 1 hour of
- 5:42
task-specific data. But you can see that
- 5:45
the robot can repetitively
- 5:47
do these tasks over and over without
- 5:49
failure. And I think this is a stepping
- 5:51
stone towards actual commercial grade
- 5:54
real world deployment because in the
- 5:56
real world physical tasks need to be
- 5:58
done by humans and by robots over and
- 6:00
over. And if your models are not
- 6:02
reliable enough, then yes, you can shoot
- 6:04
these kind of pretty demos, but it's
- 6:06
still very far away from actual
- 6:08
deployment.
- 6:09
So that actually brings us to what I
- 6:12
consider the current status quo for
- 6:14
training large-scale foundation models
- 6:16
for manipulation, which is that the
- 6:18
generous models that can do many many
- 6:20
tasks like the ones I've been showing
- 6:22
you. Even though they make pretty
- 6:23
videos, but what I would tell you is
- 6:25
that the success rate is actually
- 6:27
not super high.
- 6:29
You know, when these models are actually
- 6:31
deploying, they're about 80 to 90%
- 6:33
success rate. But if you're stuck at 80
- 6:36
to 90% success rate, then you know, the
- 6:38
chance that you can do the same task 10
- 6:40
times in a row is actually less than
- 6:42
0.1%.
- 6:44
And on on the flip side, in the history
- 6:46
of robotics, we have had very
- 6:48
specialized robots and a very
- 6:49
specialized machine learning pipelines
- 6:52
to do singular tasks like pick and
- 6:54
place. But this kind of pipeline, which
- 6:56
is what I consider specialist models,
- 6:58
aren't aren't very scalable, meaning
- 7:00
that you can't take the same pipeline to
- 7:02
just do a new task and do an arbitrary
- 7:04
task very fast. So our mission is to
- 7:06
resolve the status quo and you know,
- 7:08
this kind of like dichotomy by training
- 7:11
general purpose models that can both do
- 7:13
many many tasks and also be reliable
- 7:15
enough for commercial deployments.
- 7:18
So how do we actually do that? In
- 7:19
today's talk, I'll briefly talk about
- 7:21
some of our progress on mastering very
- 7:23
complex tasks very reliably and also
- 7:26
taking the same skills to be performant
- 7:29
not only in the environments in the
- 7:31
laboratory environments that we're
- 7:32
training, but also being able to deploy
- 7:34
to arbitrary customer sites without
- 7:36
fine-tuning or without fine-tuning
- 7:38
adaptation.
- 7:40
Okay, so let me get to this. Right, So,
- 7:43
the first question we want to answer is
- 7:45
can we take, you know, the general
- 7:47
pre-training and post-training recipe
- 7:48
that we had, but then turn these models
- 7:50
into models that can be almost close to
- 7:53
100% robust on uh any task. So, this is
- 7:57
the first research result we published
- 7:59
last year. Uh for detail, you can check
- 8:01
out our blog post. But, at a high level,
- 8:04
you know, if we want to make a model
- 8:05
100% robust on many, many tasks, I think
- 8:08
it's very good to first narrow down on a
- 8:10
commercial use case that's reasonably
- 8:12
hard, which allows you to make research
- 8:14
progress, but also has commercial value.
- 8:17
So, when we first started the company uh
- 8:19
last year, what we discovered is that if
- 8:21
you go to any, you know, uh restaurant,
- 8:23
you know, uh fancy restaurant or dim sum
- 8:25
places in the US, you'll see nicely
- 8:27
folded napkin on the table, right, for
- 8:29
you, right? If you go to a Cheesecake
- 8:31
Factory, you'll see those napkins. And
- 8:33
what happens is that in the back office
- 8:35
of the restaurant, there's usually uh
- 8:36
workers or, you know, restaurant
- 8:38
employees that's folding these napkins
- 8:40
one by one by hand. And that's a very
- 8:43
mundane process, and a lot of the
- 8:45
customers we have talked to are looking
- 8:47
into robots that can actually do the
- 8:49
same task. So, this is one of the
- 8:51
earliest case study we did on how to
- 8:53
train models to be very robust.
- 8:55
So,
- 8:56
our research result is a model called
- 8:58
the Dyna-1, which is a generalist robot
- 9:00
foundation model that's fine-tuned to do
- 9:02
napkin folding. They can actually
- 9:04
achieve 99.4% success rate over a
- 9:07
24-hour span. So, here is a time lapse
- 9:10
of the model, you know, doing the task.
- 9:12
And uh you know, it's sped up about a
- 9:14
thousand times, so you can see the clock
- 9:16
in the back running very quickly to show
- 9:18
the progress. And uh my favorite part of
- 9:21
the video is when the clock, you know,
- 9:23
hits about like right now, right? Like
- 9:25
7:00 a.m. in the morning, so the lights,
- 9:27
you know, actually come out in the
- 9:29
outside. So, the environments are
- 9:30
actually, you know, shifting over time
- 9:32
due to the lighting as the robot folds
- 9:34
the napkin, but the model's robust
- 9:36
enough, and it just keeps going.
- 9:38
So, how do we actually get to a model
- 9:40
that can do this?
- 9:41
And uh first of all, you know, this uh
- 9:43
video we put out is not a one-time
- 9:46
occurrence. The model can do this many,
- 9:48
many times
- 9:50
uh in the office, you know, so these are
- 9:51
four distinct trials of 24-hour runs.
- 9:55
So, before I dive into the technical
- 9:57
detail, just to highlight how difficult
- 9:59
the task is. So, you see that in the
- 10:01
napkin folding task, you start out with
- 10:03
a stack of napkin on the side, and what
- 10:05
the robot has to do is like be very
- 10:07
precise about picking out exactly one
- 10:09
napkin from the stack, and then fold it,
- 10:11
and then put it into a bin. And then a
- 10:13
lot of the
- 10:14
failure cases just come from the fact
- 10:15
that these, you know, parallel jaw
- 10:17
grippers we're deploying, you know, the
- 10:18
gripper may not be precise enough to be
- 10:20
able to pick out exactly one napkin. And
- 10:23
in these situations, the model has to
- 10:25
learn how to recover from the mistakes
- 10:27
of pulling out extra napkins.
- 10:29
And then secondly, the in commercial
- 10:32
environment, different from a lab demo
- 10:34
where the researchers like me are
- 10:36
thinking of task success, there's
- 10:38
actually very
- 10:40
well-defined success criteria for these
- 10:42
tasks. So, on the right, you see the
- 10:44
difference between what we consider a
- 10:46
grade five fold, which is a fold quality
- 10:48
that the restaurant would accept versus
- 10:51
a grade three, which is something that's
- 10:53
below the acceptance criteria. And what
- 10:55
you see is a barely like 1-in difference
- 10:57
in how, you know, low the you know, the
- 10:59
second fold seam of the napkin is.
- 11:03
So, how do we actually get to a model
- 11:05
that works really well? So, our internal
- 11:07
attempt was kind of like the
- 11:09
pre-training
- 11:10
many, many hours of data and then
- 11:13
post-training recipe that I told you
- 11:14
about in the beginning. And doing this
- 11:16
roughly gets you about 80% success rate.
- 11:19
And what happens is that the model, you
- 11:21
know, is doing fine in the beginning,
- 11:23
but as soon as it makes a mistake, it'll
- 11:25
typically go out of distribution, get
- 11:27
stuck, and unable to recover.
- 11:30
So, we have to do something more than
- 11:31
the standard pre-training and
- 11:32
post-training idea that's very popular
- 11:35
in robotics and also in other fields.
- 11:38
So what we did is that we developed what
- 11:40
we consider
- 11:42
reward models for complex long horizon
- 11:44
manipulation tasks. So these are models
- 11:46
that can, you know, look at a robot
- 11:48
video and accurately score its progress
- 11:51
towards solving a task. All right, so
- 11:53
let me just play these videos again. So
- 11:55
here's the same model, you know, being
- 11:57
able to score how well the robot is
- 11:59
doing these long horizon complex tasks
- 12:02
as it's, you know, going from
- 12:04
you know, starting of the task to
- 12:06
finish. So you see that as the robot's
- 12:08
completing a task, it's able to go from
- 12:10
zero to one. And if you squint at these
- 12:12
videos enough, you also see that
- 12:14
whenever the robot is actually making
- 12:16
some mistakes, there will be like slight
- 12:18
dips in the reward model. And that
- 12:20
actually becomes a very important
- 12:23
insight into how to make these models
- 12:25
very robust. So here's what happens when
- 12:28
you run such model during like a
- 12:30
autonomous run out of the robot. So you
- 12:32
see that when the robot is doing fine
- 12:34
folding napkins, the progress estimation
- 12:37
is roughly monotonic going up, right?
- 12:39
Because the robots are not messing up.
- 12:41
But what's really interesting is is that
- 12:43
let me just for fast forward a bit.
- 12:45
Whenever the model starts to make
- 12:47
mistakes, so here it is. This is the
- 12:50
ninth napkin is folding. So you see that
- 12:52
when the model is like making mistake,
- 12:54
that's when the progress estimation, you
- 12:56
know, starts to like, you know, show
- 12:57
non-monotonic sign indicating that the
- 12:59
robot is messing up. Right? And this is
- 13:01
very important because once your model
- 13:03
is like good enough in the 90% uh
- 13:06
range, then it's very inefficient for
- 13:08
humans to manually oversee the robot to
- 13:11
detect its failure and then try to
- 13:12
recover. But once we have this reward
- 13:15
model, we can actually do what I
- 13:17
consider uh scalable supervision. So you
- 13:19
can just have the robot trying to fold
- 13:21
napkins and then run this reward model
- 13:23
in the background. So whenever a model
- 13:25
does make a mistake, uh we as
- 13:27
researchers or operators can immediately
- 13:30
know the kind of video case that the
- 13:31
model is struggling on and then do very
- 13:34
targeted data collection and error
- 13:36
recovery data for the model, then we can
- 13:38
fine-tune the model again. So, the
- 13:40
overall pipeline looks like a
- 13:42
human-in-the-loop active learning
- 13:44
process where we can use the reward
- 13:46
model to help us catch the the kind of
- 13:48
mistake the model is bad at and then do
- 13:51
targeted collection to make the model
- 13:52
better and iterate.
- 13:54
And what we found is that once you
- 13:55
iterate on this
- 13:57
uh couple cycles, then you start to get
- 14:00
a model that's extremely robust and can
- 14:02
recover from all kinds of errors and
- 14:04
finally bringing us closer to a
- 14:06
commercial grade robots. So, here is
- 14:08
just uh some of the
- 14:10
uh
- 14:11
you know
- 14:12
uh
- 14:13
highlights, I guess, during the 24-hour
- 14:15
trial. So, what you saw there was the
- 14:17
robot accidentally picked out more than
- 14:19
one napkins and the model is able to,
- 14:21
you know, separate the napkins and uh
- 14:24
you know, here it's kind of doing that
- 14:26
and uh
- 14:26
be able to continue progressing. And
- 14:30
what we found very interesting is that
- 14:32
uh
- 14:33
napkin folding is a deformable object
- 14:34
manipulation task, right? So, there is
- 14:36
almost infinitely many possible states
- 14:39
or configuration that a napkin can get
- 14:41
to. So, it's impossible to exhaustively
- 14:43
collect data for all the error cases.
- 14:45
But once we had done the active learning
- 14:47
many, many times, we saw the model able
- 14:49
to generalize to new ways of recovery
- 14:52
from the mistakes made and continue to
- 14:53
make progress. And that contributed to
- 14:56
its ability to be able to uh fold that
- 14:59
uh 99.4% success rate. So, here's uh
- 15:02
what I found the most impressive bit
- 15:04
from the trial. So, typically, if you
- 15:07
look at robot videos, you know, they
- 15:08
only show you the successful cases. But
- 15:10
here, I wanted to highlight that even
- 15:12
where a model accidentally pulled the
- 15:14
entire napkins stack over, it's able to
- 15:16
demonstrate this kind of error recovery
- 15:18
behavior. That was very surprising to us
- 15:20
when we were developing the model and it
- 15:22
contribute to how it's able to just
- 15:23
continuously run, right? So, here
- 15:27
he made a
- 15:28
big mess, but he's able to just do all
- 15:30
kind of like very impressive like
- 15:32
pulling you know, stretching behavior to
- 15:35
recover from his mistake and then
- 15:36
continuously going.
- 15:39
Yeah, so you know, after developing such
- 15:41
technology, we actually are successful
- 15:44
at deploying our models at many many
- 15:46
restaurants in the US. So, here's a real
- 15:49
restaurant deployment of our robot
- 15:51
folding napkins for the customers in
- 15:53
their back office.
- 15:56
And in addition to folding napkin,
- 15:58
because the recipe is quite
- 15:59
generalizable, now we also have models
- 16:01
doing bunch of commercial tasks. So,
- 16:03
here's at a real laundromat in
- 16:06
Sacramento, and this time we fine-tune
- 16:09
our model to do
- 16:10
towel folding for the customer. So, you
- 16:13
can see you know, there's a restaurant I
- 16:15
guess a laundromat worker coming here to
- 16:17
fill the bin over and over, and the
- 16:19
robot just kept folding stacks of towels
- 16:22
to serve the customer.
- 16:27
So, now let's talk about we have a
- 16:29
recipe to master very complex tasks.
- 16:32
But, the caveat here is that in all the
- 16:34
videos I've shown you so far, we have
- 16:35
also collected data at the exact
- 16:38
customer site where location that the
- 16:40
robot is deployed. But, if you think
- 16:41
about scaling robots to any task or to
- 16:45
any customer site, then it'll be much
- 16:47
better or more ideal if the models can
- 16:49
readily generalize the environments they
- 16:51
haven't seen before. So, this is what we
- 16:54
have worked on in the
- 16:56
I guess the end of last year where we
- 16:58
figured out a data recipe to collect a
- 17:00
lot of diverse tasks to allow the robots
- 17:02
to be able to deploy at a new site
- 17:04
without any additional data while
- 17:06
maintaining the task performance. So,
- 17:09
here's a demo we did at Coral 2025. So,
- 17:12
Coral is the premier academic conference
- 17:14
on robot learning, and it's held in
- 17:16
Korea last year. So, you know, bringing
- 17:19
a robot from the US to Korea was its
- 17:21
whole challenge that I can talk about
- 17:23
offline. But, the recipe we discovered
- 17:26
was able to, you know, we just brought
- 17:28
the robot to our, you know, exhibition
- 17:30
booth and just dropped it there, and the
- 17:32
robot can start folding many, many
- 17:34
different t-shirts. And you see that the
- 17:36
robot is facing the, you know, the
- 17:38
conference attendees. So, many, many
- 17:40
times there were people just
- 17:42
deliberately trying to mess with the
- 17:43
robot, use the t-shirt to cover up the
- 17:45
robot's camera, then, you know, all kind
- 17:47
of fancy stuff that you see at academic
- 17:49
conferences. But, the model is able to
- 17:51
continuously fold t-shirts over and over
- 17:53
for 3 days straight at the conference
- 17:55
just to demonstrate the ability to, you
- 17:58
know, solve this task at environments
- 18:00
it's never seen before, bringing it, you
- 18:02
know, much closer to the kind of, you
- 18:04
know, ideal, you know, go-to-market that
- 18:07
you want to have for robotics company,
- 18:09
which is just putting the robot to a
- 18:10
site and it just starts working.
- 18:13
And recently, we have also ventured into
- 18:16
many different tasks. Like, we have a a
- 18:18
partnership with Red Bull where we're
- 18:20
opening Red Bull cans at, you know, Red
- 18:23
Bull events. So, here's a video of our
- 18:28
uh
- 18:28
robot. Again, uh I guess this video
- 18:31
won't play properly. So, but you get the
- 18:33
point. We brought the robot to a Red
- 18:35
Bull event and it's able to
- 18:37
uh open uh these Red Bull drinks for,
- 18:40
again, conference attendees over and
- 18:42
over without failure, even though in
- 18:44
this particular deployment, it's at a
- 18:46
music festival, so the lighting's always
- 18:48
changing in the background, but the
- 18:49
model can continue. Yeah, but due to
- 18:52
yeah, I guess the
- 18:54
video will not play.
- 18:56
So, just to summarize, uh our mission at
- 18:59
AINA is to be able to build
- 19:01
general-purpose models and robots that's
- 19:03
both competent at many tasks, but also
- 19:06
being able to focus and attain really
- 19:08
high performance on commercial tasks.
- 19:11
And I have demonstrated some of our
- 19:13
recent progress on mastering complex
- 19:15
tasks and also generalizing the skills
- 19:17
to environments is the robots have never
- 19:19
seen before. So, we are a series A stage
- 19:22
company actively growing and hiring and
- 19:24
if you're also interested in partnering
- 19:26
with us for deployments and other
- 19:28
things, feel free to reach out and then
- 19:30
talk to me after the talk and thank you
- 19:32
for listening.
- 19:33
>> Awesome. Thank you, Jason. Really
- 19:34
appreciate.
- 19:36
If you guys haven't seen the data robot
- 19:37
in life in real life, have you check it
- 19:39
out? I will say see as very impressive
- 19:41
performance.
- 19:43
Do we have any question or want to ask?
- 19:46
Okay, cool.
- 19:51
>> Thank you for sharing the presentation.
- 19:54
I wanted to ask if you are already doing
- 19:57
kind of like
- 19:58
providing developer kits with robots,
- 20:01
SDKs, frameworks, etc. for your
- 20:03
commercial partners?
- 20:05
>> Yeah, so we haven't been working on
- 20:07
develop developer kit, but you know,
- 20:10
at the current moment we're building our
- 20:12
own hardware stack. So, maybe at some
- 20:15
point it's on our road map, but
- 20:17
not as of not as of now. Yeah.
- 20:26
>> I'm curious if you have ventured into
- 20:28
education like teaching kids
- 20:32
elementary maths or any such use case
- 20:35
you tried.
- 20:36
>> Yeah, so we haven't looked into
- 20:37
education use case, right?
- 20:40
But I think in the future, you know,
- 20:41
once our
- 20:43
full stack robotics pipelines mature,
- 20:45
once we have our own hardware, our data
- 20:46
collection, you know, toolkit, I think
- 20:49
it's possible and I think it'll be very
- 20:50
interesting to venture into education
- 20:53
use cases because I also think robotics
- 20:55
will only get bigger in the future. So,
- 20:57
I think it'll be really really
- 20:58
interesting to get children and young
- 21:01
kids into the field and the learning how
- 21:03
to actually train models to do tasks.
- 21:07
>> Thank you for the presentation, really
- 21:08
good. Uh
- 21:09
my question is like is there do you guys
- 21:11
have a plan for uh kind of taking this
- 21:14
technology direct to consumer? Like I
- 21:16
saw the use case for the laundry folding
- 21:19
laundry.
- 21:20
A lot of people don't like folding
- 21:21
laundry. I think this is a great use
- 21:23
case. So, is that something that you
- 21:25
guys thinking about? And obviously
- 21:27
there's a cost and all those aspects,
- 21:29
but what's what's your plan in the long
- 21:31
term on this technology?
- 21:33
>> Yeah, so our long-term plan is to
- 21:35
develop, you know, like a robot model
- 21:37
plus hardware platform that can be
- 21:39
deployed anywhere for anyone. So, that
- 21:40
would include, you know, like going
- 21:42
directly to consumers. And you know, our
- 21:44
models today are able to just fold, you
- 21:46
know, uh many kinds of garments. But uh
- 21:49
in terms of like go to market strategy,
- 21:51
our current focus is on enterprise use
- 21:53
cases because I think the distribution
- 21:55
channel is like I think it's easier to
- 21:58
uh work with uh
- 22:00
enterprise customers. Like we can deploy
- 22:02
right away and iterate at customer
- 22:04
sites. But for consumer product, I think
- 22:07
uh I would imagine that it has to be
- 22:09
very very polished and it's less
- 22:10
tolerant for mistakes. And there is a
- 22:12
lot of privacy safety concerns that we
- 22:14
think uh we are
- 22:17
trying to not get into at the current
- 22:19
moment to
- 22:20
uh unblock us from deployments because
- 22:23
we think that the bottleneck for AI
- 22:25
robot is getting AI and hardware
- 22:28
co-working together very very well. And
- 22:29
I think commercial environments provide
- 22:31
ideal testing ground uh for companies
- 22:34
and for the entire field at this stage.
- 22:38
>> Thank you, Jason.
- 22:40
Uh my name is Ahmed and I work in a
- 22:43
in voice AI in in the voice space. And
- 22:46
at least for consumer robots, uh
- 22:49
where I believe humans will want to
- 22:51
issue voice commands to robots, uh can
- 22:54
you talk about how to integrate uh
- 22:57
traditional voice stacks, you know,
- 22:59
which are intent-driven and uh uh
- 23:01
different kinds of models from uh uh
- 23:04
robotic, you know, VLA models.
- 23:06
>> Yeah, that's a great question. So, we
- 23:08
think uh robot human interactions really
- 23:11
important, and it at the end of the day,
- 23:13
right? We want to develop models that
- 23:14
are teachable. So, it'd be really nice
- 23:16
if humans can speak to it, and the robot
- 23:18
does the task. So, the way I think about
- 23:20
it, so I'm not I don't have a speech or
- 23:22
audio background, so what I can think of
- 23:23
is like there is very mature
- 23:25
speech-to-text models, right? So, you
- 23:27
can use that running the on-board to
- 23:30
translate human intent or human speech
- 23:33
into text very quickly, and then, you
- 23:36
know, that's where our reasoning model
- 23:38
and our world model comes into play,
- 23:39
because both of them are able to
- 23:41
interpret text commands. So, that will
- 23:43
allow the model to translate instruction
- 23:46
into actions. But, I think in the
- 23:49
overall stack, the bottleneck is still
- 23:51
on the robot foundation models, because
- 23:53
models today our models and other
- 23:56
people's model are not ready to just
- 23:57
execute any arbitrary commands. So, I
- 23:59
think there's still a gap in terms of
- 24:01
low-level robotic manipulation to get
- 24:03
there. But, I think that's the eventual
- 24:05
goal that we want to get to. Just a
- 24:07
model that can be steerable, that can be
- 24:09
taught by anyone. Yeah.
- 24:12
Okay. Thank you.
- 24:14
>> Um how do you think about, I guess, the
- 24:16
split of intelligence between
- 24:19
the general foundation models that you
- 24:21
all are training, I guess, something
- 24:23
like a task-specific model for something
- 24:24
like laundry folding, and then kind of
- 24:26
some of the on-site specific training
- 24:29
that needs to happen for a specific
- 24:30
deployment in a particular environment?
- 24:33
>> Yeah, I think both are very very
- 24:36
important, right? So, the kind of recipe
- 24:38
that we have is, you know, we have
- 24:39
pre-training, post-training, then we
- 24:41
have some of the sort of active
- 24:43
learning, right? I I think
- 24:45
uh
- 24:45
if you think about, you know, humans,
- 24:47
right? You know, we have like basic
- 24:49
level of like semantic and physical
- 24:51
understanding that allows us to adapt
- 24:53
any physical task very fast, and I think
- 24:55
the best way to even build commercial
- 24:57
grade robots today is by doing that.
- 25:01
Because by doing a lot of generalist
- 25:03
model training have a general
- 25:04
understanding of physical world that's
- 25:06
very useful. And what we have seen, so
- 25:08
this is something I didn't talk about in
- 25:10
the talk is that when the model was
- 25:12
recovering from all kind of errors that
- 25:14
you saw, a lot of that also just came
- 25:16
from interpolating different error
- 25:18
recovery behavior that is gathered from
- 25:20
doing other tasks, right? Just because,
- 25:22
you know, the model has seen thousands
- 25:23
of tasks we had some data to recover
- 25:26
from all kinds of mistakes is able to
- 25:29
just execute that kind of like uh on
- 25:32
demand on a new task it hasn't seen
- 25:33
before. But if you're only specializing
- 25:35
a model from the get-go on one task,
- 25:37
then it's actually missing out on a lot
- 25:39
of the physical knowledge that allows
- 25:40
the model to be more robust than if you
- 25:43
only train on one task. And that's what
- 25:44
we have seen consistently. That's why,
- 25:47
you know, in the beginning of the talk I
- 25:48
also emphasize on having that general
- 25:50
pre-training backbone and then doing
- 25:52
post-training and active learning on top
- 25:54
of that to get to actual commercial
- 25:55
grade usability, not on just one task,
- 25:58
but also the same recipe that can be
- 26:00
repeated across many many tasks.
- 26:03
>> Awesome. Thank you, Jason. Really
- 26:04
appreciate. Um that's a wrap for uh data
- 26:07
session. Uh if you have more questions,
- 26:09
feel free to catch Jason after the
- 26:10
session as well. Uh but yeah, this wrap
- 26:13
our morning session. Uh this afternoon
- 26:15
we have Unity, Skydio, Zoox, uh Waymo,
- 26:17
DeepMind. Um so we'll catch you and see
- 26:19
you AI as well. So we'll see you in the
- 26:21
afternoon. Thank you, everyone.
- 26:39
>> [music]