AI Engineer World's Fair 2026
Robotics Has Been Stuck for 70 Years — Deepak Pathak, Skild AI
Read the talk
Robotics Has Been Stuck for 70 Years
Deepak Pathak explains Skild AI’s approach to a general robot brain: combine complementary data sources, share learning across different bodies, and use deployment to improve the model. The demonstrations reveal why precise grasps, unfamiliar stairs, and damaged hardware test more than a robot’s appearance.
From a talk by Deepak Pathak
At a glance
Ideas worth remembering
Judge robot data by scalability, environmental diversity, and proximity to real robot action. Large quantities of repeated experience in one setup do not satisfy all three.
Skild’s recipe assigns complementary roles to simulation and human-video pre-training, teleoperation post-training, and deployment experience returned to training.
Hardware constraints shape the required intelligence: a parallel-jaw gripper must choose an earbud grasp that preserves the orientation needed for insertion.
Visually guided stairs can demand more than spectacular body maneuvers because the robot must connect unfamiliar environmental geometry to its actions.
Supporting different bodies may also support recovery after damage. The disabled-leg and jammed-wheel examples make that benefit concrete without establishing a general safety guarantee.
An impressive demo is an old achievement
A robot looks at a picture of blocks and arranges physical blocks to match it. The task requires connecting a two-dimensional image to three-dimensional objects, then moving those objects precisely. In Deepak Pathak’s opening example, that apparently modern demonstration comes from the 1960s: the MIT copy demo. Pathak, Skild AI’s co-founder and CEO and a Carnegie Mellon professor, uses it to challenge the excitement surrounding robotics as AI’s next frontier. 1:27
The next clip reaches further back. A human controls robot manipulators through a leader–follower system in 1957. Its basic principle survives in modern teleoperation: a person supplies the motion, and a robot follows. Later examples include juggling, foosball, and a Berkeley robot that clears a table and packs objects into a box, which Pathak describes as the work of one graduate student with a single GPU machine. His dinner wager is deliberately provocative: distinguish old robotics videos from today’s.
The historical comparison supports Pathak’s diagnosis: robotics has concentrated on building particular machines without developing a general brain. That is a claim about the field’s direction, rather than proof that nothing improved. The useful distinction is between making one impressive behavior work and building intelligence that transfers to different tasks and situations.
Moravec’s paradox gives the problem its memorable phrasing: “Hard is easy, easy is hard.” Mathematical competition and chess look intellectually demanding to humans; climbing stairs and picking up chess pieces feel ordinary. Yet success at the intellectual task does not settle the physical one. Winning a board game leaves the problem of moving its pieces on an unfamiliar board.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
No internet of robot data
Large datasets and large models have been a productive recipe elsewhere in AI. Robotics lacks an equivalent internet of action data. Teleoperation supplies useful examples, but each example requires a physical robot and a human’s time. Pathak estimates roughly one minute per example. More operators can increase collection, but they do not remove its labor and hardware costs.
Skild AI’s proposed answer is an omni-bodied intelligence: “any robot, any task, one brain.” A humanoid, quadruped, conveyor-belt arm, and dexterous hand should all contribute to the same learning system. The motivation is data scarcity. Restricting training to one machine would discard experience from other hardware, tasks, and environments that the model needs. 7:22
The teaser presents robots from different companies running the Skild brain, including locomotion in new scenarios and responses to disturbances. Its intended relationship is reciprocal: shared intelligence controls many bodies, and experience from those bodies improves shared intelligence. The demonstrations and deployment results throughout the talk are Pathak’s reports; they do not establish general success rates across arbitrary robots or environments.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Three properties of useful robot data
Pathak’s retrospective moves through several attempts to collect robot experience. Each solves part of the problem and leaves another part exposed. 9:31
- Curiosity-driven exploration: Robots explore and collect their own “play data,” including forces and joint angles. This removes the need for a human to guide every action, but collection still proceeds through physical machines in the real world. Hardware cost and physical time constrain scale.
- Teleoperation: Human-guided trajectories supply data closely tied to robot action. Collection requires both robots and operators, and thousands of examples in one setup can still offer little environmental diversity. Moving the robot to a different home every day would improve coverage while making collection harder.
- Simulation: Deep reinforcement learning can train in simulation and transfer to a real robot. Simulated experience is easier to multiply, but someone must build the scenes. More trials within engineered scenes do not automatically create more kinds of environments.
- Human video: Videos offer scalable, diverse demonstrations of people doing things. Their weakness is distance from the robot’s action space: observing a human movement does not directly supply the joint commands a different body needs.
The evaluation has three axes: scalability, diversity, and closeness to the robot’s own joint-angle ground truth. Repeating a task in one room can produce a large dataset without teaching the robot much about a different room.
There is no single winning source. The useful discovery is that their weaknesses differ. Human video contributes variety that carefully engineered simulation struggles to supply. Teleoperation contributes direct robot experience that video lacks. Combining them gives each source a job instead of asking one collection method to satisfy every requirement.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Pre-training creates the base; deployment feeds it
The training recipe separates breadth from precision. Pre-training uses scalable sources such as simulation and human videos, accepting their imperfect relationship to real robot action. Post-training uses smaller amounts of high-quality teleoperation data. The analogy to language models is functional: broad initial learning gives later, narrower training something useful to adapt.
Deployment adds a third source. Once robots operate at scale, their experience can feed subsequent training. Pathak expects this data eventually to outweigh the initial sources, but that expectation depends on achieving deployment scale. A model restricted to one robot shape or hardware version would fragment the experience as machines change. An omni-bodied model is intended to keep those different deployments useful to one shared brain.
How does experience become reusable across the training stages? The diagram follows the proposed loop. Simulation and video supply the starting breadth; teleoperation helps turn that base into deployable behavior; operational experience returns to training. Supporting many bodies matters at the return path, where otherwise incompatible deployments must contribute to the same model.
Scalable experience with complementary weaknesses
The proposed flywheel combines broad pre-training, precise post-training, and experience from deployments across changing robot bodies.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The insertion starts before the gripper closes
The demonstrations begin by reconsidering what makes a manipulation task difficult. Laundry folding has a reputation for difficulty because fabric is awkward to model with hand-built physics. But a learned policy faces a different challenge: how much execution error can the task tolerate? A hand that lands a centimeter away from its intended position may still produce an acceptable fold. Pathak uses that tolerance to explain why laundry can be forgiving for learning-based control, despite being difficult for classical modeling.
Putting an AirPod-shaped earbud into its case exposes a tighter constraint. The robot has a parallel-jaw gripper, whose jaws close around the object without the in-hand manipulation available to fingers. The arm must approach and grasp the earbud in an orientation that permits the later insertion. A convenient pickup can leave the object inconveniently positioned for the next step.
The observable change is from a loose earbud to one seated in its case. Its causal sequence starts earlier: orient the arm, close the jaws around a usable grasp, carry the held object to the case, then align and insert it. The props are imitation earbuds costing roughly $5–$10 each, without magnets to help pull them into place. That detail removes a source of mechanical assistance: the robot must achieve the fit itself.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Human demonstrations, imagined scenarios, and omelets
The next learning example uses egocentric human video followed by less than one hour of additional robot data to transfer behavior to humanoids. Third-person video remains a direction under investigation in this account; successful transfer would open access to a wider range of ordinary videos. The important bridge is from watching a human do something to making a different body do it, with limited robot-specific experience.
Those humanoids have hands, but Skild is not deploying hands in factories at the time described. Pathak’s objection is mechanical reliability. A more capable brain can make a simple gripper useful, while adding fingers does not solve the problem of fragile hardware. His challenge to anyone with a better hand is practical: he would buy ten units to test.
To explain learning from little robot data, Pathak introduces the robot’s “dream”: generated scenarios inside its own model. The displayed video represents imagined experience rather than a recording of physical trials. The proposed mechanism multiplies learning across scenarios without physically collecting every one. The talk does not specify the generation objective or how imagined outcomes are checked against real physics, so this remains a high-level account of the data expansion.
Omelet making then tests the approach on a roughly $4,000 setup. Its only sensor is a camera; it has no force sensing. Contact tasks therefore depend on visual information to guide actions that apply force. Pathak does not recommend avoiding better sensors. His point is that existing hardware can support more capable behavior than its current software extracts.
The cooking behavior is described as fully end to end, without a hand-written state machine deciding when the sequence starts or ends. An empty plate in front of the robot provides visual context for making an egg. Objects can move or be replaced, and the robot continues; the pan and gas stove remain the same while other objects differ. Pathak reports less than ten hours of training data for this example and attributes the resulting tolerance to the base model. The fixed pan and stove make the demonstrated scope concrete.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Why stairs demand more than a backflip
End to end has a specific meaning here: camera observations enter the brain, and its output drives motor power. Pathak contrasts this with a traditional pipeline containing explicit mapping and planning stages. He also mentions a final control component bridging 100 Hz to 500 Hz; the captions render its name as “P,” leaving its technical identity unclear. The supported distinction is direct learned control without an explicit map-and-plan pipeline, rather than a claim that no lower-level control exists. 20:41
Locomotion returns to Moravec’s paradox. Backflips and dancing look spectacular, but Pathak describes them as comparatively easy because the robot primarily needs to control its own known body. Stair climbing adds unfamiliar geometry. The controller must see the steps, respond to their height and width, and adjust to disturbances. Knowing the body is insufficient when the next foothold depends on the surroundings.
What changes in the control problem when a robot encounters stairs? The comparison below makes the extra information visible. A familiar body supports rehearsal of a body-centered maneuver. An unfamiliar staircase requires visual information about the environment to affect action. That connection explains why an ordinary-looking ascent can test more general capability than a dramatic flip.
The stair examples use a torso camera without constructing a three-dimensional map of the surroundings. The robot intentionally steps over obstacles, traverses different stairs including fire escapes, and responds to being pulled while climbing. Pathak identifies the model as the same across these scenarios. He then shows parkour and reports that the displayed behavior takes half an hour to train, while the earlier visually guided locomotion takes much longer. The training-time comparison reinforces his warning: spectacle is a poor shortcut for judging difficulty.
Body-centered motion can be simulated
Pathak’s comparison concerns the information required for control: stairs add unfamiliar environmental geometry and disturbances.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The flywheel needs working deployments—and bodies that can change
The final stretch moves from demonstrations to deployment. After an example of LAN-cable insertion, Pathak identifies deployment itself as the hardest part of starting the data flywheel. Physical operations introduce problems beyond intelligence and hardware, so waiting for a perfect robot brain would also delay collecting the experience needed to improve it.
- GPU assembly: Pathak describes work with NVIDIA for its Houston factory and reports that the assembly system is already deployed at the time of the talk. The example combines precise manipulation with a randomized, noisy setup. Its practical aim is to operate on existing factory lines amid the mess that accompanies human work. Tolerance of disturbances alone, however, does not establish that operation alongside people is safe without guarding.
- Package delivery: The robot travels from a truck to a house’s front door. Reaching the destination requires identifying which part of the scene is the front door, so the task combines movement with an understanding of the destination. Pathak reports deliveries with partners and says the same setup also works in warehouses.
The ending gives omni-bodied learning another purpose: adaptation when the body changes unexpectedly. A robot with a broken leg has effectively become a different robot. A brain trained to support different bodies may therefore handle damage as a change in embodiment rather than assuming that its original hardware remains intact. Pathak presents transfers to previously untrained robots and recovery after a leg failure as examples of that capability.
In one example, disabling legs leaves the robot learning to walk on two legs in three trials. In another, jamming a wheel causes the robot to start walking. Pathak describes adaptation ranging from milliseconds to roughly 30 seconds across these examples, and reports a further experiment in which two halves of a robot can operate separately. These are distinct forms of recovery, rather than one universal recovery time. Pathak presents continued operation after damage as a potential safety benefit, but these adaptation examples do not establish that continued motion improves safety or provides a complete safety guarantee.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
From the talk
- Deepak Pathak’s research profileReference
A starting point for investigating the research behind the retrospective on exploration, simulation, and learning from human video.
Return to the explanation of how complementary data sources enter different training stages and how deployment experience feeds back into the shared brain.
Watch the gripper example alongside the explanation of why the initial grasp must preserve the orientation needed for insertion.
Follow the closing examples of changed robot bodies, disabled legs, and a jammed wheel that prompts a switch to walking.
Read the complete timestamped transcript
- 0:01
[music]
- 0:12
I think uh all of you know here that how
- 0:15
much progress has AI made in the last
- 0:16
five years. Uh seems like it's more than
- 0:19
uh last 100 years of progress. And what
- 0:23
is happening especially in AI is AI came
- 0:25
for language then speech audio video and
- 0:28
everybody's excitement is next is
- 0:30
robotics okay and this excitement has is
- 0:33
going on peak uh this year uh for some
- 0:37
reason which is still beyond my uh
- 0:39
understanding and you can also see uh
- 0:42
leaders like Jensen talking about
- 0:44
physical AI as the next frontier of
- 0:45
robotics. So it seems like if you open
- 0:48
Twitter or LinkedIn, it seems like
- 0:49
robots robotics is already here and
- 0:51
we'll have home robots in our homes
- 0:53
sometime by December uh as uh as it's
- 0:56
being said by multiple people. But I'm
- 0:59
here to highlight that robotics has been
- 1:02
almost here for the last 70 years
- 1:06
[sighs]
- 1:06
now. Those of you who got into robotics
- 1:09
in the last four or five years, go and
- 1:12
take any any any course in robotics and
- 1:15
if you can distinguish the videos in
- 1:17
robotics from 30 years ago versus today,
- 1:20
uh I owe you a dinner. So to give you an
- 1:24
example, let's look at this robot. So
- 1:26
this robot here is an image and the goal
- 1:28
of this robot is to look at this image
- 1:30
of blocks and arrange these blocks in
- 1:33
front of it so that it looks the same
- 1:36
pattern as in the image. Okay, this is
- 1:38
an extremely simple task. But if you
- 1:40
think about this, it's a 3D block. You
- 1:42
have to look at understand 3D from each
- 1:43
side. And the robot can do this pretty
- 1:46
precisely. Any guesses how old this
- 1:48
result is?
- 1:51
Anyone? It's a guess. You can make any
- 1:52
guess. Huh?
- 1:54
>> 40 years. Okay, that's the highest.
- 1:57
[panting]
- 1:58
This is from 1960s. This is way before
- 2:02
there were computers. Uh, okay. So, this
- 2:04
is this is called MIT copy demo. Uh,
- 2:06
this was the start of uh this is this
- 2:09
predates AI. Uh, as you know, today look
- 2:12
at this one.
- 2:17
>> One exhibit at the Nuclear Congress in
- 2:19
Philadelphia has the perfect formula for
- 2:22
>> what's happening here. The guy behind
- 2:23
the scene is controlling these robots, a
- 2:26
leader follower system. And if you look
- 2:28
at any big lab, any big company, any big
- 2:31
academic lab, how they get data, they
- 2:34
use teley operation system, which
- 2:35
follows the exact same principles.
- 2:38
Any guesses for the year for this one?
- 2:42
>> Now, everybody would correct
- 2:44
aggressively on the year, but this is
- 2:46
from 1957, 68 years ago. Okay? Ask
- 2:51
anybody in your family who is older than
- 2:53
actually they may not be around like so
- 2:56
because this is for 60 you have to ask
- 2:58
somebody who is 80 year old what was the
- 3:00
state of technology at that time like
- 3:02
RAM for few KB of RAM you would have a
- 3:05
size of a room uh like this big of a
- 3:07
setup and this is from that time when
- 3:09
people could make it work on not just
- 3:12
this you keep coming back every decade
- 3:14
the videos look as cool like here is
- 3:16
robot a humanoid juggling playing
- 3:19
foosball uh with human and this is also
- 3:23
from you can see this is all they were
- 3:24
all predate deep learning this from 8
- 3:26
years ago so this robot is a result from
- 3:28
Berkeley can clean up the whole table
- 3:30
can arrange items in a in a in a box
- 3:33
this was done by one grad student with a
- 3:35
single GPU machine nothing more than
- 3:37
that okay now this picture of and if I
- 3:40
colorize all these videos from past and
- 3:42
apply genai filter for the modern voice
- 3:45
you cannot tell that this [laughter]
- 3:47
video is from 19657 or this is from
- 3:49
today And these are not even the oldest
- 3:51
results the oldest ones go to 1940s.
- 3:54
Okay. So what is this that if you look
- 3:58
at last 70 years tech every other
- 4:00
technology every other technology has
- 4:03
come way far language understanding
- 4:05
computer vision mobile phone chips why
- 4:08
is robotics is stuck in this primitive
- 4:10
age uh for for this long and the reason
- 4:12
is uh a general brain. Robotics has
- 4:15
always been approached as a hardware
- 4:17
problem from ground up and this the
- 4:19
reason that everything around robotics
- 4:21
has progressed but robotics is still
- 4:23
stuck in the same land from the last 70
- 4:25
years. So there's a very famous paradox
- 4:26
called Morave paradox. I don't know if
- 4:28
you know about Moravec. He was also a
- 4:30
CMU professor and I'm also CMU professor
- 4:32
so have to quote him for sure. He was
- 4:34
one of the founding figures in in AI.
- 4:36
After 30 years of working in robotics he
- 4:39
arrived at a very simple conclusion.
- 4:41
Hard is easy, easy is hard. Okay.
- 4:44
Whatever human believe to be hard is
- 4:46
extremely easy for computers and vice
- 4:48
versa. And you can see it happening
- 4:50
right in front of your eyes. You would
- 4:51
say doing math is hard. Math Olympiad
- 4:54
gold medal is really really hard.
- 4:56
Climbing a stair oh super easy. Now look
- 4:59
around you in technology where we are in
- 5:01
terms of what has been solved and what's
- 5:03
not being solved. So this is robotics
- 5:05
for you. Okay. It's very it is not yet
- 5:07
another application of AI. It is what AI
- 5:10
was founded for in the very beginning
- 5:12
and has made very little progress
- 5:13
towards. This is not yet another
- 5:15
application where deep networks can come
- 5:16
in and attack the field. This is has to
- 5:19
be thought of with fundamentally first
- 5:21
principles from the ground up. This is a
- 5:23
very famous example of where you can
- 5:25
have a you know a computer beat uh Gary
- 5:28
Casper on chess in '90s but you can you
- 5:31
still don't have a a a computer or robot
- 5:33
that can pick up the chess pieces and
- 5:35
arrange them on any chess board. Now,
- 5:40
so how do we go about solving this?
- 5:41
Okay,
- 5:43
so on a more positive note, uh we have a
- 5:47
magic sauce for AI success, right? You
- 5:49
get big data set, you train big models
- 5:52
and magic happens. Okay, now we know
- 5:54
this template is working well very well
- 5:57
in many topics. Okay, now can we apply
- 6:01
the same recipe to robotics? Well, it's
- 6:03
not uh directly applicable because we
- 6:06
have no data. There is no internet of
- 6:08
robotics data and as I said earlier you
- 6:10
can go and collect data uh manually on
- 6:12
the robot called telly operation and as
- 6:15
you notice teley operation is not 5 year
- 6:17
old thing this is a 68 or 70 year old
- 6:20
thing so the the funny part is if you go
- 6:24
back and you look at the whole evolution
- 6:25
of GPD3 GPD4 models go back to GPD3 uh
- 6:30
three years ago seems like a like a
- 6:32
lifetime ago GPD3 only began working and
- 6:35
caught people's attention when you could
- 6:37
train those models on trillions of
- 6:39
tokens. So GPD3 was already trained 30
- 6:41
trillion plus token and today hundreds
- 6:43
of trillions of tokens and if you
- 6:45
collect data by manually by tell
- 6:46
operation it takes you about 1 minute to
- 6:48
get one example. You can do the math if
- 6:51
you hire all of US population it will
- 6:53
take you more than a century to reach
- 6:54
the same scale as of GPD3 uh which is
- 6:57
like uh and nobody uses GP3 today. Okay.
- 7:01
So this is very slow and expensive. So
- 7:02
in robotics I would argue nobody has
- 7:06
really scaled robotics yet and we are
- 7:08
very far from talking about scale in
- 7:10
robotics with the way other uh other
- 7:12
areas have seen scale. Okay so this is
- 7:15
where uh we have been focusing on uh our
- 7:17
scaled is about uh 3 years old but uh I
- 7:20
have been working on the problem for
- 7:21
more than a decade uh that's all I've
- 7:23
done in my career uh nothing else. Uh so
- 7:26
at scale what our thesis is is to build
- 7:28
what we call an omniodied intelligence
- 7:32
any robot any task one brain okay any
- 7:36
robot it can be a humanoid it can be a
- 7:38
quadriped it can be a robotic arm on a
- 7:40
conveyor belt or a dextrous hand doesn't
- 7:41
really matter this even this hypothesis
- 7:44
is way more general than uh one would
- 7:47
argue humans are because we control our
- 7:49
own body and why do we have to go so
- 7:51
general well the argument is in robotics
- 7:54
there there is no data anyway. So I
- 7:56
can't pick and choose which hardware do
- 7:58
I use data from. And we should be able
- 8:00
to use data from any kind of hardware,
- 8:02
any kind of task, any kind of scenario.
- 8:04
And that's the only way to truly achieve
- 8:06
the scale of uh of what language models
- 8:08
achieved uh three years ago. Okay. So
- 8:10
the goal here is this uh this whole idea
- 8:13
of any robot any task one brain. And
- 8:15
through this talk I'll hopefully
- 8:16
convince you why this is the this is the
- 8:18
way to go towards robotics. Okay. So
- 8:20
this is the uh rough intro but before I
- 8:23
go into the more details let me show you
- 8:25
just a teaser result.
- 8:28
So in this result, every single robot uh
- 8:31
it's from a different company, different
- 8:32
hardware and they're all controlled by
- 8:35
uh skilled brain whether it's humanoids
- 8:38
going up and down stairs, any kind of
- 8:40
scenarios.
- 8:42
And these are not new results. They're
- 8:44
like a couple of year old uh results in
- 8:46
here. Robust to dust disturbances.
- 8:51
You can put them zero shot in new
- 8:52
scenarios.
- 9:02
So the idea here any every single robot
- 9:06
in this video takes any action anywhere
- 9:09
in the world the brain behind the scene
- 9:12
[music] improves because it's an
- 9:14
omni-body brain.
- 9:15
>> Yeah, this is really hard right because
- 9:16
like eggs are okay.
- 9:32
So that's a teaser of what we what we
- 9:34
work on. So I want to make this talk
- 9:35
more scientific and more uh more
- 9:37
informative than a company ad. So I'll
- 9:40
talk about uh how do we scale data in
- 9:42
robotics. Okay, let's take a tour back
- 9:44
as to how have people addressed this. uh
- 9:46
I'll take a look back at my own career
- 9:48
and my hypothesis for data has been
- 9:51
changing over the years. Okay. When I
- 9:52
began uh uh working scaling things the
- 9:55
idea was you can have robots you can
- 9:58
collect data for robots manually but it
- 10:00
is too slow and robots should be allowed
- 10:02
to collect data by themselves. So we had
- 10:04
this whole idea of curiositydriven
- 10:05
exploration. Uh allow robots to explore
- 10:08
in a curious way in the environment and
- 10:10
it has a lot of good parts like no human
- 10:13
required. Robots can go and keep playing
- 10:15
collect more data. We called it play
- 10:16
data at the time. It's very rich because
- 10:18
it has force sensors and joint angles.
- 10:21
But the difficulty is it is only in the
- 10:23
physical world. So it's very difficult
- 10:25
to scale. E Google was at it at the
- 10:28
time. Uh h having hundreds of robots.
- 10:31
Even for Google it's too expensive and
- 10:33
too slow to scale. The other idea is
- 10:35
again telly operation. As I said it's an
- 10:37
old idea but again the same issues
- 10:39
impossible to scale because now you
- 10:41
don't need robot you also need humans.
- 10:43
Uh so it's even more expensive and
- 10:45
diversity is very limited because even
- 10:47
if you put a robot in one setup with a
- 10:49
human you can get thousand examples but
- 10:52
they'll all be in the same setup. So
- 10:54
what you ideally want carry the robot to
- 10:56
new home every day in new scenario and
- 10:58
that's just impossibly hard to scale.
- 11:01
Then we had major breakthrough in
- 11:03
learning from simulation. This was one
- 11:05
of the first result where uh one could
- 11:08
show deep reinforcement learning uh
- 11:10
trained in simulation transfer to real
- 11:12
robot. This is one of one of the
- 11:14
award-winning paper uh at the time and
- 11:17
now it's used in every humanoid uh every
- 11:19
company out there. uh this was from our
- 11:21
lab at Berkeley and CMU. But again they
- 11:23
there are pros and cons. It's very easy
- 11:26
to scale but diversity is hard to get
- 11:28
because you have to engineer every scene
- 11:29
in simulation manually. So if you're
- 11:33
listening to this like there is no
- 11:34
single answer I'm coming to. Uh there is
- 11:36
no uh no no uh no single solution. This
- 11:39
is uh another work we did earlier. This
- 11:41
is all before skilled learning from
- 11:43
human videos. You watch the human do
- 11:45
things and then robot copies this. Again
- 11:47
high diversity. You can use videos from
- 11:49
YouTube etc. Highly scalable but it's a
- 11:52
very poor form of data because it's very
- 11:55
far from robot. Okay. So what is the
- 11:58
solution here? The answer is there is no
- 12:00
there is no golden path. You have to
- 12:02
think about data in the context of
- 12:05
different features. And in my opinion
- 12:08
there are only three features that
- 12:09
matter in in robotics.
- 12:12
scalability, diversity and how close you
- 12:15
are to your robot uh uh robot joint
- 12:17
angles. Okay, scalability means can I
- 12:20
quickly scale it across scenarios.
- 12:22
Diversity means can I get diverse data
- 12:25
because just having 100 trillion tokens
- 12:27
is completely useless if they're all in
- 12:30
the same environment and same scenarios.
- 12:31
Okay. And closeness to robot mean how
- 12:34
far are you from the ground truth of
- 12:36
robot own joint angles. So if you look
- 12:38
at simulation very scalable low diverse
- 12:42
uh uh but moderately close to robot
- 12:44
human videos are very scalable and
- 12:45
diverse but very far from robot data you
- 12:47
have to learn to map the human to robot
- 12:49
this other two teleop and the and the
- 12:51
manipulation interfaces they are kind of
- 12:53
in between and you can see here uh right
- 12:56
now around the world if you look at
- 12:57
different companies they're all focusing
- 12:59
on one of these approaches most of them
- 13:02
are on teleop or or yumi setups very few
- 13:05
on everything else and but there There
- 13:07
is no golden answer like they every each
- 13:09
one of them have a downside but the
- 13:11
golden light here is that they are all
- 13:14
complemented to each other like they are
- 13:16
not their cons are not exactly matching
- 13:19
from each other. So this is where uh the
- 13:21
recipe that we have converged to over
- 13:23
the years is to realize
- 13:26
separate the training into two parts
- 13:27
pre-training and post- training but
- 13:29
that's not surprise right but how do we
- 13:30
pre-train you want to pre-train using
- 13:32
data which is highly scalable and
- 13:33
diverse but maybe low quality so
- 13:35
simulation human videos etc. So that's
- 13:38
how you pre-train and then for post
- 13:40
training you use data from telly
- 13:41
operation. Now what is teleoperation
- 13:43
data? It's very high quality but low in
- 13:45
amount. Okay, high quality low in
- 13:48
amount. It exactly reminds us of the
- 13:50
recipe in language models. You pre-train
- 13:52
data on the internet then you post train
- 13:54
for coding for uh uh for your own
- 13:57
company etc.
- 13:58
But in robotics there is one more bucket
- 14:00
which is deployment data and deployment
- 14:03
data is highly scalable once deployment
- 14:06
scaled. So over time this data will take
- 14:09
over everything else. And what we are
- 14:12
trying to build is what we call this
- 14:14
data flywheel which goes which takes
- 14:17
this data from post training time and
- 14:18
puts back in pre-training. And now you
- 14:20
can see why this idea of omniodied brain
- 14:23
is extremely important because this is
- 14:25
very unlikely that you have only one
- 14:27
robot only one version deployed forever
- 14:30
in every task around the world. This
- 14:32
never happens in any area of hardware
- 14:35
except chips because they are very hard
- 14:37
to manufacture. Uh and you can still see
- 14:39
even in the chips uh inference chips are
- 14:41
coming left and right for many companies
- 14:43
these days. So in hardware this is never
- 14:45
the case. you have only one version of
- 14:46
the robot, one shape, which is why
- 14:48
omniodied intelligence is the enabler of
- 14:52
what we call a deployment data flywheel.
- 14:54
Okay, so let's now look at a few results
- 14:57
uh very quick. Now this brain is
- 15:00
extremely general. So we can do variety
- 15:02
of tasks very quickly. You may have you
- 15:04
have se you may have seen many tasks
- 15:05
like laundry folding etc. It's very
- 15:07
popular task in Silicon Valley for some
- 15:09
reason. uh and uh and the and the
- 15:11
argument here is uh when have you ever
- 15:14
thought while folding a t-shirt that if
- 15:17
I miss my hand by 1 cm my t-shirt fold
- 15:21
will be a blunder you don't think like
- 15:23
this people don't even think while
- 15:24
folding they just hold anywhere you just
- 15:26
do something so that the tolerance for
- 15:28
error is extremely extremely high then
- 15:32
why is this task hard
- 15:34
anybody why why do people get why do
- 15:37
people get fooled into believing this
- 15:38
task is
- 15:39
Let me put it that way.
- 15:42
Fabric, right? Why is fabric hard?
- 15:46
Simulation. But nobody's using
- 15:47
simulation anyway. This is all from
- 15:48
telly operation. Why is this hard? You
- 15:51
know why is it hard? Because it is hard
- 15:53
for classical version of robotics.
- 15:56
Classically in robotics, people would
- 15:58
model the whole physics, create models
- 16:00
by hand and then do this. It's very hard
- 16:01
for that. But for deep learning based
- 16:03
robotics, this is the easiest task
- 16:05
possible because it has high tolerance
- 16:07
for error. So it is completely uh now so
- 16:10
many companies focusing on this task.
- 16:12
Now what is hard is I would say
- 16:13
something about like let's say this
- 16:14
task.
- 16:17
If I ask you before seeing this video is
- 16:19
this task doable without having hands
- 16:22
very likely half the people will say no
- 16:24
because it requires putting like how
- 16:26
many of you have lost airports? Uh like
- 16:29
and in our company there's a channel
- 16:31
called right airpod because people keep
- 16:32
losing their right. Now in this case the
- 16:35
robot does not have hand it has gripper.
- 16:37
So here the task is very hard for the
- 16:40
gripper. So it requires a higher level
- 16:42
of intelligence because the arm has to
- 16:44
go and orient itself to pick up the
- 16:47
airpod in the right manner such that it
- 16:50
can be inserted because the grippers can
- 16:52
only close up and like like this.
- 16:54
They're parallejo grippers. So you
- 16:56
cannot turn the uh the the the airpod at
- 16:59
the very end. Now what you are seeing
- 17:00
here these are not real airpods. They
- 17:02
are fake ones from teu uh like $5 $10
- 17:05
each. So they don't have a magnet
- 17:07
inside. So the robot has to really work
- 17:09
hard to put this inside properly because
- 17:11
there is no magnet to pull it uh easily.
- 17:13
So these are all fake uh fake ones. This
- 17:15
is even harder than it appears uh in the
- 17:17
video. You can also uh like once you
- 17:21
build the general brain behind the
- 17:23
scene, you can also go and learn it from
- 17:25
variety of just human videos without
- 17:26
actually having any finetuning data at
- 17:28
tell operation time. So in this scenario
- 17:30
what we did, we trained on human videos
- 17:32
like this like egocentric videos. we are
- 17:34
now trying to transfer it to more third
- 17:36
person uh uh videos and if third person
- 17:39
works you can learn from YouTube any any
- 17:41
kind of open source data. So in this
- 17:43
case we see the video and then we add
- 17:46
less than 1 hour of robot data. So very
- 17:48
quick uh transformation and then the it
- 17:50
can transfer to humanoids. Now these
- 17:53
ones have hands. Hands or no hands is a
- 17:56
big debate. People often use hands as an
- 17:58
excuse as to why robotics is not here
- 18:01
but that's not the case. It's always the
- 18:03
intelligence uh behind the scene.
- 18:07
We don't deploy hands right now because
- 18:09
there are none available which can be
- 18:10
deployed in factories. They all break
- 18:12
within 100 yards. Uh pick anyone.
- 18:17
If you have better hand uh I would love
- 18:19
to buy 10 units ASAP to test. So it's
- 18:23
robust to scenarios. Now the idea is you
- 18:26
can do many tasks by watching humans.
- 18:28
And the reason we can work with very
- 18:30
less data which is less than one hour of
- 18:32
robot data is because robot imagines
- 18:34
things in its own head and tries to
- 18:36
multiply the learning from for many
- 18:38
scenarios. So what you see here is a
- 18:40
completely fake video uh you can call it
- 18:42
robot's dream. So it's inside the
- 18:44
robot's own model where it can imagine
- 18:46
scenarios uh in so this is not real
- 18:48
data. This is all fake uh from own
- 18:51
robot's own model. You can transfer it
- 18:53
to you know more complex task and even
- 18:56
lower cost hardware. So this is a egg uh
- 18:59
uh egg making like omelette making task.
- 19:02
Uh now here the these this whole setup
- 19:05
costs about $4,000. So extremely cheap
- 19:07
arms uh compared to the uh what you see
- 19:10
out there. And if you notice here this
- 19:14
setup is too cheap to even have any
- 19:16
sensors. So the only sensor here is just
- 19:18
a camera. No force nothing else. And the
- 19:20
robot can do task which require forces
- 19:23
uh from vision. Now I'm not saying
- 19:25
that's a future like you should not be
- 19:27
doing this. I'm sure the sensors will
- 19:29
improve but
- 19:31
>> from the existing sensors we are way
- 19:33
behind than where we can be from just
- 19:36
intelligence perspective.
- 19:37
>> Okay. So this can keep on going. This is
- 19:40
my adviser from Berkeley. He did not
- 19:42
believe the robots can do it. He just
- 19:44
kept standing for like uh 1 hour and the
- 19:46
robot kept making omelette. Uh and and
- 19:48
the interesting part here is this is
- 19:50
fully end to end system. There is no
- 19:51
state machine nothing. So the reason
- 19:54
robot is looking egg because there is an
- 19:55
empty plate in the front. I don't know
- 19:57
why it's fluctuating. It's not in the
- 19:58
there's no cut in the video. I I assure
- 20:00
you of that. It's HDMI problem. So as
- 20:03
soon as you remove the plate uh sorry I
- 20:06
don't know why this is happening.
- 20:08
Yeah. So there's no state machine when
- 20:10
the robot starts when the robot ends.
- 20:11
It's all automated and it's very robust
- 20:13
to even disturbances. You can change
- 20:15
objects around. Uh you can add these are
- 20:18
all unseen and unseen objects. Uh only
- 20:20
the pan and the gas stove is same.
- 20:22
Everything else is different and the
- 20:24
robot can keep on going. So you can get
- 20:26
basic robustness. So this was this is
- 20:28
all trained with less than 10 hours of
- 20:30
data. Uh so it's very low data to be
- 20:32
learning this robustness and it's coming
- 20:34
from the base model behind the scene.
- 20:37
I'm skipping uh this in the interest of
- 20:39
time. So unlike traditional robotics
- 20:42
pipelines where you have planning,
- 20:43
mapping etc. This is an end toend brain.
- 20:46
Now end to end is heavily used term in
- 20:49
in many areas of AI. But when I mean end
- 20:51
to end, I really mean end to end. It
- 20:53
reads directly from the cameras and it
- 20:55
applies power directly to motors. So we
- 20:59
use nothing uh in between except just a
- 21:02
P at the very end to go from 100 to 500
- 21:04
Hz. But P is not robotics contribution.
- 21:06
This is before it predates to like World
- 21:08
War II or something. Okay. Now the the
- 21:11
the interesting part is vision for us is
- 21:14
yet another input. So nothing that crazy
- 21:16
about it. So if you have seen, you may
- 21:19
have seen a lot of lot of results of
- 21:20
robots dancing, doing karate, kung fu,
- 21:22
backflip.
- 21:24
Go back and think how many results have
- 21:27
you seen of robots going up and down
- 21:29
stairs. Very few. Like even the top
- 21:32
companies out there in humanoid, they
- 21:34
have not shown beyond one sample stair
- 21:37
just to check the tick tick box. You do.
- 21:39
And people talk about humanoids are
- 21:41
coming, China versus US, all of this is
- 21:44
sort of BS. If humanoids cannot climb
- 21:46
stairs, what's the point of having legs?
- 21:49
All right, that's the only reason why
- 21:50
you have legs. Now, this is actually a
- 21:52
Moex paradox add action. Left one
- 21:56
actually is much easier for a robot. A
- 21:58
backflip, dancing, etc. is much easier
- 22:00
for the robot. And you may think, okay,
- 22:02
why is this paradox exist? Think one
- 22:04
level deeper. When a robot is doing a
- 22:07
back flip, it has to only know about its
- 22:10
own body, nothing else. Right? So it's
- 22:14
and when everything is known or fully
- 22:16
observed that's what computers are good
- 22:19
at because they can do they can simulate
- 22:22
every possible uh uh setup but the right
- 22:25
one even simply climbing on stairs
- 22:27
requires seeing the stair at the first
- 22:29
at the first place or what's the height
- 22:31
what's the width you don't measure
- 22:33
height and width exactly but you adjust
- 22:34
to all the disturbances and that
- 22:36
requires vision so right one requires
- 22:38
understanding and that's why it is hard
- 22:39
and you don't see any of this but for us
- 22:41
it's just get another result it doesn't
- 22:42
really matter so we And uh we put this
- 22:44
result out one and a half year ago. We
- 22:46
shot this like two and a half years ago.
- 22:48
Uh but uh like this robot can go any
- 22:51
scenario. It is not stumbling on things.
- 22:53
It intentionally steps over things and
- 22:55
there is no mapping or planning. It
- 22:57
doesn't make any 3D map of the
- 22:58
surroundings. It's all operating from
- 23:00
camera on the torso. And you can put
- 23:02
this in any kind of setup, any kind of
- 23:04
stairs. Uh I'm going faster here. These
- 23:07
are, you know, uh fire escape stairs.
- 23:11
It's a bit weird. Fire escape stairs are
- 23:13
to be used when there is fire in urgency
- 23:15
and they are the hardest in every
- 23:16
building. Uh so we are testing in all in
- 23:19
all those setups but you can put this in
- 23:21
anywhere you can disturb it and stay
- 23:22
doesn't really matter. This is
- 23:23
superhuman capability because robot
- 23:25
cannot see it's being pulled. So it's a
- 23:28
surprise factor for the robot that
- 23:29
you're pulling the violet while
- 23:30
climbing. And again it's the same model
- 23:32
everywhere in all these scenarios. So
- 23:35
very easy although easy parkour does
- 23:37
look fun uh uh does look fun and people
- 23:40
often find like uh this is uh you know
- 23:45
the robots can do all this but this is
- 23:48
so much easier than what I showed
- 23:50
earlier and even if you see these videos
- 23:53
right now many of you will find this
- 23:55
more impressive even though I'm telling
- 23:56
you this is easy it takes half an hour
- 23:58
to train this while the previous one it
- 24:00
takes much longer uh and much more
- 24:01
difficult to train
- 24:03
so you can take this Based model and you
- 24:05
can transfer to variety of tasks very
- 24:07
quickly. So here this is a very old
- 24:09
result from one and a half year ago uh
- 24:11
for inserting LAN cables etc. We have
- 24:13
very advanced systems now but we are
- 24:15
deploying these models already across
- 24:17
variety of application. So for instance
- 24:19
because to to build the data flywhe the
- 24:21
hardest part is the deployment itself.
- 24:24
There are so many other issues beside
- 24:25
intelligence and hardware that come up
- 24:27
when you deploy robotics. It is not the
- 24:29
same. This is not same same as deploying
- 24:32
chat GP over an app because you have to
- 24:35
work around many of the and I think
- 24:36
Skyio gave a talk before this and they
- 24:39
can talk to you about all day about the
- 24:40
hardness of deployment. So we have to
- 24:42
start deployment now so that we can have
- 24:45
this data flywheel in a foreseeable
- 24:46
future and I'll give you a few examples
- 24:48
of deployment we are doing. So one of
- 24:50
the examples with Nvidia like Nvidia is
- 24:52
opening their first factory in Houston.
- 24:54
uh and their GPUs what you use right now
- 24:56
they're all built outside US in in
- 24:58
Taiwan mainly and it's all done manually
- 25:00
over there and uh humos are really
- 25:02
efficient and really uh really good at
- 25:04
making these things but sustaining cost
- 25:07
and scalability and throughput here in
- 25:08
US it's very hard to maintain without
- 25:10
this labor force so here we are
- 25:12
automating uh this GPU assembly for them
- 25:15
this was we did a live demo in Nvidia
- 25:18
GTC this is deployed in factory already
- 25:20
uh last week so it's already live but
- 25:23
the interesting part I want to highlight
- 25:25
this is a very traditional factory task
- 25:28
but if you look at the the whole setup
- 25:32
this is extremely randomized extremely
- 25:34
noisy and you go to any factory
- 25:37
traditional setup they're extremely
- 25:39
clean so here the robot the same brain
- 25:42
can not only do very precise task it can
- 25:44
be very robust to any kind of
- 25:46
disturbance which means now these robots
- 25:48
do not need to be in a cage uh and they
- 25:51
can be uh just deployed as is on on
- 25:53
these factory lines without any change
- 25:55
to any factory line alongside humans and
- 25:58
wherever there is humans there is mess.
- 26:00
Okay, you can put the same brain to
- 26:01
delivery applications. So here the robot
- 26:03
delivers from the from the truck to the
- 26:05
to the door. It has to find where the
- 26:07
front door is. So it's a common sense
- 26:09
problem. It's not mobility problem. And
- 26:10
we already have been delivering uh
- 26:12
packages to with working behind the
- 26:14
scene with many partners to delivering
- 26:16
packages to front door of the houses uh
- 26:17
in these areas. same setup works across
- 26:20
warehouses and all this. Um now just to
- 26:24
close the talk uh close the topic I
- 26:27
mentioned this is an omnibody brain and
- 26:29
I hope it's very clear why it's
- 26:30
extremely essential to have an omnibody
- 26:32
brain to build data flywheel but there's
- 26:34
one extra benefit and the benefit is
- 26:36
your robots may break over time and uh
- 26:39
or the safety uh scenarios and safety is
- 26:42
a byproduct of this omnibody brain
- 26:44
because if you have a we can put the I'm
- 26:46
just skipping very fast here so you can
- 26:47
see the online this is all online uh but
- 26:49
we can put the same brain across each of
- 26:51
these systems each of these robots
- 26:55
and in this case we did not even train
- 26:57
on these robots. This is completely zero
- 26:59
shot transfer to all these humanoids and
- 27:01
all these things and even if the robot
- 27:03
breaks like here the leg breaks it can
- 27:06
recover within milliseconds because when
- 27:08
your leg breaks your three-legged robot
- 27:10
is a new robot. So it doesn't really
- 27:12
matter what's the shape of the robot is
- 27:14
anything changing here we disabled the
- 27:16
legs of the robot it learns to learn it
- 27:18
learns to walk on two legs in just three
- 27:20
trials. So what you are seeing on screen
- 27:22
is all the training for the robot that
- 27:24
they are that that is happening in these
- 27:26
systems. So it's it adapts in few
- 27:28
milliseconds to uh uh 30 seconds or so.
- 27:32
And like one example here it's it's
- 27:33
going on wheels. You jam the wheel it
- 27:36
starts walking. So this is a another
- 27:38
byproduct of omnibody brain. Safety
- 27:42
takes a very different meaning uh with
- 27:44
these kind of models. when the robot can
- 27:46
fly like or sorry when the robots can
- 27:48
walk or or operate with only half the
- 27:50
body. I removed some of the gory scenes
- 27:52
from this but we also cut the robot in
- 27:53
half and it can still work uh with both
- 27:56
the two halves separately but for that
- 27:57
you can go to YouTube uh and that's all
- 27:59
I have. Thank you.