AI Engineer World's Fair 2026

Robotics Has Been Stuck for 70 Years — Deepak Pathak, Skild AI

Read the talk

Robotics Has Been Stuck for 70 Years

Deepak Pathak explains Skild AI’s approach to a general robot brain: combine complementary data sources, share learning across different bodies, and use deployment to improve the model. The demonstrations reveal why precise grasps, unfamiliar stairs, and damaged hardware test more than a robot’s appearance.

From a talk by Deepak Pathak

At a glance

Ideas worth remembering

  • Judge robot data by scalability, environmental diversity, and proximity to real robot action. Large quantities of repeated experience in one setup do not satisfy all three.

  • Skild’s recipe assigns complementary roles to simulation and human-video pre-training, teleoperation post-training, and deployment experience returned to training.

  • Hardware constraints shape the required intelligence: a parallel-jaw gripper must choose an earbud grasp that preserves the orientation needed for insertion.

  • Visually guided stairs can demand more than spectacular body maneuvers because the robot must connect unfamiliar environmental geometry to its actions.

  • Supporting different bodies may also support recovery after damage. The disabled-leg and jammed-wheel examples make that benefit concrete without establishing a general safety guarantee.

An impressive demo is an old achievement

A robot looks at a picture of blocks and arranges physical blocks to match it. The task requires connecting a two-dimensional image to three-dimensional objects, then moving those objects precisely. In Deepak Pathak’s opening example, that apparently modern demonstration comes from the 1960s: the MIT copy demo. Pathak, Skild AI’s co-founder and CEO and a Carnegie Mellon professor, uses it to challenge the excitement surrounding robotics as AI’s next frontier. 1:27

Source frame: An impressive demo is an old achievement
Source frame: An impressive demo is an old achievement

The next clip reaches further back. A human controls robot manipulators through a leader–follower system in 1957. Its basic principle survives in modern teleoperation: a person supplies the motion, and a robot follows. Later examples include juggling, foosball, and a Berkeley robot that clears a table and packs objects into a box, which Pathak describes as the work of one graduate student with a single GPU machine. His dinner wager is deliberately provocative: distinguish old robotics videos from today’s.

The historical comparison supports Pathak’s diagnosis: robotics has concentrated on building particular machines without developing a general brain. That is a claim about the field’s direction, rather than proof that nothing improved. The useful distinction is between making one impressive behavior work and building intelligence that transfers to different tasks and situations.

Moravec’s paradox gives the problem its memorable phrasing: “Hard is easy, easy is hard.” Mathematical competition and chess look intellectually demanding to humans; climbing stairs and picking up chess pieces feel ordinary. Yet success at the intellectual task does not settle the physical one. Winning a board game leaves the problem of moving its pieces on an unfamiliar board.

Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:12 · section reference included

No internet of robot data

Large datasets and large models have been a productive recipe elsewhere in AI. Robotics lacks an equivalent internet of action data. Teleoperation supplies useful examples, but each example requires a physical robot and a human’s time. Pathak estimates roughly one minute per example. More operators can increase collection, but they do not remove its labor and hardware costs.

Source frame: No internet of robot data
Source frame: No internet of robot data

Skild AI’s proposed answer is an omni-bodied intelligence: “any robot, any task, one brain.” A humanoid, quadruped, conveyor-belt arm, and dexterous hand should all contribute to the same learning system. The motivation is data scarcity. Restricting training to one machine would discard experience from other hardware, tasks, and environments that the model needs. 7:22

The teaser presents robots from different companies running the Skild brain, including locomotion in new scenarios and responses to disturbances. Its intended relationship is reciprocal: shared intelligence controls many bodies, and experience from those bodies improves shared intelligence. The demonstrations and deployment results throughout the talk are Pathak’s reports; they do not establish general success rates across arbitrary robots or environments.

Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

5:40 · section reference included

Three properties of useful robot data

Pathak’s retrospective moves through several attempts to collect robot experience. Each solves part of the problem and leaves another part exposed. 9:31

Source frame: Three properties of useful robot data
Source frame: Three properties of useful robot data
  • Curiosity-driven exploration: Robots explore and collect their own “play data,” including forces and joint angles. This removes the need for a human to guide every action, but collection still proceeds through physical machines in the real world. Hardware cost and physical time constrain scale.
  • Teleoperation: Human-guided trajectories supply data closely tied to robot action. Collection requires both robots and operators, and thousands of examples in one setup can still offer little environmental diversity. Moving the robot to a different home every day would improve coverage while making collection harder.
  • Simulation: Deep reinforcement learning can train in simulation and transfer to a real robot. Simulated experience is easier to multiply, but someone must build the scenes. More trials within engineered scenes do not automatically create more kinds of environments.
  • Human video: Videos offer scalable, diverse demonstrations of people doing things. Their weakness is distance from the robot’s action space: observing a human movement does not directly supply the joint commands a different body needs.

The evaluation has three axes: scalability, diversity, and closeness to the robot’s own joint-angle ground truth. Repeating a task in one room can produce a large dataset without teaching the robot much about a different room.

There is no single winning source. The useful discovery is that their weaknesses differ. Human video contributes variety that carefully engineered simulation struggles to supply. Teleoperation contributes direct robot experience that video lacks. Combining them gives each source a job instead of asking one collection method to satisfy every requirement.

Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

9:32 · section reference included

Pre-training creates the base; deployment feeds it

The training recipe separates breadth from precision. Pre-training uses scalable sources such as simulation and human videos, accepting their imperfect relationship to real robot action. Post-training uses smaller amounts of high-quality teleoperation data. The analogy to language models is functional: broad initial learning gives later, narrower training something useful to adapt.

Source frame: Pre-training creates the base; deployment feeds it
Source frame: Pre-training creates the base; deployment feeds it

Deployment adds a third source. Once robots operate at scale, their experience can feed subsequent training. Pathak expects this data eventually to outweigh the initial sources, but that expectation depends on achieving deployment scale. A model restricted to one robot shape or hardware version would fragment the experience as machines change. An omni-bodied model is intended to keep those different deployments useful to one shared brain.

How does experience become reusable across the training stages? The diagram follows the proposed loop. Simulation and video supply the starting breadth; teleoperation helps turn that base into deployable behavior; operational experience returns to training. Supporting many bodies matters at the return path, where otherwise incompatible deployments must contribute to the same model.

How it fits togetherHow different data sources feed one robot brain

Scalable experience with complementary weaknesses

The proposed flywheel combines broad pre-training, precise post-training, and experience from deployments across changing robot bodies.

13:2113:23
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

13:21 · section reference included

The insertion starts before the gripper closes

The demonstrations begin by reconsidering what makes a manipulation task difficult. Laundry folding has a reputation for difficulty because fabric is awkward to model with hand-built physics. But a learned policy faces a different challenge: how much execution error can the task tolerate? A hand that lands a centimeter away from its intended position may still produce an acceptable fold. Pathak uses that tolerance to explain why laundry can be forgiving for learning-based control, despite being difficult for classical modeling.

Source frame: The insertion starts before the gripper closes
Source frame: The insertion starts before the gripper closes

Putting an AirPod-shaped earbud into its case exposes a tighter constraint. The robot has a parallel-jaw gripper, whose jaws close around the object without the in-hand manipulation available to fingers. The arm must approach and grasp the earbud in an orientation that permits the later insertion. A convenient pickup can leave the object inconveniently positioned for the next step.

The observable change is from a loose earbud to one seated in its case. Its causal sequence starts earlier: orient the arm, close the jaws around a usable grasp, carry the held object to the case, then align and insert it. The props are imitation earbuds costing roughly $5–$10 each, without magnets to help pull them into place. That detail removes a source of mechanical assistance: the robot must achieve the fit itself.

15:0515:11
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

14:54 · section reference included

Human demonstrations, imagined scenarios, and omelets

The next learning example uses egocentric human video followed by less than one hour of additional robot data to transfer behavior to humanoids. Third-person video remains a direction under investigation in this account; successful transfer would open access to a wider range of ordinary videos. The important bridge is from watching a human do something to making a different body do it, with limited robot-specific experience.

Source frame: Human demonstrations, imagined scenarios, and omelets
Source frame: Human demonstrations, imagined scenarios, and omelets

Those humanoids have hands, but Skild is not deploying hands in factories at the time described. Pathak’s objection is mechanical reliability. A more capable brain can make a simple gripper useful, while adding fingers does not solve the problem of fragile hardware. His challenge to anyone with a better hand is practical: he would buy ten units to test.

To explain learning from little robot data, Pathak introduces the robot’s “dream”: generated scenarios inside its own model. The displayed video represents imagined experience rather than a recording of physical trials. The proposed mechanism multiplies learning across scenarios without physically collecting every one. The talk does not specify the generation objective or how imagined outcomes are checked against real physics, so this remains a high-level account of the data expansion.

Omelet making then tests the approach on a roughly $4,000 setup. Its only sensor is a camera; it has no force sensing. Contact tasks therefore depend on visual information to guide actions that apply force. Pathak does not recommend avoiding better sensors. His point is that existing hardware can support more capable behavior than its current software extracts.

The cooking behavior is described as fully end to end, without a hand-written state machine deciding when the sequence starts or ends. An empty plate in front of the robot provides visual context for making an egg. Objects can move or be replaced, and the robot continues; the pan and gas stove remain the same while other objects differ. Pathak reports less than ten hours of training data for this example and attributes the resulting tolerance to the base model. The fixed pan and stove make the demonstrated scope concrete.

17:1717:21
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

17:17 · section reference included

Why stairs demand more than a backflip

End to end has a specific meaning here: camera observations enter the brain, and its output drives motor power. Pathak contrasts this with a traditional pipeline containing explicit mapping and planning stages. He also mentions a final control component bridging 100 Hz to 500 Hz; the captions render its name as “P,” leaving its technical identity unclear. The supported distinction is direct learned control without an explicit map-and-plan pipeline, rather than a claim that no lower-level control exists. 20:41

Source frame: Why stairs demand more than a backflip
Source frame: Why stairs demand more than a backflip

Locomotion returns to Moravec’s paradox. Backflips and dancing look spectacular, but Pathak describes them as comparatively easy because the robot primarily needs to control its own known body. Stair climbing adds unfamiliar geometry. The controller must see the steps, respond to their height and width, and adjust to disturbances. Knowing the body is insufficient when the next foothold depends on the surroundings.

What changes in the control problem when a robot encounters stairs? The comparison below makes the extra information visible. A familiar body supports rehearsal of a body-centered maneuver. An unfamiliar staircase requires visual information about the environment to affect action. That connection explains why an ordinary-looking ascent can test more general capability than a dramatic flip.

The stair examples use a torso camera without constructing a three-dimensional map of the surroundings. The robot intentionally steps over obstacles, traverses different stairs including fire escapes, and responds to being pulled while climbing. Pathak identifies the model as the same across these scenarios. He then shows parkour and reports that the displayed behavior takes half an hour to train, while the earlier visually guided locomotion takes much longer. The training-time comparison reinforces his warning: spectacle is a poor shortcut for judging difficulty.

Compare the ideasKnown-body motion versus visually guided locomotion

Body-centered motion can be simulated

Pathak’s comparison concerns the information required for control: stairs add unfamiliar environmental geometry and disturbances.

Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

20:37 · section reference included

The flywheel needs working deployments—and bodies that can change

The final stretch moves from demonstrations to deployment. After an example of LAN-cable insertion, Pathak identifies deployment itself as the hardest part of starting the data flywheel. Physical operations introduce problems beyond intelligence and hardware, so waiting for a perfect robot brain would also delay collecting the experience needed to improve it.

Source frame: The flywheel needs working deployments—and bodies that can change
Source frame: The flywheel needs working deployments—and bodies that can change
  • GPU assembly: Pathak describes work with NVIDIA for its Houston factory and reports that the assembly system is already deployed at the time of the talk. The example combines precise manipulation with a randomized, noisy setup. Its practical aim is to operate on existing factory lines amid the mess that accompanies human work. Tolerance of disturbances alone, however, does not establish that operation alongside people is safe without guarding.
  • Package delivery: The robot travels from a truck to a house’s front door. Reaching the destination requires identifying which part of the scene is the front door, so the task combines movement with an understanding of the destination. Pathak reports deliveries with partners and says the same setup also works in warehouses.

The ending gives omni-bodied learning another purpose: adaptation when the body changes unexpectedly. A robot with a broken leg has effectively become a different robot. A brain trained to support different bodies may therefore handle damage as a change in embodiment rather than assuming that its original hardware remains intact. Pathak presents transfers to previously untrained robots and recovery after a leg failure as examples of that capability.

In one example, disabling legs leaves the robot learning to walk on two legs in three trials. In another, jamming a wheel causes the robot to start walking. Pathak describes adaptation ranging from milliseconds to roughly 30 seconds across these examples, and reports a further experiment in which two halves of a robot can operate separately. These are distinct forms of recovery, rather than one universal recovery time. Pathak presents continued operation after damage as a potential safety benefit, but these adaptation examples do not establish that continued motion improves safety or provides a complete safety guarantee.

24:0324:05
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

24:03 · section reference included

Resources

From the talk

Read the complete timestamped transcript
  1. 0:01

    [music]

  2. 0:12

    I think uh all of you know here that how

  3. 0:15

    much progress has AI made in the last

  4. 0:16

    five years. Uh seems like it's more than

  5. 0:19

    uh last 100 years of progress. And what

  6. 0:23

    is happening especially in AI is AI came

  7. 0:25

    for language then speech audio video and

  8. 0:28

    everybody's excitement is next is

  9. 0:30

    robotics okay and this excitement has is

  10. 0:33

    going on peak uh this year uh for some

  11. 0:37

    reason which is still beyond my uh

  12. 0:39

    understanding and you can also see uh

  13. 0:42

    leaders like Jensen talking about

  14. 0:44

    physical AI as the next frontier of

  15. 0:45

    robotics. So it seems like if you open

  16. 0:48

    Twitter or LinkedIn, it seems like

  17. 0:49

    robots robotics is already here and

  18. 0:51

    we'll have home robots in our homes

  19. 0:53

    sometime by December uh as uh as it's

  20. 0:56

    being said by multiple people. But I'm

  21. 0:59

    here to highlight that robotics has been

  22. 1:02

    almost here for the last 70 years

  23. 1:06

    [sighs]

  24. 1:06

    now. Those of you who got into robotics

  25. 1:09

    in the last four or five years, go and

  26. 1:12

    take any any any course in robotics and

  27. 1:15

    if you can distinguish the videos in

  28. 1:17

    robotics from 30 years ago versus today,

  29. 1:20

    uh I owe you a dinner. So to give you an

  30. 1:24

    example, let's look at this robot. So

  31. 1:26

    this robot here is an image and the goal

  32. 1:28

    of this robot is to look at this image

  33. 1:30

    of blocks and arrange these blocks in

  34. 1:33

    front of it so that it looks the same

  35. 1:36

    pattern as in the image. Okay, this is

  36. 1:38

    an extremely simple task. But if you

  37. 1:40

    think about this, it's a 3D block. You

  38. 1:42

    have to look at understand 3D from each

  39. 1:43

    side. And the robot can do this pretty

  40. 1:46

    precisely. Any guesses how old this

  41. 1:48

    result is?

  42. 1:51

    Anyone? It's a guess. You can make any

  43. 1:52

    guess. Huh?

  44. 1:54

    >> 40 years. Okay, that's the highest.

  45. 1:57

    [panting]

  46. 1:58

    This is from 1960s. This is way before

  47. 2:02

    there were computers. Uh, okay. So, this

  48. 2:04

    is this is called MIT copy demo. Uh,

  49. 2:06

    this was the start of uh this is this

  50. 2:09

    predates AI. Uh, as you know, today look

  51. 2:12

    at this one.

  52. 2:17

    >> One exhibit at the Nuclear Congress in

  53. 2:19

    Philadelphia has the perfect formula for

  54. 2:22

    >> what's happening here. The guy behind

  55. 2:23

    the scene is controlling these robots, a

  56. 2:26

    leader follower system. And if you look

  57. 2:28

    at any big lab, any big company, any big

  58. 2:31

    academic lab, how they get data, they

  59. 2:34

    use teley operation system, which

  60. 2:35

    follows the exact same principles.

  61. 2:38

    Any guesses for the year for this one?

  62. 2:42

    >> Now, everybody would correct

  63. 2:44

    aggressively on the year, but this is

  64. 2:46

    from 1957, 68 years ago. Okay? Ask

  65. 2:51

    anybody in your family who is older than

  66. 2:53

    actually they may not be around like so

  67. 2:56

    because this is for 60 you have to ask

  68. 2:58

    somebody who is 80 year old what was the

  69. 3:00

    state of technology at that time like

  70. 3:02

    RAM for few KB of RAM you would have a

  71. 3:05

    size of a room uh like this big of a

  72. 3:07

    setup and this is from that time when

  73. 3:09

    people could make it work on not just

  74. 3:12

    this you keep coming back every decade

  75. 3:14

    the videos look as cool like here is

  76. 3:16

    robot a humanoid juggling playing

  77. 3:19

    foosball uh with human and this is also

  78. 3:23

    from you can see this is all they were

  79. 3:24

    all predate deep learning this from 8

  80. 3:26

    years ago so this robot is a result from

  81. 3:28

    Berkeley can clean up the whole table

  82. 3:30

    can arrange items in a in a in a box

  83. 3:33

    this was done by one grad student with a

  84. 3:35

    single GPU machine nothing more than

  85. 3:37

    that okay now this picture of and if I

  86. 3:40

    colorize all these videos from past and

  87. 3:42

    apply genai filter for the modern voice

  88. 3:45

    you cannot tell that this [laughter]

  89. 3:47

    video is from 19657 or this is from

  90. 3:49

    today And these are not even the oldest

  91. 3:51

    results the oldest ones go to 1940s.

  92. 3:54

    Okay. So what is this that if you look

  93. 3:58

    at last 70 years tech every other

  94. 4:00

    technology every other technology has

  95. 4:03

    come way far language understanding

  96. 4:05

    computer vision mobile phone chips why

  97. 4:08

    is robotics is stuck in this primitive

  98. 4:10

    age uh for for this long and the reason

  99. 4:12

    is uh a general brain. Robotics has

  100. 4:15

    always been approached as a hardware

  101. 4:17

    problem from ground up and this the

  102. 4:19

    reason that everything around robotics

  103. 4:21

    has progressed but robotics is still

  104. 4:23

    stuck in the same land from the last 70

  105. 4:25

    years. So there's a very famous paradox

  106. 4:26

    called Morave paradox. I don't know if

  107. 4:28

    you know about Moravec. He was also a

  108. 4:30

    CMU professor and I'm also CMU professor

  109. 4:32

    so have to quote him for sure. He was

  110. 4:34

    one of the founding figures in in AI.

  111. 4:36

    After 30 years of working in robotics he

  112. 4:39

    arrived at a very simple conclusion.

  113. 4:41

    Hard is easy, easy is hard. Okay.

  114. 4:44

    Whatever human believe to be hard is

  115. 4:46

    extremely easy for computers and vice

  116. 4:48

    versa. And you can see it happening

  117. 4:50

    right in front of your eyes. You would

  118. 4:51

    say doing math is hard. Math Olympiad

  119. 4:54

    gold medal is really really hard.

  120. 4:56

    Climbing a stair oh super easy. Now look

  121. 4:59

    around you in technology where we are in

  122. 5:01

    terms of what has been solved and what's

  123. 5:03

    not being solved. So this is robotics

  124. 5:05

    for you. Okay. It's very it is not yet

  125. 5:07

    another application of AI. It is what AI

  126. 5:10

    was founded for in the very beginning

  127. 5:12

    and has made very little progress

  128. 5:13

    towards. This is not yet another

  129. 5:15

    application where deep networks can come

  130. 5:16

    in and attack the field. This is has to

  131. 5:19

    be thought of with fundamentally first

  132. 5:21

    principles from the ground up. This is a

  133. 5:23

    very famous example of where you can

  134. 5:25

    have a you know a computer beat uh Gary

  135. 5:28

    Casper on chess in '90s but you can you

  136. 5:31

    still don't have a a a computer or robot

  137. 5:33

    that can pick up the chess pieces and

  138. 5:35

    arrange them on any chess board. Now,

  139. 5:40

    so how do we go about solving this?

  140. 5:41

    Okay,

  141. 5:43

    so on a more positive note, uh we have a

  142. 5:47

    magic sauce for AI success, right? You

  143. 5:49

    get big data set, you train big models

  144. 5:52

    and magic happens. Okay, now we know

  145. 5:54

    this template is working well very well

  146. 5:57

    in many topics. Okay, now can we apply

  147. 6:01

    the same recipe to robotics? Well, it's

  148. 6:03

    not uh directly applicable because we

  149. 6:06

    have no data. There is no internet of

  150. 6:08

    robotics data and as I said earlier you

  151. 6:10

    can go and collect data uh manually on

  152. 6:12

    the robot called telly operation and as

  153. 6:15

    you notice teley operation is not 5 year

  154. 6:17

    old thing this is a 68 or 70 year old

  155. 6:20

    thing so the the funny part is if you go

  156. 6:24

    back and you look at the whole evolution

  157. 6:25

    of GPD3 GPD4 models go back to GPD3 uh

  158. 6:30

    three years ago seems like a like a

  159. 6:32

    lifetime ago GPD3 only began working and

  160. 6:35

    caught people's attention when you could

  161. 6:37

    train those models on trillions of

  162. 6:39

    tokens. So GPD3 was already trained 30

  163. 6:41

    trillion plus token and today hundreds

  164. 6:43

    of trillions of tokens and if you

  165. 6:45

    collect data by manually by tell

  166. 6:46

    operation it takes you about 1 minute to

  167. 6:48

    get one example. You can do the math if

  168. 6:51

    you hire all of US population it will

  169. 6:53

    take you more than a century to reach

  170. 6:54

    the same scale as of GPD3 uh which is

  171. 6:57

    like uh and nobody uses GP3 today. Okay.

  172. 7:01

    So this is very slow and expensive. So

  173. 7:02

    in robotics I would argue nobody has

  174. 7:06

    really scaled robotics yet and we are

  175. 7:08

    very far from talking about scale in

  176. 7:10

    robotics with the way other uh other

  177. 7:12

    areas have seen scale. Okay so this is

  178. 7:15

    where uh we have been focusing on uh our

  179. 7:17

    scaled is about uh 3 years old but uh I

  180. 7:20

    have been working on the problem for

  181. 7:21

    more than a decade uh that's all I've

  182. 7:23

    done in my career uh nothing else. Uh so

  183. 7:26

    at scale what our thesis is is to build

  184. 7:28

    what we call an omniodied intelligence

  185. 7:32

    any robot any task one brain okay any

  186. 7:36

    robot it can be a humanoid it can be a

  187. 7:38

    quadriped it can be a robotic arm on a

  188. 7:40

    conveyor belt or a dextrous hand doesn't

  189. 7:41

    really matter this even this hypothesis

  190. 7:44

    is way more general than uh one would

  191. 7:47

    argue humans are because we control our

  192. 7:49

    own body and why do we have to go so

  193. 7:51

    general well the argument is in robotics

  194. 7:54

    there there is no data anyway. So I

  195. 7:56

    can't pick and choose which hardware do

  196. 7:58

    I use data from. And we should be able

  197. 8:00

    to use data from any kind of hardware,

  198. 8:02

    any kind of task, any kind of scenario.

  199. 8:04

    And that's the only way to truly achieve

  200. 8:06

    the scale of uh of what language models

  201. 8:08

    achieved uh three years ago. Okay. So

  202. 8:10

    the goal here is this uh this whole idea

  203. 8:13

    of any robot any task one brain. And

  204. 8:15

    through this talk I'll hopefully

  205. 8:16

    convince you why this is the this is the

  206. 8:18

    way to go towards robotics. Okay. So

  207. 8:20

    this is the uh rough intro but before I

  208. 8:23

    go into the more details let me show you

  209. 8:25

    just a teaser result.

  210. 8:28

    So in this result, every single robot uh

  211. 8:31

    it's from a different company, different

  212. 8:32

    hardware and they're all controlled by

  213. 8:35

    uh skilled brain whether it's humanoids

  214. 8:38

    going up and down stairs, any kind of

  215. 8:40

    scenarios.

  216. 8:42

    And these are not new results. They're

  217. 8:44

    like a couple of year old uh results in

  218. 8:46

    here. Robust to dust disturbances.

  219. 8:51

    You can put them zero shot in new

  220. 8:52

    scenarios.

  221. 9:02

    So the idea here any every single robot

  222. 9:06

    in this video takes any action anywhere

  223. 9:09

    in the world the brain behind the scene

  224. 9:12

    [music] improves because it's an

  225. 9:14

    omni-body brain.

  226. 9:15

    >> Yeah, this is really hard right because

  227. 9:16

    like eggs are okay.

  228. 9:32

    So that's a teaser of what we what we

  229. 9:34

    work on. So I want to make this talk

  230. 9:35

    more scientific and more uh more

  231. 9:37

    informative than a company ad. So I'll

  232. 9:40

    talk about uh how do we scale data in

  233. 9:42

    robotics. Okay, let's take a tour back

  234. 9:44

    as to how have people addressed this. uh

  235. 9:46

    I'll take a look back at my own career

  236. 9:48

    and my hypothesis for data has been

  237. 9:51

    changing over the years. Okay. When I

  238. 9:52

    began uh uh working scaling things the

  239. 9:55

    idea was you can have robots you can

  240. 9:58

    collect data for robots manually but it

  241. 10:00

    is too slow and robots should be allowed

  242. 10:02

    to collect data by themselves. So we had

  243. 10:04

    this whole idea of curiositydriven

  244. 10:05

    exploration. Uh allow robots to explore

  245. 10:08

    in a curious way in the environment and

  246. 10:10

    it has a lot of good parts like no human

  247. 10:13

    required. Robots can go and keep playing

  248. 10:15

    collect more data. We called it play

  249. 10:16

    data at the time. It's very rich because

  250. 10:18

    it has force sensors and joint angles.

  251. 10:21

    But the difficulty is it is only in the

  252. 10:23

    physical world. So it's very difficult

  253. 10:25

    to scale. E Google was at it at the

  254. 10:28

    time. Uh h having hundreds of robots.

  255. 10:31

    Even for Google it's too expensive and

  256. 10:33

    too slow to scale. The other idea is

  257. 10:35

    again telly operation. As I said it's an

  258. 10:37

    old idea but again the same issues

  259. 10:39

    impossible to scale because now you

  260. 10:41

    don't need robot you also need humans.

  261. 10:43

    Uh so it's even more expensive and

  262. 10:45

    diversity is very limited because even

  263. 10:47

    if you put a robot in one setup with a

  264. 10:49

    human you can get thousand examples but

  265. 10:52

    they'll all be in the same setup. So

  266. 10:54

    what you ideally want carry the robot to

  267. 10:56

    new home every day in new scenario and

  268. 10:58

    that's just impossibly hard to scale.

  269. 11:01

    Then we had major breakthrough in

  270. 11:03

    learning from simulation. This was one

  271. 11:05

    of the first result where uh one could

  272. 11:08

    show deep reinforcement learning uh

  273. 11:10

    trained in simulation transfer to real

  274. 11:12

    robot. This is one of one of the

  275. 11:14

    award-winning paper uh at the time and

  276. 11:17

    now it's used in every humanoid uh every

  277. 11:19

    company out there. uh this was from our

  278. 11:21

    lab at Berkeley and CMU. But again they

  279. 11:23

    there are pros and cons. It's very easy

  280. 11:26

    to scale but diversity is hard to get

  281. 11:28

    because you have to engineer every scene

  282. 11:29

    in simulation manually. So if you're

  283. 11:33

    listening to this like there is no

  284. 11:34

    single answer I'm coming to. Uh there is

  285. 11:36

    no uh no no uh no single solution. This

  286. 11:39

    is uh another work we did earlier. This

  287. 11:41

    is all before skilled learning from

  288. 11:43

    human videos. You watch the human do

  289. 11:45

    things and then robot copies this. Again

  290. 11:47

    high diversity. You can use videos from

  291. 11:49

    YouTube etc. Highly scalable but it's a

  292. 11:52

    very poor form of data because it's very

  293. 11:55

    far from robot. Okay. So what is the

  294. 11:58

    solution here? The answer is there is no

  295. 12:00

    there is no golden path. You have to

  296. 12:02

    think about data in the context of

  297. 12:05

    different features. And in my opinion

  298. 12:08

    there are only three features that

  299. 12:09

    matter in in robotics.

  300. 12:12

    scalability, diversity and how close you

  301. 12:15

    are to your robot uh uh robot joint

  302. 12:17

    angles. Okay, scalability means can I

  303. 12:20

    quickly scale it across scenarios.

  304. 12:22

    Diversity means can I get diverse data

  305. 12:25

    because just having 100 trillion tokens

  306. 12:27

    is completely useless if they're all in

  307. 12:30

    the same environment and same scenarios.

  308. 12:31

    Okay. And closeness to robot mean how

  309. 12:34

    far are you from the ground truth of

  310. 12:36

    robot own joint angles. So if you look

  311. 12:38

    at simulation very scalable low diverse

  312. 12:42

    uh uh but moderately close to robot

  313. 12:44

    human videos are very scalable and

  314. 12:45

    diverse but very far from robot data you

  315. 12:47

    have to learn to map the human to robot

  316. 12:49

    this other two teleop and the and the

  317. 12:51

    manipulation interfaces they are kind of

  318. 12:53

    in between and you can see here uh right

  319. 12:56

    now around the world if you look at

  320. 12:57

    different companies they're all focusing

  321. 12:59

    on one of these approaches most of them

  322. 13:02

    are on teleop or or yumi setups very few

  323. 13:05

    on everything else and but there There

  324. 13:07

    is no golden answer like they every each

  325. 13:09

    one of them have a downside but the

  326. 13:11

    golden light here is that they are all

  327. 13:14

    complemented to each other like they are

  328. 13:16

    not their cons are not exactly matching

  329. 13:19

    from each other. So this is where uh the

  330. 13:21

    recipe that we have converged to over

  331. 13:23

    the years is to realize

  332. 13:26

    separate the training into two parts

  333. 13:27

    pre-training and post- training but

  334. 13:29

    that's not surprise right but how do we

  335. 13:30

    pre-train you want to pre-train using

  336. 13:32

    data which is highly scalable and

  337. 13:33

    diverse but maybe low quality so

  338. 13:35

    simulation human videos etc. So that's

  339. 13:38

    how you pre-train and then for post

  340. 13:40

    training you use data from telly

  341. 13:41

    operation. Now what is teleoperation

  342. 13:43

    data? It's very high quality but low in

  343. 13:45

    amount. Okay, high quality low in

  344. 13:48

    amount. It exactly reminds us of the

  345. 13:50

    recipe in language models. You pre-train

  346. 13:52

    data on the internet then you post train

  347. 13:54

    for coding for uh uh for your own

  348. 13:57

    company etc.

  349. 13:58

    But in robotics there is one more bucket

  350. 14:00

    which is deployment data and deployment

  351. 14:03

    data is highly scalable once deployment

  352. 14:06

    scaled. So over time this data will take

  353. 14:09

    over everything else. And what we are

  354. 14:12

    trying to build is what we call this

  355. 14:14

    data flywheel which goes which takes

  356. 14:17

    this data from post training time and

  357. 14:18

    puts back in pre-training. And now you

  358. 14:20

    can see why this idea of omniodied brain

  359. 14:23

    is extremely important because this is

  360. 14:25

    very unlikely that you have only one

  361. 14:27

    robot only one version deployed forever

  362. 14:30

    in every task around the world. This

  363. 14:32

    never happens in any area of hardware

  364. 14:35

    except chips because they are very hard

  365. 14:37

    to manufacture. Uh and you can still see

  366. 14:39

    even in the chips uh inference chips are

  367. 14:41

    coming left and right for many companies

  368. 14:43

    these days. So in hardware this is never

  369. 14:45

    the case. you have only one version of

  370. 14:46

    the robot, one shape, which is why

  371. 14:48

    omniodied intelligence is the enabler of

  372. 14:52

    what we call a deployment data flywheel.

  373. 14:54

    Okay, so let's now look at a few results

  374. 14:57

    uh very quick. Now this brain is

  375. 15:00

    extremely general. So we can do variety

  376. 15:02

    of tasks very quickly. You may have you

  377. 15:04

    have se you may have seen many tasks

  378. 15:05

    like laundry folding etc. It's very

  379. 15:07

    popular task in Silicon Valley for some

  380. 15:09

    reason. uh and uh and the and the

  381. 15:11

    argument here is uh when have you ever

  382. 15:14

    thought while folding a t-shirt that if

  383. 15:17

    I miss my hand by 1 cm my t-shirt fold

  384. 15:21

    will be a blunder you don't think like

  385. 15:23

    this people don't even think while

  386. 15:24

    folding they just hold anywhere you just

  387. 15:26

    do something so that the tolerance for

  388. 15:28

    error is extremely extremely high then

  389. 15:32

    why is this task hard

  390. 15:34

    anybody why why do people get why do

  391. 15:37

    people get fooled into believing this

  392. 15:38

    task is

  393. 15:39

    Let me put it that way.

  394. 15:42

    Fabric, right? Why is fabric hard?

  395. 15:46

    Simulation. But nobody's using

  396. 15:47

    simulation anyway. This is all from

  397. 15:48

    telly operation. Why is this hard? You

  398. 15:51

    know why is it hard? Because it is hard

  399. 15:53

    for classical version of robotics.

  400. 15:56

    Classically in robotics, people would

  401. 15:58

    model the whole physics, create models

  402. 16:00

    by hand and then do this. It's very hard

  403. 16:01

    for that. But for deep learning based

  404. 16:03

    robotics, this is the easiest task

  405. 16:05

    possible because it has high tolerance

  406. 16:07

    for error. So it is completely uh now so

  407. 16:10

    many companies focusing on this task.

  408. 16:12

    Now what is hard is I would say

  409. 16:13

    something about like let's say this

  410. 16:14

    task.

  411. 16:17

    If I ask you before seeing this video is

  412. 16:19

    this task doable without having hands

  413. 16:22

    very likely half the people will say no

  414. 16:24

    because it requires putting like how

  415. 16:26

    many of you have lost airports? Uh like

  416. 16:29

    and in our company there's a channel

  417. 16:31

    called right airpod because people keep

  418. 16:32

    losing their right. Now in this case the

  419. 16:35

    robot does not have hand it has gripper.

  420. 16:37

    So here the task is very hard for the

  421. 16:40

    gripper. So it requires a higher level

  422. 16:42

    of intelligence because the arm has to

  423. 16:44

    go and orient itself to pick up the

  424. 16:47

    airpod in the right manner such that it

  425. 16:50

    can be inserted because the grippers can

  426. 16:52

    only close up and like like this.

  427. 16:54

    They're parallejo grippers. So you

  428. 16:56

    cannot turn the uh the the the airpod at

  429. 16:59

    the very end. Now what you are seeing

  430. 17:00

    here these are not real airpods. They

  431. 17:02

    are fake ones from teu uh like $5 $10

  432. 17:05

    each. So they don't have a magnet

  433. 17:07

    inside. So the robot has to really work

  434. 17:09

    hard to put this inside properly because

  435. 17:11

    there is no magnet to pull it uh easily.

  436. 17:13

    So these are all fake uh fake ones. This

  437. 17:15

    is even harder than it appears uh in the

  438. 17:17

    video. You can also uh like once you

  439. 17:21

    build the general brain behind the

  440. 17:23

    scene, you can also go and learn it from

  441. 17:25

    variety of just human videos without

  442. 17:26

    actually having any finetuning data at

  443. 17:28

    tell operation time. So in this scenario

  444. 17:30

    what we did, we trained on human videos

  445. 17:32

    like this like egocentric videos. we are

  446. 17:34

    now trying to transfer it to more third

  447. 17:36

    person uh uh videos and if third person

  448. 17:39

    works you can learn from YouTube any any

  449. 17:41

    kind of open source data. So in this

  450. 17:43

    case we see the video and then we add

  451. 17:46

    less than 1 hour of robot data. So very

  452. 17:48

    quick uh transformation and then the it

  453. 17:50

    can transfer to humanoids. Now these

  454. 17:53

    ones have hands. Hands or no hands is a

  455. 17:56

    big debate. People often use hands as an

  456. 17:58

    excuse as to why robotics is not here

  457. 18:01

    but that's not the case. It's always the

  458. 18:03

    intelligence uh behind the scene.

  459. 18:07

    We don't deploy hands right now because

  460. 18:09

    there are none available which can be

  461. 18:10

    deployed in factories. They all break

  462. 18:12

    within 100 yards. Uh pick anyone.

  463. 18:17

    If you have better hand uh I would love

  464. 18:19

    to buy 10 units ASAP to test. So it's

  465. 18:23

    robust to scenarios. Now the idea is you

  466. 18:26

    can do many tasks by watching humans.

  467. 18:28

    And the reason we can work with very

  468. 18:30

    less data which is less than one hour of

  469. 18:32

    robot data is because robot imagines

  470. 18:34

    things in its own head and tries to

  471. 18:36

    multiply the learning from for many

  472. 18:38

    scenarios. So what you see here is a

  473. 18:40

    completely fake video uh you can call it

  474. 18:42

    robot's dream. So it's inside the

  475. 18:44

    robot's own model where it can imagine

  476. 18:46

    scenarios uh in so this is not real

  477. 18:48

    data. This is all fake uh from own

  478. 18:51

    robot's own model. You can transfer it

  479. 18:53

    to you know more complex task and even

  480. 18:56

    lower cost hardware. So this is a egg uh

  481. 18:59

    uh egg making like omelette making task.

  482. 19:02

    Uh now here the these this whole setup

  483. 19:05

    costs about $4,000. So extremely cheap

  484. 19:07

    arms uh compared to the uh what you see

  485. 19:10

    out there. And if you notice here this

  486. 19:14

    setup is too cheap to even have any

  487. 19:16

    sensors. So the only sensor here is just

  488. 19:18

    a camera. No force nothing else. And the

  489. 19:20

    robot can do task which require forces

  490. 19:23

    uh from vision. Now I'm not saying

  491. 19:25

    that's a future like you should not be

  492. 19:27

    doing this. I'm sure the sensors will

  493. 19:29

    improve but

  494. 19:31

    >> from the existing sensors we are way

  495. 19:33

    behind than where we can be from just

  496. 19:36

    intelligence perspective.

  497. 19:37

    >> Okay. So this can keep on going. This is

  498. 19:40

    my adviser from Berkeley. He did not

  499. 19:42

    believe the robots can do it. He just

  500. 19:44

    kept standing for like uh 1 hour and the

  501. 19:46

    robot kept making omelette. Uh and and

  502. 19:48

    the interesting part here is this is

  503. 19:50

    fully end to end system. There is no

  504. 19:51

    state machine nothing. So the reason

  505. 19:54

    robot is looking egg because there is an

  506. 19:55

    empty plate in the front. I don't know

  507. 19:57

    why it's fluctuating. It's not in the

  508. 19:58

    there's no cut in the video. I I assure

  509. 20:00

    you of that. It's HDMI problem. So as

  510. 20:03

    soon as you remove the plate uh sorry I

  511. 20:06

    don't know why this is happening.

  512. 20:08

    Yeah. So there's no state machine when

  513. 20:10

    the robot starts when the robot ends.

  514. 20:11

    It's all automated and it's very robust

  515. 20:13

    to even disturbances. You can change

  516. 20:15

    objects around. Uh you can add these are

  517. 20:18

    all unseen and unseen objects. Uh only

  518. 20:20

    the pan and the gas stove is same.

  519. 20:22

    Everything else is different and the

  520. 20:24

    robot can keep on going. So you can get

  521. 20:26

    basic robustness. So this was this is

  522. 20:28

    all trained with less than 10 hours of

  523. 20:30

    data. Uh so it's very low data to be

  524. 20:32

    learning this robustness and it's coming

  525. 20:34

    from the base model behind the scene.

  526. 20:37

    I'm skipping uh this in the interest of

  527. 20:39

    time. So unlike traditional robotics

  528. 20:42

    pipelines where you have planning,

  529. 20:43

    mapping etc. This is an end toend brain.

  530. 20:46

    Now end to end is heavily used term in

  531. 20:49

    in many areas of AI. But when I mean end

  532. 20:51

    to end, I really mean end to end. It

  533. 20:53

    reads directly from the cameras and it

  534. 20:55

    applies power directly to motors. So we

  535. 20:59

    use nothing uh in between except just a

  536. 21:02

    P at the very end to go from 100 to 500

  537. 21:04

    Hz. But P is not robotics contribution.

  538. 21:06

    This is before it predates to like World

  539. 21:08

    War II or something. Okay. Now the the

  540. 21:11

    the interesting part is vision for us is

  541. 21:14

    yet another input. So nothing that crazy

  542. 21:16

    about it. So if you have seen, you may

  543. 21:19

    have seen a lot of lot of results of

  544. 21:20

    robots dancing, doing karate, kung fu,

  545. 21:22

    backflip.

  546. 21:24

    Go back and think how many results have

  547. 21:27

    you seen of robots going up and down

  548. 21:29

    stairs. Very few. Like even the top

  549. 21:32

    companies out there in humanoid, they

  550. 21:34

    have not shown beyond one sample stair

  551. 21:37

    just to check the tick tick box. You do.

  552. 21:39

    And people talk about humanoids are

  553. 21:41

    coming, China versus US, all of this is

  554. 21:44

    sort of BS. If humanoids cannot climb

  555. 21:46

    stairs, what's the point of having legs?

  556. 21:49

    All right, that's the only reason why

  557. 21:50

    you have legs. Now, this is actually a

  558. 21:52

    Moex paradox add action. Left one

  559. 21:56

    actually is much easier for a robot. A

  560. 21:58

    backflip, dancing, etc. is much easier

  561. 22:00

    for the robot. And you may think, okay,

  562. 22:02

    why is this paradox exist? Think one

  563. 22:04

    level deeper. When a robot is doing a

  564. 22:07

    back flip, it has to only know about its

  565. 22:10

    own body, nothing else. Right? So it's

  566. 22:14

    and when everything is known or fully

  567. 22:16

    observed that's what computers are good

  568. 22:19

    at because they can do they can simulate

  569. 22:22

    every possible uh uh setup but the right

  570. 22:25

    one even simply climbing on stairs

  571. 22:27

    requires seeing the stair at the first

  572. 22:29

    at the first place or what's the height

  573. 22:31

    what's the width you don't measure

  574. 22:33

    height and width exactly but you adjust

  575. 22:34

    to all the disturbances and that

  576. 22:36

    requires vision so right one requires

  577. 22:38

    understanding and that's why it is hard

  578. 22:39

    and you don't see any of this but for us

  579. 22:41

    it's just get another result it doesn't

  580. 22:42

    really matter so we And uh we put this

  581. 22:44

    result out one and a half year ago. We

  582. 22:46

    shot this like two and a half years ago.

  583. 22:48

    Uh but uh like this robot can go any

  584. 22:51

    scenario. It is not stumbling on things.

  585. 22:53

    It intentionally steps over things and

  586. 22:55

    there is no mapping or planning. It

  587. 22:57

    doesn't make any 3D map of the

  588. 22:58

    surroundings. It's all operating from

  589. 23:00

    camera on the torso. And you can put

  590. 23:02

    this in any kind of setup, any kind of

  591. 23:04

    stairs. Uh I'm going faster here. These

  592. 23:07

    are, you know, uh fire escape stairs.

  593. 23:11

    It's a bit weird. Fire escape stairs are

  594. 23:13

    to be used when there is fire in urgency

  595. 23:15

    and they are the hardest in every

  596. 23:16

    building. Uh so we are testing in all in

  597. 23:19

    all those setups but you can put this in

  598. 23:21

    anywhere you can disturb it and stay

  599. 23:22

    doesn't really matter. This is

  600. 23:23

    superhuman capability because robot

  601. 23:25

    cannot see it's being pulled. So it's a

  602. 23:28

    surprise factor for the robot that

  603. 23:29

    you're pulling the violet while

  604. 23:30

    climbing. And again it's the same model

  605. 23:32

    everywhere in all these scenarios. So

  606. 23:35

    very easy although easy parkour does

  607. 23:37

    look fun uh uh does look fun and people

  608. 23:40

    often find like uh this is uh you know

  609. 23:45

    the robots can do all this but this is

  610. 23:48

    so much easier than what I showed

  611. 23:50

    earlier and even if you see these videos

  612. 23:53

    right now many of you will find this

  613. 23:55

    more impressive even though I'm telling

  614. 23:56

    you this is easy it takes half an hour

  615. 23:58

    to train this while the previous one it

  616. 24:00

    takes much longer uh and much more

  617. 24:01

    difficult to train

  618. 24:03

    so you can take this Based model and you

  619. 24:05

    can transfer to variety of tasks very

  620. 24:07

    quickly. So here this is a very old

  621. 24:09

    result from one and a half year ago uh

  622. 24:11

    for inserting LAN cables etc. We have

  623. 24:13

    very advanced systems now but we are

  624. 24:15

    deploying these models already across

  625. 24:17

    variety of application. So for instance

  626. 24:19

    because to to build the data flywhe the

  627. 24:21

    hardest part is the deployment itself.

  628. 24:24

    There are so many other issues beside

  629. 24:25

    intelligence and hardware that come up

  630. 24:27

    when you deploy robotics. It is not the

  631. 24:29

    same. This is not same same as deploying

  632. 24:32

    chat GP over an app because you have to

  633. 24:35

    work around many of the and I think

  634. 24:36

    Skyio gave a talk before this and they

  635. 24:39

    can talk to you about all day about the

  636. 24:40

    hardness of deployment. So we have to

  637. 24:42

    start deployment now so that we can have

  638. 24:45

    this data flywheel in a foreseeable

  639. 24:46

    future and I'll give you a few examples

  640. 24:48

    of deployment we are doing. So one of

  641. 24:50

    the examples with Nvidia like Nvidia is

  642. 24:52

    opening their first factory in Houston.

  643. 24:54

    uh and their GPUs what you use right now

  644. 24:56

    they're all built outside US in in

  645. 24:58

    Taiwan mainly and it's all done manually

  646. 25:00

    over there and uh humos are really

  647. 25:02

    efficient and really uh really good at

  648. 25:04

    making these things but sustaining cost

  649. 25:07

    and scalability and throughput here in

  650. 25:08

    US it's very hard to maintain without

  651. 25:10

    this labor force so here we are

  652. 25:12

    automating uh this GPU assembly for them

  653. 25:15

    this was we did a live demo in Nvidia

  654. 25:18

    GTC this is deployed in factory already

  655. 25:20

    uh last week so it's already live but

  656. 25:23

    the interesting part I want to highlight

  657. 25:25

    this is a very traditional factory task

  658. 25:28

    but if you look at the the whole setup

  659. 25:32

    this is extremely randomized extremely

  660. 25:34

    noisy and you go to any factory

  661. 25:37

    traditional setup they're extremely

  662. 25:39

    clean so here the robot the same brain

  663. 25:42

    can not only do very precise task it can

  664. 25:44

    be very robust to any kind of

  665. 25:46

    disturbance which means now these robots

  666. 25:48

    do not need to be in a cage uh and they

  667. 25:51

    can be uh just deployed as is on on

  668. 25:53

    these factory lines without any change

  669. 25:55

    to any factory line alongside humans and

  670. 25:58

    wherever there is humans there is mess.

  671. 26:00

    Okay, you can put the same brain to

  672. 26:01

    delivery applications. So here the robot

  673. 26:03

    delivers from the from the truck to the

  674. 26:05

    to the door. It has to find where the

  675. 26:07

    front door is. So it's a common sense

  676. 26:09

    problem. It's not mobility problem. And

  677. 26:10

    we already have been delivering uh

  678. 26:12

    packages to with working behind the

  679. 26:14

    scene with many partners to delivering

  680. 26:16

    packages to front door of the houses uh

  681. 26:17

    in these areas. same setup works across

  682. 26:20

    warehouses and all this. Um now just to

  683. 26:24

    close the talk uh close the topic I

  684. 26:27

    mentioned this is an omnibody brain and

  685. 26:29

    I hope it's very clear why it's

  686. 26:30

    extremely essential to have an omnibody

  687. 26:32

    brain to build data flywheel but there's

  688. 26:34

    one extra benefit and the benefit is

  689. 26:36

    your robots may break over time and uh

  690. 26:39

    or the safety uh scenarios and safety is

  691. 26:42

    a byproduct of this omnibody brain

  692. 26:44

    because if you have a we can put the I'm

  693. 26:46

    just skipping very fast here so you can

  694. 26:47

    see the online this is all online uh but

  695. 26:49

    we can put the same brain across each of

  696. 26:51

    these systems each of these robots

  697. 26:55

    and in this case we did not even train

  698. 26:57

    on these robots. This is completely zero

  699. 26:59

    shot transfer to all these humanoids and

  700. 27:01

    all these things and even if the robot

  701. 27:03

    breaks like here the leg breaks it can

  702. 27:06

    recover within milliseconds because when

  703. 27:08

    your leg breaks your three-legged robot

  704. 27:10

    is a new robot. So it doesn't really

  705. 27:12

    matter what's the shape of the robot is

  706. 27:14

    anything changing here we disabled the

  707. 27:16

    legs of the robot it learns to learn it

  708. 27:18

    learns to walk on two legs in just three

  709. 27:20

    trials. So what you are seeing on screen

  710. 27:22

    is all the training for the robot that

  711. 27:24

    they are that that is happening in these

  712. 27:26

    systems. So it's it adapts in few

  713. 27:28

    milliseconds to uh uh 30 seconds or so.

  714. 27:32

    And like one example here it's it's

  715. 27:33

    going on wheels. You jam the wheel it

  716. 27:36

    starts walking. So this is a another

  717. 27:38

    byproduct of omnibody brain. Safety

  718. 27:42

    takes a very different meaning uh with

  719. 27:44

    these kind of models. when the robot can

  720. 27:46

    fly like or sorry when the robots can

  721. 27:48

    walk or or operate with only half the

  722. 27:50

    body. I removed some of the gory scenes

  723. 27:52

    from this but we also cut the robot in

  724. 27:53

    half and it can still work uh with both

  725. 27:56

    the two halves separately but for that

  726. 27:57

    you can go to YouTube uh and that's all

  727. 27:59

    I have. Thank you.