AI Engineer World's Fair 2026

World Models Need Causality, Not Pretty Pixels — Christopher Manning, Moonlake AI

Read the talk

World Models Need Causality, Not Pretty Pixels

Christopher Manning traces the history of AI and language models, then develops Moonlake AI’s approach to physical intelligence: reconstruct a world from observations, give its objects behavior in code, and refine the simulation against reality.

From a talk by Christopher Manning

At a glance

Ideas worth remembering

  • Language-model progress needed model flexibility alongside data and compute; substantial text scale existed before architectures could use it with today’s breadth.

  • An action-conditioned world model represents state and predicts how actions change it. Attractive generated observations alone do not establish that capability.

  • The tea-box example adds capability in layers: separate movable objects, reconstruct hidden contents using retrieved information, then make those contents independently manipulable.

  • Moonlake combines generated code, textures and physics models, then proposes refining simulations through comparisons with real observations and behavior.

  • Simulation fidelity should follow the intended task. Simulated training and discovery remain useful only insofar as the model captures the real-world details that matter.

AI began with language, feedback and machines that could act

Christopher Manning opens with a goal for Moonlake AI: simulation infrastructure for practical physical AI. Embodied intelligence means an intelligence that can operate in an environment, rather than only describe it. But the route to that goal begins with a deliberately slow historical build. Language, control and robots have been part of AI’s story from the beginning.

Source frame: AI began with language, feedback and machines that could act
Source frame: AI began with language, feedback and machines that could act

The 1956 Dartmouth summer research project supplies the familiar starting point: John McCarthy coined the term artificial intelligence and gathered researchers including Claude Shannon. Earlier cybernetics work had already connected communication, control and feedback in living things and computers. That tradition matters here because acting intelligently requires a loop: observe an environment, act on it, and use the resulting change to decide what comes next.

Natural language processing also predates the name AI. Manning revisits a 1954 public demonstration of Russian-to-English machine translation, accompanied by predictions that computers would replace most human translators. The demonstration establishes an early ambition that would take decades to mature: language processing as useful machine intelligence, rather than a peripheral application.

McCarthy founded the Stanford Artificial Intelligence Lab starting in 1963. Although his own work centered on mathematical logic, he supported a wider effort to build embodied intelligence. The Stanford Cart and, more directly relevant to this talk, Shakey brought perception and action together. Shakey could perceive its environment, move around and move boxes. Understanding the world was already tied to doing something in it.

0:210:28
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:21 · section reference included

Language models were useful long before they became the center of AI

A playful Stanford-versus-Berkeley detour places this history in its institutional setting. Manning emphasizes Stanford’s early AI work and network connections, while acknowledging Berkeley’s important systems work and its later growth in AI. He then returns to his own field: natural language processing, where language models had long been a central technology even when much of the rest of AI paid them little attention.

Source frame: Language models were useful long before they became the center of AI
Source frame: Language models were useful long before they became the center of AI

The lineage runs from character-level models of language to Shannon’s word and character n-gram models in the late 1940s, then to probabilistic text models developed at IBM in the 1970s. These models gave speech and NLP systems a way to judge which sequences of words were plausible. Their value was practical: they helped power speech recognition, spelling correction and machine translation, including Google’s translation work around 2007.

Useful did not yet mean general. In Manning’s recollection, language models were components for particular tasks; researchers still expected broader intelligence to require separate memories, knowledge representations, planning systems and reasoning systems. That expectation makes the later rise of language models surprising even to someone who spent his career building them.

8:318:34
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

8:07 · section reference included

Data scale needed a flexible model

Manning’s language-model history separates three ingredients that are easy to collapse into one story: data, compute and model flexibility. An early neural language model around 2000 used a 32 million-token corpus and a 31,000-word vocabulary. Compute constrained what it could do. By contrast, he cites Google’s 2007 language model built on two trillion tokens: substantial data scale existed well before today’s neural models. 14:58

Source frame: Data scale needed a flexible model
Source frame: Data scale needed a flexible model

The missing ingredient in that large earlier model was flexibility. More text could not supply the expressive power of neural architectures that had not yet arrived. Transformer-based language models appeared in 2018, but early versions again used comparatively small datasets. In this account, the ingredients came together with GPT-3 in 2020 and subsequent models: enough data and compute, joined to an architecture capable of using them.

The result was a surprising victory for NLP. Language models moved from a relatively marginal part of AI to the technology many people now mean when they say AI. Manning’s surprise is personal: after 30 years in the field, he still finds it remarkable that a model can work through difficult mathematics using test-time thinking and reach answers he could not produce himself. That success sets a demanding starting point for the next question: what does intelligence still need to act in the physical world?

Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

14:11 · section reference included

A world model predicts what an action changes

Text-based descriptions have supported far more intelligence than many researchers expected. Physical intelligence adds another demand: an agent must operate in the world around it. One route is to learn directly from hardware experiments, recorded robot behavior or real-world video. Manning points to roughly 10,000 hours of human teleoperation as an example of the data collection burden. Robots move at physical speed, and people must spend time guiding them.

Source frame: A world model predicts what an action changes
Source frame: A world model predicts what an action changes

Shakey supplies an older alternative: keep an internal representation of the world and use it to consider actions before executing them. Its representation recorded locations and object attributes in a logical grid world. The modern version keeps the same purpose while seeking richer environments. An action-conditioned world model starts from a representation of the current state and predicts the state that will result from a chosen action.

The important distinction is between an observation and the state behind it. A picture shows what a camera can see. A semantic state represents things the agent needs to reason about: objects, their attributes and how actions affect them. Planning needs predictions about those changes, so the quality of generated images alone is an insufficient test.

A fluid Genie 3 generative-video example by Riley Goodside makes the distinction visible. Manning credits its visual appeal, then criticizes the demonstrated approach for relying on simulated observations without the semantics needed for dependable planning. This is his assessment of the demonstrated approach, rather than a measured comparison of planning performance. Moonlake’s proposed advantage is causal structure: a simulator should explain how acting on something produces a change that an agent can anticipate. 23:56

Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

17:39 · section reference included

From a room you can view to a tea box you can open

Moonlake’s starting input is an image or a short video: a partial observation of an underlying world. The desired output is a model in which an agent can act and inspect the consequences. A room reconstruction provides a concrete test. Manning describes a Marble reconstruction that supports walking around and seeing the scene from different angles. For that task, it works well. But the room also contains a kettle, a box and a cup, and the demonstrated reconstruction does not let the user manipulate them.

Source frame: From a room you can view to a tea box you can open
Source frame: From a room you can view to a tea box you can open

The first change is to separate the scene into background and foreground objects. The background can retain a representation suited to viewing the environment. Foreground objects become separate things that can move. The observable difference is simple: the tea box stops being part of the room’s appearance and becomes an object the agent can act on.

Opening the box exposes a harder problem. The original photo shows it closed, so its contents cannot be recovered from visible pixels alone. Moonlake adds web retrieval in a retrieval-augmented fashion: find images of the product open, read its description and obtain dimensions. That information supports a reconstruction with tea bags inside. It describes the product’s expected contents; it cannot establish which contents are actually present in this particular closed box. 29:04

The next change is equally important: the contents must become objects too. Tea bags that merely appear inside the box do not yet support the task. They need their own size and manipulable representation so an agent can lift one and take it out. Each step adds a capability required by the intended action, rather than indiscriminately adding detail to the whole room.

What changes between viewing the tea box and using it? The diagram follows the added representations. A movable box needs a separate identity; an openable box needs hidden structure; removing a tea bag needs the contents to have their own identities and behavior. Visual reconstruction supplies the scene, while these additions supply the actions.

How it fits togetherAdding the structure needed to use the tea box

The observation shows the exterior and hides the contents.

The same scene gains progressively richer actions as the box and its contents become separately represented objects.

Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

25:19 · section reference included

Code makes the reconstructed world editable

The controllable parts of Moonlake’s worlds have code underneath them. This uses the ability of language models to work with symbolic material—including mathematics and code—as well as natural language. Generated code supplies an editable representation of objects and behavior. Neural generation and symbolic control meet in what Manning calls a neurosymbolic world model.

Source frame: Code makes the reconstructed world editable
Source frame: Code makes the reconstructed world editable

The construction process uses an iterative coding loop. Code renders objects, diffusion-based generation supplies textures, and the rendered result is assessed against physical reality. Differences guide revisions to the code, producing a new render for the next comparison. The practical advantage of the symbolic layer is that people and other applications can interface with it, control it, edit it and maintain it. 31:40

How does a reconstruction improve after its first attempt? The loop below makes the feedback path explicit: reality provides a comparison target, the comparison identifies deviations, and code revisions change the next simulated result. This is the mechanism Moonlake proposes for improving its generated worlds, rather than treating the first reconstruction as final.

Manning extends the idea that software will eat the world to physical processes. Software already manages records, suppliers and other virtual representations of activity. His hope is that verifiable simulations powered by code will let it reach further into the physical activity itself. The link is prediction: a useful simulation gives an agent a place to test what an action will do before it acts outside the computer.

How it fits togetherRefining a generated simulation

Code represents the controllable parts of the world.

The comparison feeds back into code changes, which produce the next result to assess.

Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

30:50 · section reference included

Model the details the task needs

The tea example continues beyond opening the box. Making tea calls for removing a bag from its foil container, putting it in a cup, boiling water and pouring it. These actions demand additional structure and behavior. A reconstruction adequate for moving a closed box may be inadequate for handling its packaging or liquid. Simulation quality therefore depends on the task that must transfer to reality.

Source frame: Model the details the task needs
Source frame: Model the details the task needs

A simulator need not reproduce every detail of the world. Human world models also leave most things unmodeled while representing what matters to the current purpose. The engineering choice is where to spend fidelity: include the objects, properties and dynamics that affect the intended action, and retain control over which parts of the world receive that detail.

The closing application is an industrial process, later identified in Q&A as a conveyor belt system. A short video becomes a 3D simulation with objects and movement represented in sufficient detail to generate training data. A robotic system can then explore the simulated environment and learn how to operate in it. This is the proposed replacement for collecting 10,000 hours of real-world teleoperation.

Manning describes obtaining 10,000 hours of simulation “for free.” The useful economic claim is that repeated simulated experience can avoid the corresponding human teleoperation effort. The presentation does not quantify simulation compute costs or establish measured real-world transfer performance; cheap, effective training remains the intended payoff of building a sufficiently accurate model.

34:1534:19
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

34:06 · section reference included

Games, symbolic interfaces and the long tail of situations

The first questions extend the approach beyond the industrial example. Several applications and representation choices share the same need for an agent to practice actions in a modeled environment:

  • Games: Millions of players provide abundant behavior data, and their goals are often fairly clear. A reinforcement learning loop can use those goals to learn effective actions. Moonlake has explored gaming, although Manning says its recent emphasis is physical infrastructure.
  • Symbolic interfaces: Code brings back some of the function of older knowledge representations without requiring a return to old-style ontologies. A purely neural latent representation may be powerful, but people and other applications have difficulty connecting to it. Moonlake’s bet is that symbolic representations offer a practical interface for physical AI.
  • Rare situations: Humanoid robotics, space robotics and other new forms of automation need experience beyond routine cases. Collecting that experience physically is costly, especially in the long tail of unusual scenarios. Good simulation provides a way to explore those cases.

The symbolic choice is a practical bet about integration, rather than a claim that neural representations cannot work. A code-based world gives neural systems something they can reason over and plan with, while giving humans an accessible way to change the model. The older ambition of knowledge representation returns in executable form.

38:4438:57
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

38:09 · section reference included

Physics supplies dynamics; reality tests the model

Does the simulation learn all physical behavior from video? Manning answers that it uses physics engines and knowledge of physics to generate and control movement. Water is his example: starting from an image does not mean the system must infer all fluid behavior from that image. A physics model supplies predictions about how it will behave. The neurosymbolic approach therefore combines generated representations with established physical modeling.

The next question presses on the simulation-to-reality gap: mechanical details and tolerances can make real behavior diverge from a simulation. Moonlake’s proposed response extends the earlier coding loop from appearance to behavior. Compare simulated behavior against real-world video, then use neural optimization to improve the simulation. The aim is to shrink discrepancies automatically rather than depend entirely on people hand-writing and adjusting a physics simulation. 45:52

Physical dynamics are only one layer of a working environment. An audience question asks about modeling people, processes and the decision makers who control objects. Manning says Moonlake has not really been addressing that sociotechnical layer. He sees potential in language models as human-behavior simulators because they absorb large amounts of human behavior data, but presents that as a neighboring area of development, rather than a capability established by the physical-world examples.

The final question asks whether exploration in latent space can reveal connections humans have not already identified. A sufficiently good simulation can support surprising discoveries, Manning replies, but correctness inside the simulated world does not ensure that every relevant fact about reality has been captured. Reality can also contain relationships absent from the model. The value of simulation is wider exploration with a useful approximation; its discoveries still depend on what the approximation represents. 49:16

Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

44:26 · section reference included

Resources

From the talk

  • Revisit Riley Goodside’s generative-video example and the specific distinction Manning draws between visual observations and semantic state for planning.

  • The company behind the simulation infrastructure described here; a starting point for investigating its work on controllable worlds for physical AI.

Read the complete timestamped transcript
  1. 0:01

    [music]

  2. 0:12

    Okay. Hi everyone. [applause]

  3. 0:16

    Um, and thanks a lot for making it down

  4. 0:19

    to the second floor and finding my

  5. 0:21

    session here. Okay. So I'm Chris Manning

  6. 0:24

    and what I want to do today is present

  7. 0:26

    something about Moonlakes's approach to

  8. 0:28

    producing a simulation infrastructure

  9. 0:30

    for practical physical AI. The overall

  10. 0:33

    goal here is that a north star for AI

  11. 0:37

    and actually cognitive science as well

  12. 0:39

    has always been to understand and work

  13. 0:42

    out how to build embodied intelligence.

  14. 0:45

    So today I'm going to tell us about some

  15. 0:47

    of the recent work that we've been doing

  16. 0:49

    at approaching a practical form of

  17. 0:52

    embodied artificial general

  18. 0:53

    intelligence.

  19. 0:55

    But you know I'm not actually going to

  20. 0:58

    start there because uh you know um Swick

  21. 1:03

    said no you shouldn't just do that. Um,

  22. 1:07

    you should have a double slot. And first

  23. 1:10

    of all, since I've got you here today, I

  24. 1:13

    should have people tell you about the

  25. 1:15

    history. You should tell people about

  26. 1:16

    the history of AI and your journey

  27. 1:19

    through it and how you ended up here.

  28. 1:21

    Um, so, uh, sit back, um, get ready for

  29. 1:26

    story time and we're going to have the

  30. 1:29

    really slow build and we're going to

  31. 1:31

    hear about all about the history of AI

  32. 1:33

    for 20 minutes and then we're going to

  33. 1:35

    hear the details about what Moon Lake is

  34. 1:38

    doing. Um, he looks very persuasive

  35. 1:41

    there, doesn't he? Whereas at the time

  36. 1:43

    we were having this conversation, I was

  37. 1:45

    having a very bad hair day. So I had

  38. 1:48

    very little um choice um to agree to

  39. 1:51

    this task. Um so here I am with some

  40. 1:54

    long-term remarks on how did I and all

  41. 1:57

    of us get to this moment from the far

  42. 2:00

    ago days um when AI didn't work. Um so

  43. 2:04

    the very beginning of AI was the

  44. 2:07

    Dartmouth summer research project in

  45. 2:10

    1956. It was for the holding of this um

  46. 2:13

    kind of summer group project was when um

  47. 2:18

    John McCarthy coined the term AI and got

  48. 2:22

    together um this group of people. Um so

  49. 2:25

    that one on the back right there, that's

  50. 2:28

    John McCarthy, there's Minsky and both

  51. 2:31

    most of them look like a real bunch of

  52. 2:33

    geeks as you can see. Um but the

  53. 2:35

    attractive one on right on the right end

  54. 2:38

    is you know my personal hero as more of

  55. 2:40

    a language guy um Claude Shannon.

  56. 2:43

    So this is sort of the start of AI 1956.

  57. 2:48

    It's where the the term AI came from.

  58. 2:50

    But really there's other stuff that came

  59. 2:53

    earlier than 1956.

  60. 2:56

    So really starting in the 40s and early

  61. 2:58

    in the 50s there was work on

  62. 3:00

    cybernetics. So cybernetics sought to

  63. 3:03

    tie together communications control and

  64. 3:06

    feedback in living things and computers.

  65. 3:08

    Um so uh you know kubernetes is really

  66. 3:12

    the same word as cybernetics just less

  67. 3:15

    anglicized um but has very different

  68. 3:18

    meaning. Um yeah so this led to the

  69. 3:20

    earliest work on neural networks and the

  70. 3:23

    perceptron network. So here's Frank

  71. 3:25

    Rosenlat's perceptron um that was you

  72. 3:29

    know predated the term artificial

  73. 3:31

    intelligence and as you can see like in

  74. 3:34

    those days you actually used to wire

  75. 3:36

    neural networks now this modern matrix

  76. 3:38

    multiplication but another thing that

  77. 3:41

    actually started before the term

  78. 3:44

    artificial intelligence and has been my

  79. 3:47

    long-term area of interest is doing

  80. 3:50

    things in natural language processing

  81. 3:52

    and natural language processing

  82. 3:54

    began as machine translation. And so

  83. 3:57

    here we are again in 1954,

  84. 4:01

    two years before the term artificial

  85. 4:03

    intelligence was coined. And here's the

  86. 4:06

    front page of the New York Times. Now,

  87. 4:09

    there's a way in which the front page of

  88. 4:10

    the New York Times in 1954 is

  89. 4:14

    surprisingly reminiscent of the issues

  90. 4:16

    that you see in the front page of the

  91. 4:18

    New York Times um in 2026 because it's

  92. 4:23

    exactly the same issue. President

  93. 4:25

    proposing ending citizenship for people.

  94. 4:28

    Um not much has changed in the um

  95. 4:31

    intervening 75 years. Um allegations of

  96. 4:35

    various kinds of behavior akin to

  97. 4:37

    treason. All sounds very familiar. Um,

  98. 4:40

    but this was the top half of the page.

  99. 4:41

    And if you went down to the bottom half

  100. 4:43

    of the page, um, you found this article.

  101. 4:46

    Russian is turned into English by a fast

  102. 4:49

    electronic, um, translator. A public

  103. 4:52

    demonstration of what's believed to be

  104. 4:54

    the first successful use of a machine to

  105. 4:56

    translate meaningful text from one

  106. 4:58

    language to another took place here

  107. 5:00

    yesterday afternoon. Um, and here's a

  108. 5:03

    little bit of video of showing that

  109. 5:05

    system.

  110. 5:11

    into

  111. 5:16

    one of the first nonmerica

  112. 5:26

    were made that the computer would

  113. 5:27

    replace most human translators.

  114. 5:32

    >> Okay. Um, and we'll come back to that

  115. 5:35

    again in a little bit. Um, but coming

  116. 5:38

    off of the founding of AI, um, by John

  117. 5:42

    McCarthy, fairly soon after that, John

  118. 5:45

    McCarthy moved to Stanford and founded

  119. 5:48

    Stanford Artificial Intelligence, the

  120. 5:50

    Stanford Artificial Intelligence Lab,

  121. 5:53

    starting from 1963.

  122. 5:56

    And the original Stanford AI lab was in

  123. 5:59

    this um, building up in the foothills.

  124. 6:01

    Um if you down if you know the South Bay

  125. 6:04

    well um if you've ever been to a

  126. 6:06

    restraero and you look over next to a

  127. 6:09

    restraero um where there's the Portola

  128. 6:12

    pastures um horse area or if you ride a

  129. 6:15

    horse um that's where this AI lab used

  130. 6:18

    to be and that was the site of many of

  131. 6:20

    the um founding um work in artificial

  132. 6:24

    intelligence and also a lot of other

  133. 6:26

    stuff. Um the Stanford AI lab was the

  134. 6:29

    where the very first video game

  135. 6:31

    tournament was played in 1972.

  136. 6:35

    Um now McCarthy himself um was um

  137. 6:40

    mathematical logician and nearly all of

  138. 6:43

    his work was sort of building out these

  139. 6:45

    not notions of mathematical logic but in

  140. 6:48

    his thinking he was very wide ranging

  141. 6:52

    and so he

  142. 6:55

    he liked the idea of how could we build

  143. 6:58

    an embodied artificial general

  144. 7:00

    intelligence and was very happy to sort

  145. 7:03

    of provide the facilities and the money

  146. 7:06

    for other people to start to explore

  147. 7:08

    this. Um so this was the Stanford AI

  148. 7:11

    Labs um very first robot um the Stanford

  149. 7:15

    cart not very fancy um the old um sale

  150. 7:19

    building um but more relevant to what

  151. 7:22

    we're going to talk about today for

  152. 7:25

    simulation infrastructures for AI um was

  153. 7:28

    another robot um Stanford SRRI shaky

  154. 7:32

    robot

  155. 7:33

    >> mach

  156. 7:37

    audio Um, so this was a robot that could

  157. 7:41

    um perceive the environment, move

  158. 7:44

    around, move boxes and so on in the

  159. 7:48

    environment and do things like that.

  160. 7:52

    Okay. And so this sort of started to

  161. 7:56

    explore this idea of having an embodied

  162. 8:00

    in artificial intelligence that had an

  163. 8:03

    understanding of a world and its

  164. 8:05

    environment.

  165. 8:07

    Um, just one little distraction that's

  166. 8:10

    not really about AI. Um, who here has a

  167. 8:13

    connection to Stanford? Stanford

  168. 8:16

    connections. Who here has a connection

  169. 8:19

    to Berkeley? Berkeley connections. Less

  170. 8:22

    Berkeley people. Anybody with a

  171. 8:23

    connection to both places? There are

  172. 8:25

    some people who've been to both of them.

  173. 8:27

    No. Okay. Um, so I just have to put in

  174. 8:31

    my advertisement for early Stanford. Um

  175. 8:34

    if you're an artificial intelligence

  176. 8:37

    person you know the story in the early

  177. 8:40

    days is all Stanford. I mean there you

  178. 8:43

    know there are two different stories I

  179. 8:45

    can tell right you know if you're in the

  180. 8:47

    history of west coast universities

  181. 8:50

    really in the first half of the 20th

  182. 8:52

    century all the prominent stuff was

  183. 8:55

    Berkeley and that Stanford was um this

  184. 8:58

    sort of pokey regional school but once

  185. 9:01

    you get to the second half of the 20th

  186. 9:04

    century which is the era of computers

  187. 9:06

    and AI um it all happened at Stanford um

  188. 9:12

    So that you know if you look at things

  189. 9:14

    like the beginnings of the internet, the

  190. 9:16

    Arpanet, um here it is in 1972

  191. 9:20

    and it connects up some UC campuses,

  192. 9:23

    Berke um UCLA and Santa Barbara and

  193. 9:27

    connects to Stanford, MIT, Harvard, but

  194. 9:30

    no Berkeley. Go ahead to 1980. Um

  195. 9:34

    Berkeley still isn't part of the

  196. 9:36

    internet. you know, it's made it as far

  197. 9:38

    as Hawaii and London and Berkeley is

  198. 9:40

    still not there. Um, and really in all

  199. 9:43

    of this period, there just wasn't AI at

  200. 9:47

    Berkeley. I mean, there was other stuff

  201. 9:49

    that went on at Berkeley, I should be

  202. 9:51

    fair. There was important systems work.

  203. 9:53

    There was ingress database and BSD Unix,

  204. 9:56

    which some of you probably remember if

  205. 9:58

    you have gray hair like me. Um, but you

  206. 10:01

    know, essentially through the 60s and

  207. 10:04

    70s and the first half of the 80s, there

  208. 10:07

    was just no AI at Berkeley. And so AI at

  209. 10:10

    Berkeley really only got underway in

  210. 10:12

    1986

  211. 10:14

    um when two fresh Stanford PhD grads um

  212. 10:17

    Jendra Malik and Stuart Russell um moved

  213. 10:21

    from Stanford um to Berkeley to take up

  214. 10:24

    professorships. Now of course that's 40

  215. 10:26

    years ago now. So, they've had they've

  216. 10:28

    had um AI for a while at Berkeley now,

  217. 10:30

    but not in the old days. Um, if you'd

  218. 10:33

    like to know more about this history of

  219. 10:35

    the old days, um, a couple of years ago,

  220. 10:38

    we actually put together, um, a story of

  221. 10:41

    the first 60 years of, um, AI at

  222. 10:43

    Stanford. Conveniently, um, the 60 years

  223. 10:46

    stopped just before the arrival of chat

  224. 10:49

    GPT and large language models. Um, and

  225. 10:53

    you can, um, find it on YouTube.

  226. 10:56

    Um, okay. So, that was my very brief

  227. 10:59

    potted history of AI. Um, what did I do

  228. 11:03

    for my life? Well, what I did for my

  229. 11:06

    life was to be a natural language

  230. 11:09

    processing person. And so, a natural

  231. 11:11

    language processing that was where we

  232. 11:14

    developed language models. And language

  233. 11:18

    models actually in some form go back a

  234. 11:23

    very very long way. So the first

  235. 11:26

    language model was proposed by Andre

  236. 11:29

    Marov who invented Markoff models right.

  237. 11:32

    So in developing the idea of Markoff

  238. 11:35

    models he actually did it with language

  239. 11:38

    um taking a novel of Pushkin's Eugene

  240. 11:41

    Anagen and developed a character level

  241. 11:44

    language model over that. Um but the

  242. 11:47

    famous thing is then back to Claude

  243. 11:49

    Shannon who we saw earlier who sort of

  244. 11:51

    formalized information theory and

  245. 11:54

    started to build um word and character

  246. 11:57

    engram language models in the late 1940s

  247. 12:01

    which is the dominant stuff that we went

  248. 12:03

    on and using and in particular um then

  249. 12:06

    in 1975

  250. 12:08

    again so back in the sort of early part

  251. 12:11

    of AI um a famous group at IBM Fred

  252. 12:15

    Gelan's group at IBM sort of defined

  253. 12:18

    well they came up with the term language

  254. 12:20

    model that we still use today and they

  255. 12:23

    defined this idea of probabilistic

  256. 12:26

    models of text that could be used as a

  257. 12:29

    basis for all kinds of speech and

  258. 12:31

    natural language processing

  259. 12:33

    applications. So really from very early

  260. 12:36

    on in the speech and NLP tradition,

  261. 12:39

    language models were seen as a central

  262. 12:41

    technology that other people in other

  263. 12:43

    areas of AI and um machine learning just

  264. 12:48

    didn't really know about. And it was

  265. 12:50

    what enabled

  266. 12:53

    um interesting good things to happen

  267. 12:55

    early in speech and various areas of

  268. 12:59

    NLP. So both early speech recognition

  269. 13:02

    systems but also the kind of um machine

  270. 13:06

    translation that you got at Google

  271. 13:09

    starting about 2007. It was considerably

  272. 13:12

    powered by language models. And so I

  273. 13:16

    think there's actually a kind of an

  274. 13:17

    interesting story here of the

  275. 13:19

    development of our modern AI

  276. 13:23

    um sorry which is um slightly different

  277. 13:26

    to the story most people remember which

  278. 13:30

    um is a story that's dominated by the

  279. 13:33

    vision story because the vision sort of

  280. 13:35

    came in most people's head to be seen as

  281. 13:39

    the sort of entry place of neural

  282. 13:41

    networks in a large scale. Um but

  283. 13:44

    nevertheless I'll point out the

  284. 13:45

    limitation of this early work in

  285. 13:47

    language models. It was used for

  286. 13:49

    spelling correction, machine

  287. 13:50

    translation, all of these things. But at

  288. 13:53

    this time, nobody thought of language

  289. 13:56

    models as these were going to solve

  290. 13:57

    artificial intelligence. We still all

  291. 14:00

    bought the old AI story of that we're

  292. 14:02

    going to need memories, knowledge

  293. 14:04

    representations, planning systems,

  294. 14:06

    reasoning systems, all of these classic

  295. 14:08

    AI things.

  296. 14:11

    Yeah. So here's the sort of history of

  297. 14:13

    large language models from an NLP

  298. 14:16

    suspect perspective. So the very first

  299. 14:18

    mention of the term large language

  300. 14:20

    models was in 1998 as far as I can tell.

  301. 14:24

    But you know really that was sort of a

  302. 14:26

    how are we going to store all this text

  303. 14:29

    um to build a model. So the first

  304. 14:31

    interesting um connection of large

  305. 14:33

    language models was in 2000 when Joshua

  306. 14:36

    Benjio and colleagues defined new

  307. 14:39

    probabilistic language models. the first

  308. 14:41

    neural language model. Um, but as you

  309. 14:44

    can see from those stats of 32 million

  310. 14:47

    um token corpus, 31,000word vocabulary,

  311. 14:50

    this was a teeny model because at that

  312. 14:53

    point they just didn't have enough

  313. 14:55

    compute to do anything interesting. Um,

  314. 14:58

    but interestingly as early as 2007

  315. 15:02

    at Google they were able to solve the

  316. 15:05

    compute problem and the data problem.

  317. 15:08

    Um, so most people forget this now, but

  318. 15:11

    in 2007,

  319. 15:13

    um, Google had a language model that was

  320. 15:16

    built on two trillion tokens of text.

  321. 15:19

    So, you know, that's a bit smaller than

  322. 15:22

    the state-of-the-art models now that

  323. 15:23

    might be trained on 15 trillions of

  324. 15:25

    text, but it's actually the same order

  325. 15:27

    of magnitude, right? So we all already

  326. 15:30

    in sort of two decades ago there was the

  327. 15:34

    scale of data and compute to build a

  328. 15:37

    form of language model. The big problem

  329. 15:39

    was they didn't have enough model

  330. 15:42

    flexibility that the kind of modern

  331. 15:44

    powerful neural networks that we use

  332. 15:46

    today hadn't yet been invented.

  333. 15:48

    Um and so it was then only in 2018 that

  334. 15:52

    modern large language models started to

  335. 15:55

    appear using transformers. But the early

  336. 15:58

    versions of those um were back to small

  337. 16:02

    amounts of data. So the early GPT model

  338. 16:05

    only used 3.3 billion tokens. So down

  339. 16:08

    three orders of magnitude in size again.

  340. 16:11

    So we reverted to not enough data. And

  341. 16:14

    then it was eventually only in 2020 with

  342. 16:17

    GPT3 forward that all of those things

  343. 16:20

    came together and language models really

  344. 16:22

    took off as we know about it now.

  345. 16:26

    Um and so that led to this surprising

  346. 16:29

    victory of natural language processing.

  347. 16:31

    I mean it actually gives me a bit of a

  348. 16:33

    laugh um that if in the sort of mass

  349. 16:36

    media if you see a reference to AI these

  350. 16:39

    days with pretty high probability

  351. 16:41

    they're actually going to be talking

  352. 16:42

    about a large language model which

  353. 16:45

    wasn't the way it used to be where NLP

  354. 16:47

    used to be a fairly marginal area of

  355. 16:49

    artificial intelligence. And as we all

  356. 16:51

    know, these large language models have

  357. 16:54

    allowed us to do amazing amazing things.

  358. 16:57

    And so the ability of these systems to

  359. 17:01

    reason and solve complex problems like

  360. 17:04

    math problems is just completely

  361. 17:06

    stunning. I mean, you know, even for

  362. 17:09

    someone like me who's spent my 30 years

  363. 17:11

    working in natural language processing,

  364. 17:13

    I kind of find it hard to believe that

  365. 17:16

    you can pick really difficult math

  366. 17:19

    problems of a kind I certainly could not

  367. 17:21

    solve myself and just feed it into a

  368. 17:24

    large language model and somehow it can

  369. 17:27

    um chunk along doing its test time

  370. 17:29

    thinking and come up with right answers.

  371. 17:31

    I mean it's just been a sort of a

  372. 17:33

    stunning breakthrough from an unexpected

  373. 17:36

    direction. Um so that's been most of my

  374. 17:39

    world but then here we are um now

  375. 17:43

    getting back into the main topic of the

  376. 17:45

    talk of well where are we um why is

  377. 17:50

    embodied artificial general intelligence

  378. 17:53

    a north star that we should be looking

  379. 17:55

    at and the reason for that is even

  380. 17:58

    though it's been so amazing what can be

  381. 18:02

    done with large language models they're

  382. 18:06

    still this textbased description of the

  383. 18:09

    world. And it turns out that you can do

  384. 18:13

    a lot of stuff with a textbased

  385. 18:15

    description of the world. I think it's

  386. 18:17

    true that we can do just way more than

  387. 18:21

    almost anybody believed possible with a

  388. 18:23

    textbased description of the world.

  389. 18:26

    Certainly a lot of prominent people in

  390. 18:28

    robotics and computer vision spent a

  391. 18:30

    decade saying, "Oh, you'll never be able

  392. 18:32

    to do that with a large language model."

  393. 18:35

    where actually we've been able to do a

  394. 18:36

    lot of that with a large language model.

  395. 18:39

    But still at the end of the day um we do

  396. 18:42

    actually want to deal with the world

  397. 18:44

    around us and have artificial

  398. 18:46

    intelligence that can operate in our

  399. 18:48

    world. And so the question is then how

  400. 18:51

    can we build these embodied artificial

  401. 18:53

    general intelligences?

  402. 18:55

    And so gradually I started to um get a

  403. 18:58

    bit more interested in well how can we

  404. 19:01

    start to incorporate the visual world

  405. 19:04

    and I started to look at things like

  406. 19:06

    visual question answering and how you

  407. 19:09

    could connect between um text and then

  408. 19:12

    generative models of visual worlds and

  409. 19:15

    reason about that. And so today I'm

  410. 19:18

    going to be telling you more about how

  411. 19:20

    you can build out that line of work and

  412. 19:23

    be using simulation as a way to start to

  413. 19:26

    approach um embodied artificial general

  414. 19:29

    intelligence. So why do we want

  415. 19:32

    simulation? Um there are sort of two

  416. 19:35

    ways that you can go about um starting

  417. 19:38

    to build intelligent models of the

  418. 19:41

    physical world. One way of doing it is

  419. 19:44

    that you're actually going to directly

  420. 19:46

    learn in the real world. So you can um

  421. 19:50

    set up hardware or collect your YouTube

  422. 19:54

    videos in the real world and start

  423. 19:57

    learning intelligent policies for how to

  424. 20:00

    act. So this was the kind of approach

  425. 20:02

    that was used in Google X's QOP um work

  426. 20:05

    that you've got your um rows of robot

  427. 20:08

    arms that are doing things. you're

  428. 20:10

    recording what they're doing and you're

  429. 20:12

    in the physical world starting to learn

  430. 20:15

    a policy. Um, that's a really unpleasant

  431. 20:20

    way to try and make progress. So, when

  432. 20:22

    people are do making progress in this

  433. 20:25

    kind of world, you're getting about

  434. 20:28

    10,000 hours of teleyop, that means a

  435. 20:31

    human is moving the robot around um data

  436. 20:35

    to start to train a model. um it's not a

  437. 20:38

    very appealing picture. If we want to

  438. 20:40

    start having generally good robotics,

  439. 20:43

    we're just not going to get very far if

  440. 20:46

    we're sort of um chugging along at the

  441. 20:49

    speed of these robots um moving.

  442. 20:52

    But interestingly, if we go back to

  443. 20:55

    Shaky in the 1970s,

  444. 20:57

    Shaky had a different answer. Shaky

  445. 21:00

    said, "Well, we shouldn't just have the

  446. 21:05

    real world that's surrounding us. We

  447. 21:08

    should also have in the robot's head a

  448. 21:12

    world model that has an internal

  449. 21:16

    representation of what the world is

  450. 21:18

    like. So this idea of a world model as

  451. 21:22

    an abstracted internal representation of

  452. 21:24

    the world which can be used for

  453. 21:26

    simulation and planning that's an idea

  454. 21:29

    that goes back to cognitive science. So

  455. 21:31

    it was first proposed by Kenneth Craig

  456. 21:34

    um in the 1940s. So he argued that

  457. 21:38

    humans and other creatures have world

  458. 21:41

    models inside their own heads. Um so um

  459. 21:45

    that they can think about alternatives

  460. 21:48

    and how they're likely to play to play

  461. 21:50

    out and therefore they can think of good

  462. 21:53

    plans um before acting. And so that's

  463. 21:57

    also what we'd like our robots to do is

  464. 22:01

    to have the same kind of abilities. Um

  465. 22:04

    so this was Shaky's world model. So,

  466. 22:06

    Shiki had a kind of a blocks world where

  467. 22:08

    it was actually a grid world as you

  468. 22:10

    might remember from your early AI

  469. 22:12

    textbooks, but it was representing um

  470. 22:15

    what was in different places in the

  471. 22:17

    world and what were the attributes of

  472. 22:20

    different things and it was all stored

  473. 22:22

    in this kind of logical representation

  474. 22:25

    in those days. So how now can we start

  475. 22:30

    um building a 21st century version of an

  476. 22:34

    embodied artificial general

  477. 22:36

    intelligence? And I think the right way

  478. 22:38

    to do it is to work out how to build

  479. 22:41

    good action condition world models. And

  480. 22:44

    today I want to talk about a practical

  481. 22:48

    effective way to build a simulation

  482. 22:51

    infrastructure that can be used um for

  483. 22:54

    as a basis from embodied AGI.

  484. 22:58

    So the idea of a world model is now

  485. 23:00

    normally formalized in terms of

  486. 23:02

    reinforcement learning ideas that what

  487. 23:05

    we do is we have observations of a world

  488. 23:08

    but we assume that underlying those we

  489. 23:11

    have a semantic abstracted

  490. 23:14

    representation of the state of the world

  491. 23:16

    in our head. And what the world model

  492. 23:19

    does is gives us an ability to try and

  493. 23:23

    predict um when an action is taken in

  494. 23:26

    one state, what new state is going to

  495. 23:29

    emerge. And so the crucial thing there

  496. 23:32

    is that the world model is abstracted

  497. 23:36

    and has more semantics. Quite a lot of

  498. 23:39

    the time when people have talked about

  499. 23:42

    world models, they haven't really been

  500. 23:45

    thinking about the abstraction and the

  501. 23:47

    semantics, they've just been talking

  502. 23:49

    about can you produce beautiful

  503. 23:52

    generative AI video. Um so here's um

  504. 23:56

    Genie 3. Um this is kind of a cute

  505. 23:59

    example by Riley Goodside who was the

  506. 24:01

    same person who got a lot of fame in the

  507. 24:04

    early LLM days by being a good prompt

  508. 24:07

    engineer. um and he's now playing around

  509. 24:10

    um here with Genie 3. And you know the

  510. 24:13

    the kind of things you can do with Genie

  511. 24:15

    3, you know, it looks beautiful, it's

  512. 24:18

    wonderful, it's pretty fluid, it seems

  513. 24:23

    great, but you know, most of visual AI

  514. 24:26

    for this period has been just judged by

  515. 24:29

    the pixels. If you've got beautiful

  516. 24:32

    pixels, you've got a beautiful piece of

  517. 24:34

    software, but these pixels are trying to

  518. 24:37

    simulate observations. They don't

  519. 24:40

    actually have or represent any of the

  520. 24:43

    semantics behind the world. And

  521. 24:45

    therefore, they aren't good at realism.

  522. 24:48

    They aren't good for giving a basis to

  523. 24:50

    plan. They aren't good for actually

  524. 24:52

    having a robust simulation

  525. 24:54

    infrastructure of the real world. And so

  526. 24:57

    we want this kind of simulation

  527. 24:59

    infrastructure because with a good

  528. 25:01

    simulator we actually have causal

  529. 25:03

    knowledge of how the world works which

  530. 25:06

    allows us to predict and plan how things

  531. 25:09

    will work in any situation. And so

  532. 25:13

    that's the kind of world um that we're

  533. 25:16

    wanting to have and make available at

  534. 25:19

    Moon Lake. So the starting point is a

  535. 25:23

    real world observation. Um, so given an

  536. 25:27

    image or a bit of video, we want to be

  537. 25:30

    able to interpret this image as an

  538. 25:33

    observation which is always partial of

  539. 25:36

    what's actually in the underlying world.

  540. 25:39

    And then what we want to do is

  541. 25:41

    reconstruct a model of this world which

  542. 25:45

    actually allows us to do stuff in it and

  543. 25:48

    see how it reacts. that this will give

  544. 25:50

    us a basis of being able to work out

  545. 25:54

    causality, work out how to plan and

  546. 25:57

    reason in a repres representation

  547. 26:00

    condition way which gives us the basis

  548. 26:02

    for intelligent robot actions inside

  549. 26:06

    this world.

  550. 26:08

    So what might one do here? Well, the

  551. 26:11

    first thing you can do is take that um

  552. 26:14

    picture and feed it into one of our

  553. 26:17

    well-known um other companies products

  554. 26:20

    and say, "Okay, make a simulation of

  555. 26:22

    this world." Um and so this is what you

  556. 26:26

    get from Marble. And it's a pretty good

  557. 26:28

    simulation of the world. And if what you

  558. 26:31

    want to do is just uh walk around in

  559. 26:34

    this world and see it from different

  560. 26:36

    angles, this works pretty well. Um the

  561. 26:42

    question is is this a good simulation

  562. 26:44

    and whether it's a good simulation

  563. 26:47

    depends on

  564. 26:50

    what problem you want to solve in the

  565. 26:52

    real world. And then is does this give

  566. 26:55

    you sufficient information to solve the

  567. 26:58

    problem in the real world? And if all

  568. 27:01

    you want to do in the real world is to

  569. 27:03

    be able to wander around and see the

  570. 27:05

    view from different angles, then this is

  571. 27:08

    great. We're done. But a lot of the time

  572. 27:11

    what you'd like to do in the real world

  573. 27:15

    is understand the objects are here and

  574. 27:18

    to be able to do things with them.

  575. 27:21

    Right? So there's some objects here.

  576. 27:22

    There's a kettle and there's a box and

  577. 27:25

    there's a cup. But in the marble world,

  578. 27:29

    there's nothing you can actually do with

  579. 27:31

    these things. All you can do is sort of

  580. 27:33

    wander around as a disembodied figure.

  581. 27:36

    Um, and so for a lot of purposes such as

  582. 27:39

    doing things with robotics or other

  583. 27:42

    physically accurate worlds, um, we need

  584. 27:45

    to be able to have more understanding in

  585. 27:47

    our model of the world, a better, more

  586. 27:50

    detailed simulation. And so at Moon

  587. 27:53

    Lake, we're wanting to work out what

  588. 27:55

    people actually want to do with their

  589. 27:57

    simulation, and then to build the kind

  590. 27:59

    of simulation that will power that. So

  591. 28:02

    you might think, oh, I actually want to

  592. 28:04

    be able to move around the objects in

  593. 28:06

    the world. And so that's the kind of

  594. 28:09

    starting point of the kind of thing that

  595. 28:11

    we're trying to do at Moon Lake. So if

  596. 28:14

    we want to um have a world in which we

  597. 28:18

    can move things around then we're saying

  598. 28:20

    okay kind of like um image blaster a

  599. 28:23

    system like that we actually need to

  600. 28:26

    take this world and understand what's in

  601. 28:29

    it and then have objects that can be

  602. 28:31

    manipulated. So we separate out a

  603. 28:34

    background and then objects inside that

  604. 28:38

    background with then being able to sort

  605. 28:40

    of have the background world in a

  606. 28:43

    representation that's similar to marble.

  607. 28:46

    But then in the foreground there are

  608. 28:48

    various kinds of objects that we can

  609. 28:50

    manipulate and move around. Now, well,

  610. 28:53

    that's a start, but um in the previous

  611. 28:57

    picture, um the tea box was always

  612. 29:00

    closed. And we might wonder if you can

  613. 29:03

    open up the tea box and see what's

  614. 29:06

    inside it. Um and well, this is sort of

  615. 29:09

    a part of how observations of a world

  616. 29:12

    are always partial. Um for what we could

  617. 29:15

    see in the actual photo, there was a

  618. 29:17

    closed tea box. We couldn't even see the

  619. 29:19

    tea box very well. How could we possibly

  620. 29:22

    know what's inside it? And well, the

  621. 29:26

    answer to the way we can know what's

  622. 29:28

    inside it in the 21st century is we go

  623. 29:32

    off and do a little bit re of research.

  624. 29:34

    We fire up um getting information um off

  625. 29:38

    the web um in a

  626. 29:42

    a rag style fashion um to work out

  627. 29:45

    what's there. And then we can find

  628. 29:48

    images of this um tea um box on the web

  629. 29:52

    which show it in more detail. They show

  630. 29:54

    a picture of what it's like when it's

  631. 29:56

    open. There are descriptions of it.

  632. 29:58

    Organic luxury tea bag collection,

  633. 30:00

    leather gift box. It gives its size. We

  634. 30:03

    can find all about it. So therefore, we

  635. 30:05

    can build um in our simulation um a tea

  636. 30:09

    box with an understanding of the

  637. 30:11

    contents of that tea box. Well, that's

  638. 30:14

    really good. Um, so how now we have um

  639. 30:17

    the tea box with tea bags inside it. But

  640. 30:20

    if we do nothing else, um, they're just

  641. 30:24

    sort of sitting there and movable. So

  642. 30:26

    we'd like to realize the fact, well,

  643. 30:28

    wait a minute, tea bags you can lift up

  644. 30:31

    and you can take out of a tea box. So

  645. 30:35

    then we need to start having these

  646. 30:37

    teaags also um be objects that are

  647. 30:40

    modeled in our simulation that they have

  648. 30:42

    a size and an ability to be manipulated

  649. 30:46

    as well. Um so how are we doing all of

  650. 30:50

    this? And so a distinctive part of

  651. 30:52

    what's happening at Moon Lake is

  652. 30:55

    building although part of this is in the

  653. 30:59

    world of vision

  654. 31:01

    a lot of what we're doing is actually

  655. 31:03

    back in the world of code. So this is

  656. 31:06

    picking up on the idea um that what's

  657. 31:10

    normally referred to as language models

  658. 31:12

    but these days are really symbolic

  659. 31:14

    models which work on not only human

  660. 31:17

    languages but also math code and things

  661. 31:20

    like that that they that has been just a

  662. 31:23

    very powerful substrate with which to

  663. 31:26

    make progress. And so we are producing

  664. 31:29

    controllable manipulable world models by

  665. 31:33

    generating the controllable parts of

  666. 31:36

    these worlds by putting code under them.

  667. 31:40

    And in particular, um, we can train

  668. 31:43

    these code models using the same kind of

  669. 31:46

    loop engineering that many of you will

  670. 31:48

    have seen, um, with Claude that we're

  671. 31:51

    having a loop where we're writing code

  672. 31:55

    to render objects,

  673. 31:58

    um, and put on that um, diffusionbased

  674. 32:01

    textures, etc. We can then assess how

  675. 32:05

    good our render is against physical

  676. 32:08

    reality and then we can work out the

  677. 32:11

    deviations. We can revise the code and

  678. 32:15

    make better and better um renders of

  679. 32:17

    what goes along. And so that's the way

  680. 32:20

    that so we're building this detailed

  681. 32:24

    action conditional world model so that

  682. 32:26

    we can take actions on the objects in

  683. 32:29

    the world.

  684. 32:30

    Um so that means that we have a kind of

  685. 32:33

    neuros symbolic representation of the

  686. 32:36

    world here and neuros symbolic

  687. 32:38

    representations have a big advantage

  688. 32:41

    that the symbolic representation can be

  689. 32:44

    easily interfaced controlled edited and

  690. 32:47

    maintained in our human world. Um, and

  691. 32:51

    so the message I'd like to give as a

  692. 32:56

    little delta of a message here, um, is

  693. 32:59

    most of you are probably familiar with

  694. 33:02

    Mark Andre's famous statement 15 years

  695. 33:04

    ago, software will eat the world. And

  696. 33:08

    you know, that was mostly right. Um but

  697. 33:12

    I think it wasn't completely right

  698. 33:15

    because it's just not the case that

  699. 33:18

    software ate the physical world. Um it,

  700. 33:22

    you know, a lot of the world went

  701. 33:24

    virtual and um a lot of the world is

  702. 33:27

    inside computers now and yeah it could

  703. 33:30

    eat all of the processes of managing,

  704. 33:34

    counting, supplying all um records of

  705. 33:37

    suppliers, all the stuff that's in the

  706. 33:40

    virtual world. But we hadn't had the

  707. 33:43

    power for software to eat the physical

  708. 33:45

    world. Whereas the hope is with this new

  709. 33:49

    ability of being able to build

  710. 33:52

    verifiable simulations powered by code

  711. 33:55

    of the sort that I've just sketched for

  712. 33:57

    a moment there that this will allow us

  713. 34:00

    to actually have verifiable simulation

  714. 34:03

    which will also eat the physical world.

  715. 34:06

    And so that's the kind of thing we

  716. 34:08

    built. And we can go on from here and um

  717. 34:11

    you know keep on making further steps of

  718. 34:14

    this. Right? So we maybe don't want to

  719. 34:15

    only have tea bags um in foil

  720. 34:19

    containers, but we'd like to be able to

  721. 34:21

    um take them out of the foil container

  722. 34:24

    and we don't we somehow want to actually

  723. 34:27

    get the tea bag inside the cup because

  724. 34:29

    that's a useful step for making tea. And

  725. 34:32

    then of course we also want to have um

  726. 34:35

    the jug of um water which we want to be

  727. 34:38

    able to boil and then be able to pour

  728. 34:40

    that into the tea. Um and once we have a

  729. 34:45

    good world simulator like that, the idea

  730. 34:48

    of this is that we are building in all

  731. 34:52

    of the parts of the simulated world

  732. 34:56

    which will allow effective transfer into

  733. 35:00

    the real world. And it's important to

  734. 35:02

    think about that as to sort of what's

  735. 35:05

    necessary for different kinds of

  736. 35:07

    transfer. And the argument um that I

  737. 35:12

    we'd like to make is that you know

  738. 35:14

    normally uh simulation isn't complete

  739. 35:19

    and accurate in every detail because

  740. 35:22

    there are sort of parts of the world

  741. 35:24

    that are important to you and parts of

  742. 35:27

    the world that aren't important to you.

  743. 35:29

    I mean, this is the same with human

  744. 35:31

    world models, right? That a lot of the

  745. 35:33

    world we're not actually modeling at any

  746. 35:35

    time, but we're modeling the bits of it

  747. 35:37

    that are important to us. And it's

  748. 35:39

    having that control of which things you

  749. 35:42

    need to model and use is what needed

  750. 35:45

    then to give you the actual simulations

  751. 35:48

    that will be allowed to be effectively

  752. 35:50

    used um in the real world. So, what kind

  753. 35:53

    of applications

  754. 35:55

    um is this going to allow us to do? So

  755. 35:58

    the hope is that um rather than having

  756. 36:02

    to collect 10,000 hours of data in by

  757. 36:07

    teley op in the real world, we can

  758. 36:09

    instead build accurate simulations which

  759. 36:12

    will allow transfer to real. So let me

  760. 36:15

    show you the kind of progress that we've

  761. 36:17

    been making on this. So this is um not

  762. 36:20

    highuting robotics um but is the kind of

  763. 36:23

    physical AI that you find everywhere in

  764. 36:27

    the real world that there are things

  765. 36:29

    happening in physical processes um which

  766. 36:32

    you'd like to be able to um understand

  767. 36:37

    and automate. And so in particular, um,

  768. 36:40

    if we're going to realize any of the

  769. 36:42

    dreams of bringing back America as a

  770. 36:45

    great manufacturing economy, it seems

  771. 36:47

    like we have to work out how to be able

  772. 36:49

    to automate much more in the physical

  773. 36:52

    world. Well, um, using the Moonlake AI

  774. 36:56

    technology, what we can then do is say,

  775. 36:59

    so from this short video, we can

  776. 37:02

    automatically produce an accurate

  777. 37:05

    simulation of that world, turning it

  778. 37:07

    into a 3D world with enough detail about

  779. 37:11

    how things move and what objects are in

  780. 37:14

    the world that we can start to build on

  781. 37:17

    this and use it as a basis of simulated

  782. 37:20

    data which transfers refers accurately

  783. 37:23

    to the real world. So we then have this

  784. 37:26

    model which allows us to get 10,000

  785. 37:28

    hours of simulation um for free. And so

  786. 37:32

    on the basis of that um we can then

  787. 37:35

    train up a robotic system that can

  788. 37:38

    operate um to actually um explore what

  789. 37:44

    you can do in this world and learn the

  790. 37:47

    way to operate in the world so it works.

  791. 37:50

    and then we can have the effective and

  792. 37:53

    cheap training of robotic systems.

  793. 37:57

    Okay. Um so that's the story. Um thanks

  794. 38:00

    a lot everyone.

  795. 38:02

    [applause]

  796. 38:09

    And I do have time for questions I

  797. 38:11

    believe.

  798. 38:16

    >> Thank you very much for the

  799. 38:17

    conversation. Um, super deep. My first

  800. 38:20

    question is, would this be applied to

  801. 38:23

    something like the gaming industry and

  802. 38:26

    how could they use it for we're seeing

  803. 38:28

    GTA 6? Everyone is talking about this

  804. 38:31

    new game. Um, I could see a very

  805. 38:34

    application there where you could have

  806. 38:36

    endless interaction with the world model

  807. 38:39

    that they're building. U, have you been

  808. 38:41

    working with the gaming industry at all

  809. 38:42

    or this is more applied to robotics?

  810. 38:44

    >> Um, so yeah, absolutely. Um, another

  811. 38:49

    hugely good area for this is in the

  812. 38:51

    gaming industry. And you know, the real

  813. 38:55

    world is then kind of a virtual world,

  814. 38:57

    but you can have a simulation of your

  815. 38:59

    gaming world um, and then be building

  816. 39:02

    the same kind of embodied intelligence.

  817. 39:06

    And the gaming industry has some really

  818. 39:08

    appealing attributes. Well, you know,

  819. 39:10

    there are millions and millions of game

  820. 39:12

    players, so there's lots of easy data to

  821. 39:14

    collect. Um the people who are gaming

  822. 39:18

    have goals which are fairly clearly

  823. 39:20

    known. So you can sort of learn in a

  824. 39:22

    reinforcement learning loop good ways to

  825. 39:25

    act. And so absolutely one of the

  826. 39:27

    applications that Moon Lake has explored

  827. 39:30

    is in the gaming industry and there's

  828. 39:32

    likely to be more of that. Um but

  829. 39:34

    recently we've been actually

  830. 39:37

    particularly emphasizing physical

  831. 39:39

    infrastructure and doing simulations in

  832. 39:40

    the physical infrastructure world.

  833. 39:42

    >> Thank you so much.

  834. 39:55

    Uh so it's fascinating to me that that

  835. 39:58

    you're turning back to neurosy symbolic

  836. 40:00

    representations and um the part of the

  837. 40:04

    original kind of old school AI involved

  838. 40:07

    a lot of ontology development and

  839. 40:09

    knowledge representation and logic and

  840. 40:11

    representation of that. Do you see a

  841. 40:13

    role for uh deepening the neuros

  842. 40:16

    symbolic representation using kind of

  843. 40:19

    like what we how ontologies have been

  844. 40:21

    developed previously to model um a

  845. 40:24

    representation of the real world? Um and

  846. 40:27

    do you see that as being a kind of an

  847. 40:28

    area of expansion and development for

  848. 40:31

    the type of tools that you're working

  849. 40:32

    in? Yeah, I mean I don't think we're

  850. 40:36

    quite going to go back to old style

  851. 40:38

    ontologies and knowledge bases, but I

  852. 40:41

    mean effectively

  853. 40:43

    that is what we're doing, you know,

  854. 40:46

    apart from it's in new clothing of um

  855. 40:49

    having it being um code that I mean I I

  856. 40:55

    do you know this is an interesting space

  857. 40:58

    and we can talk about it in more general

  858. 41:00

    right there is you know there's sort a

  859. 41:03

    purely neural approach in which your

  860. 41:05

    latent representation of the world is

  861. 41:07

    purely neural and so that's the kind of

  862. 41:09

    thing that the jeoper architecture is

  863. 41:12

    after I mean in some ways that's a good

  864. 41:15

    pure neural approach but on the other

  865. 41:18

    hand that's very hard for humans to

  866. 41:24

    connect to in any way or for other

  867. 41:26

    applications to connect to in any way

  868. 41:28

    and so our bet is that the sort of

  869. 41:31

    practical ical way to have a physical AI

  870. 41:34

    simulation infrastructure for the

  871. 41:36

    foreseeable future is to base it on

  872. 41:39

    symbolic representations.

  873. 41:41

    Taking advantage of sort of the huge

  874. 41:43

    power of code that we see everywhere

  875. 41:47

    around us in the sort of um codeex cla

  876. 41:50

    code era that those kind of symbolic

  877. 41:53

    representations allow um neural eye

  878. 41:56

    systems to reason, plan and do all of

  879. 41:59

    these things excellently well. And so to

  880. 42:01

    some extent yes that will be bringing

  881. 42:03

    back notions like older fashion

  882. 42:06

    knowledge representation.

  883. 42:14

    Mic keeps on.

  884. 42:21

    Hello. Hello. Oh hi. Um so a lot of our

  885. 42:25

    existing um it seems like industrial

  886. 42:28

    robotics use cases or to sort of

  887. 42:30

    automate existing sort of human manual

  888. 42:34

    intensive processes. Um would investment

  889. 42:38

    in

  890. 42:40

    development of like world models would

  891. 42:42

    that allow us to go explore novel new uh

  892. 42:46

    processes or activities where they're

  893. 42:49

    currently not accessible by human labor

  894. 42:52

    or by even the kind of existing um

  895. 42:55

    industrial manufacturing processes that

  896. 42:57

    we have.

  897. 42:58

    >> I mean sure absolutely. Yeah. So for the

  898. 43:01

    example at the end I showed a very old

  899. 43:04

    school um conveyor belt system but I

  900. 43:08

    mean you know we're also in this world

  901. 43:11

    of amazing things happening in humanoid

  902. 43:14

    robotics and for any of the new forms of

  903. 43:20

    robotics automation

  904. 43:23

    other cases you could think about doing

  905. 43:25

    things in space well you'll have exactly

  906. 43:27

    the same problem that You want to train

  907. 43:31

    up AI agents to be able to act in

  908. 43:35

    different scenarios and it's extremely

  909. 43:40

    costly and difficult to do that training

  910. 43:43

    in the real world. Um and it's very hard

  911. 43:47

    when training in the real world to get

  912. 43:49

    out into the sort of tale of rare cases.

  913. 43:53

    I mean you've seen that um for

  914. 43:55

    autonomous driving, right? There's a

  915. 43:57

    reason why autonomous driving kind of

  916. 43:59

    took 20 years to arrive. Um from you

  917. 44:02

    know Stanley breakthroughs of yay

  918. 44:05

    autonomous car wins the race to actually

  919. 44:07

    having Whimos um widely deployed is

  920. 44:10

    because you have this enormous tail and

  921. 44:12

    the way to get out to that tail for all

  922. 44:14

    of these new applications with um

  923. 44:17

    humanoid robotics, space robotics, etc.

  924. 44:20

    is to have good simulation.

  925. 44:26

    >> Hey, uh thank you for your talk. I love

  926. 44:28

    the energy. Uh my question is you showed

  927. 44:32

    a video of water. Uh so is everything

  928. 44:35

    like learned from videos or is like

  929. 44:37

    physics ingrained in anything like for

  930. 44:39

    example does your model know any

  931. 44:41

    physical properties of water like

  932. 44:42

    viscosity, Bernali's laws or like for

  933. 44:45

    example friction. uh for some use cases

  934. 44:48

    like self-driving maybe those things are

  935. 44:50

    like not important right you just need

  936. 44:52

    to know the velocity and stuff but like

  937. 44:54

    for other use cases I would imagine

  938. 44:55

    these physical properties are important

  939. 44:57

    so

  940. 44:57

    >> and we absolutely know physics I mean

  941. 44:59

    this is the sense in which this

  942. 45:02

    neurosyolic approach is you know you

  943. 45:06

    call a conservative approach if you will

  944. 45:09

    you know that this is absolutely using

  945. 45:11

    physics engines and knowledge of physics

  946. 45:15

    um to in its generation and control of

  947. 45:19

    movement, right? That if you're what you

  948. 45:22

    know when you've just started off with

  949. 45:24

    an image of water or and you want to

  950. 45:27

    understand how that's behaves, you're

  951. 45:30

    using a physics model to predict how

  952. 45:32

    it's going to behave.

  953. 45:36

    Okay.

  954. 45:40

    Um I know is there time for Oh, there's

  955. 45:44

    someone else at the mic.

  956. 45:47

    Hi.

  957. 45:47

    >> Hey. Uh, Professor Manning, thank you

  958. 45:49

    for this presentation. Um, I have a

  959. 45:52

    question about seem to row gap. So, I

  960. 45:55

    mean, simulation is great. It's cheap,

  961. 45:57

    is scalable. Um do you have a suggestion

  962. 46:01

    for um how we close the simulation to

  963. 46:04

    real gaps

  964. 46:06

    [snorts] encounter you know like the

  965. 46:09

    like fix part that we are not able to

  966. 46:12

    accurately simulate and like mechanical

  967. 46:16

    tolerancing back latches like all those

  968. 46:19

    things that we are not able to put into

  969. 46:23

    the simulations.

  970. 46:24

    Um yeah so traditionally the problem has

  971. 46:27

    always been the simtoreal gap and the

  972. 46:30

    simtoreal gap seeming too large so that

  973. 46:34

    a lot of the training has had to happen

  974. 46:38

    in the real world

  975. 46:41

    and we believe that the answer to that

  976. 46:44

    is you know effectively claude code loop

  977. 46:49

    we're now in this world in which we can

  978. 46:52

    do neural optimization.

  979. 46:55

    So we can have the simulation

  980. 46:59

    compared to behavior in the real world

  981. 47:02

    with video and we can automatically

  982. 47:06

    learn to shrink the sim to real gap in a

  983. 47:10

    way that just wasn't possible when you

  984. 47:12

    had people trying to handr write a

  985. 47:16

    physics simulation of something.

  986. 47:22

    This this was a great discussion. Thank

  987. 47:24

    you. Um tangential to the question that

  988. 47:27

    was just asked, how do you see like what

  989. 47:30

    do you see the biggest challenges are to

  990. 47:32

    adding like the sociote techchnical

  991. 47:36

    layer, the pieces, the the processes,

  992. 47:38

    the people, the authorities that make

  993. 47:40

    decisions on the objects that you're

  994. 47:42

    looking to simulate. What are the

  995. 47:44

    biggest challenges to developing a

  996. 47:46

    system that has accurate representation?

  997. 47:49

    So for example, if we want to model uh

  998. 47:53

    EV tolls or um doing you know

  999. 47:56

    forecasting of energy at airports uh you

  1000. 47:59

    know things like that not necessarily

  1001. 48:00

    extending beyond the robotics into these

  1002. 48:04

    other um use cases areas.

  1003. 48:06

    >> Um yeah that's a great area. Um in all

  1004. 48:09

    honesty it's not something we've been um

  1005. 48:11

    really dealing with at Moon Lake. Um,

  1006. 48:14

    you know, at the there are still some

  1007. 48:17

    limitations, but you know, large

  1008. 48:20

    language models with their slurping up

  1009. 48:23

    of enormous amounts of human behavior

  1010. 48:26

    data that they're actually getting

  1011. 48:30

    better and better at being able to

  1012. 48:32

    simulate how human beings, different

  1013. 48:35

    kinds of human beings are going to

  1014. 48:38

    behave and react in different

  1015. 48:40

    circumstances.

  1016. 48:42

    So I think we are starting to approach

  1017. 48:45

    the point in which we can have fairly

  1018. 48:49

    good human behavior simulators that are

  1019. 48:53

    being powered by the knowledge of large

  1020. 48:55

    language models and there are a couple

  1021. 48:57

    of companies that are now starting to

  1022. 48:58

    look at that.

  1023. 49:02

    >> Cool. Last question. Yeah.

  1024. 49:09

    person over here has one for a long

  1025. 49:11

    time. [laughter]

  1026. 49:16

    >> Thank you. Um so if you look at um

  1027. 49:20

    analog chips for instance right the

  1028. 49:22

    physics is understood but as you go

  1029. 49:24

    higher up the chemical operations then

  1030. 49:27

    the heat and those things are still

  1031. 49:29

    people are working on it. So that's not

  1032. 49:31

    what you call completely

  1033. 49:33

    um understood. Now I'll give you another

  1034. 49:36

    example of say I'm looking at kidney

  1035. 49:38

    related uh literature and stuff and then

  1036. 49:41

    liver related stuff

  1037. 49:43

    these things it gives a pretty good

  1038. 49:45

    answer the kidney stuff it gives answer

  1039. 49:47

    but when you look at the relations

  1040. 49:49

    between the two people have somewhere

  1041. 49:52

    u doctors have made a link and that's

  1042. 49:54

    why you are able to see my take is if we

  1043. 49:58

    operated completely in the

  1044. 50:02

    latent space and we did not know the

  1045. 50:04

    connection connections at a human level.

  1046. 50:06

    Can we explore in the latent space

  1047. 50:08

    completely and be able to find hidden

  1048. 50:10

    connections over there that could

  1049. 50:12

    translate into the real world or am I

  1050. 50:15

    thinking something crazy?

  1051. 50:16

    >> Yeah, I mean absolutely. I mean to the

  1052. 50:19

    extent that we have a pretty good

  1053. 50:21

    simulation, we can hope to find

  1054. 50:26

    surprising discoveries that turn out to

  1055. 50:28

    be correct in the simulated world. I

  1056. 50:31

    mean, you know, things go both ways,

  1057. 50:35

    right? There are also likely to be

  1058. 50:37

    things that turn out to be true in the

  1059. 50:39

    real world that weren't in our

  1060. 50:41

    simulation, right? There's this famous

  1061. 50:42

    statement about all models are wrong,

  1062. 50:44

    but some models are useful, right? And

  1063. 50:47

    that so anybody's world model,

  1064. 50:51

    regardless of whether it's the one we're

  1065. 50:52

    building or the world model in a human

  1066. 50:55

    head, right, they're not always right.

  1067. 50:57

    Sometimes we think a person's going to

  1068. 50:59

    react in one way and they react in

  1069. 51:00

    another way. But nevertheless, a lot of

  1070. 51:03

    the time it can let us explore much more

  1071. 51:06

    widely and discover new facts and new

  1072. 51:08

    connections that we weren't aware of in

  1073. 51:10

    the real world.

  1074. 51:13

    Okay,

  1075. 51:16

    >> awesome. That's a wrap everyone. Thank

  1076. 51:17

    you so much, Chris.