AI Engineer World's Fair 2026
World Models Need Causality, Not Pretty Pixels — Christopher Manning, Moonlake AI
Read the talk
World Models Need Causality, Not Pretty Pixels
Christopher Manning traces the history of AI and language models, then develops Moonlake AI’s approach to physical intelligence: reconstruct a world from observations, give its objects behavior in code, and refine the simulation against reality.
From a talk by Christopher Manning
At a glance
Ideas worth remembering
Language-model progress needed model flexibility alongside data and compute; substantial text scale existed before architectures could use it with today’s breadth.
An action-conditioned world model represents state and predicts how actions change it. Attractive generated observations alone do not establish that capability.
The tea-box example adds capability in layers: separate movable objects, reconstruct hidden contents using retrieved information, then make those contents independently manipulable.
Moonlake combines generated code, textures and physics models, then proposes refining simulations through comparisons with real observations and behavior.
Simulation fidelity should follow the intended task. Simulated training and discovery remain useful only insofar as the model captures the real-world details that matter.
AI began with language, feedback and machines that could act
Christopher Manning opens with a goal for Moonlake AI: simulation infrastructure for practical physical AI. Embodied intelligence means an intelligence that can operate in an environment, rather than only describe it. But the route to that goal begins with a deliberately slow historical build. Language, control and robots have been part of AI’s story from the beginning.
The 1956 Dartmouth summer research project supplies the familiar starting point: John McCarthy coined the term artificial intelligence and gathered researchers including Claude Shannon. Earlier cybernetics work had already connected communication, control and feedback in living things and computers. That tradition matters here because acting intelligently requires a loop: observe an environment, act on it, and use the resulting change to decide what comes next.
Natural language processing also predates the name AI. Manning revisits a 1954 public demonstration of Russian-to-English machine translation, accompanied by predictions that computers would replace most human translators. The demonstration establishes an early ambition that would take decades to mature: language processing as useful machine intelligence, rather than a peripheral application.
McCarthy founded the Stanford Artificial Intelligence Lab starting in 1963. Although his own work centered on mathematical logic, he supported a wider effort to build embodied intelligence. The Stanford Cart and, more directly relevant to this talk, Shakey brought perception and action together. Shakey could perceive its environment, move around and move boxes. Understanding the world was already tied to doing something in it.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Language models were useful long before they became the center of AI
A playful Stanford-versus-Berkeley detour places this history in its institutional setting. Manning emphasizes Stanford’s early AI work and network connections, while acknowledging Berkeley’s important systems work and its later growth in AI. He then returns to his own field: natural language processing, where language models had long been a central technology even when much of the rest of AI paid them little attention.
The lineage runs from character-level models of language to Shannon’s word and character n-gram models in the late 1940s, then to probabilistic text models developed at IBM in the 1970s. These models gave speech and NLP systems a way to judge which sequences of words were plausible. Their value was practical: they helped power speech recognition, spelling correction and machine translation, including Google’s translation work around 2007.
Useful did not yet mean general. In Manning’s recollection, language models were components for particular tasks; researchers still expected broader intelligence to require separate memories, knowledge representations, planning systems and reasoning systems. That expectation makes the later rise of language models surprising even to someone who spent his career building them.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Data scale needed a flexible model
Manning’s language-model history separates three ingredients that are easy to collapse into one story: data, compute and model flexibility. An early neural language model around 2000 used a 32 million-token corpus and a 31,000-word vocabulary. Compute constrained what it could do. By contrast, he cites Google’s 2007 language model built on two trillion tokens: substantial data scale existed well before today’s neural models. 14:58
The missing ingredient in that large earlier model was flexibility. More text could not supply the expressive power of neural architectures that had not yet arrived. Transformer-based language models appeared in 2018, but early versions again used comparatively small datasets. In this account, the ingredients came together with GPT-3 in 2020 and subsequent models: enough data and compute, joined to an architecture capable of using them.
The result was a surprising victory for NLP. Language models moved from a relatively marginal part of AI to the technology many people now mean when they say AI. Manning’s surprise is personal: after 30 years in the field, he still finds it remarkable that a model can work through difficult mathematics using test-time thinking and reach answers he could not produce himself. That success sets a demanding starting point for the next question: what does intelligence still need to act in the physical world?
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A world model predicts what an action changes
Text-based descriptions have supported far more intelligence than many researchers expected. Physical intelligence adds another demand: an agent must operate in the world around it. One route is to learn directly from hardware experiments, recorded robot behavior or real-world video. Manning points to roughly 10,000 hours of human teleoperation as an example of the data collection burden. Robots move at physical speed, and people must spend time guiding them.
Shakey supplies an older alternative: keep an internal representation of the world and use it to consider actions before executing them. Its representation recorded locations and object attributes in a logical grid world. The modern version keeps the same purpose while seeking richer environments. An action-conditioned world model starts from a representation of the current state and predicts the state that will result from a chosen action.
The important distinction is between an observation and the state behind it. A picture shows what a camera can see. A semantic state represents things the agent needs to reason about: objects, their attributes and how actions affect them. Planning needs predictions about those changes, so the quality of generated images alone is an insufficient test.
A fluid Genie 3 generative-video example by Riley Goodside makes the distinction visible. Manning credits its visual appeal, then criticizes the demonstrated approach for relying on simulated observations without the semantics needed for dependable planning. This is his assessment of the demonstrated approach, rather than a measured comparison of planning performance. Moonlake’s proposed advantage is causal structure: a simulator should explain how acting on something produces a change that an agent can anticipate. 23:56
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
From a room you can view to a tea box you can open
Moonlake’s starting input is an image or a short video: a partial observation of an underlying world. The desired output is a model in which an agent can act and inspect the consequences. A room reconstruction provides a concrete test. Manning describes a Marble reconstruction that supports walking around and seeing the scene from different angles. For that task, it works well. But the room also contains a kettle, a box and a cup, and the demonstrated reconstruction does not let the user manipulate them.
The first change is to separate the scene into background and foreground objects. The background can retain a representation suited to viewing the environment. Foreground objects become separate things that can move. The observable difference is simple: the tea box stops being part of the room’s appearance and becomes an object the agent can act on.
Opening the box exposes a harder problem. The original photo shows it closed, so its contents cannot be recovered from visible pixels alone. Moonlake adds web retrieval in a retrieval-augmented fashion: find images of the product open, read its description and obtain dimensions. That information supports a reconstruction with tea bags inside. It describes the product’s expected contents; it cannot establish which contents are actually present in this particular closed box. 29:04
The next change is equally important: the contents must become objects too. Tea bags that merely appear inside the box do not yet support the task. They need their own size and manipulable representation so an agent can lift one and take it out. Each step adds a capability required by the intended action, rather than indiscriminately adding detail to the whole room.
What changes between viewing the tea box and using it? The diagram follows the added representations. A movable box needs a separate identity; an openable box needs hidden structure; removing a tea bag needs the contents to have their own identities and behavior. Visual reconstruction supplies the scene, while these additions supply the actions.
The observation shows the exterior and hides the contents.
The same scene gains progressively richer actions as the box and its contents become separately represented objects.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Code makes the reconstructed world editable
The controllable parts of Moonlake’s worlds have code underneath them. This uses the ability of language models to work with symbolic material—including mathematics and code—as well as natural language. Generated code supplies an editable representation of objects and behavior. Neural generation and symbolic control meet in what Manning calls a neurosymbolic world model.
The construction process uses an iterative coding loop. Code renders objects, diffusion-based generation supplies textures, and the rendered result is assessed against physical reality. Differences guide revisions to the code, producing a new render for the next comparison. The practical advantage of the symbolic layer is that people and other applications can interface with it, control it, edit it and maintain it. 31:40
How does a reconstruction improve after its first attempt? The loop below makes the feedback path explicit: reality provides a comparison target, the comparison identifies deviations, and code revisions change the next simulated result. This is the mechanism Moonlake proposes for improving its generated worlds, rather than treating the first reconstruction as final.
Manning extends the idea that software will eat the world to physical processes. Software already manages records, suppliers and other virtual representations of activity. His hope is that verifiable simulations powered by code will let it reach further into the physical activity itself. The link is prediction: a useful simulation gives an agent a place to test what an action will do before it acts outside the computer.
Code represents the controllable parts of the world.
The comparison feeds back into code changes, which produce the next result to assess.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Model the details the task needs
The tea example continues beyond opening the box. Making tea calls for removing a bag from its foil container, putting it in a cup, boiling water and pouring it. These actions demand additional structure and behavior. A reconstruction adequate for moving a closed box may be inadequate for handling its packaging or liquid. Simulation quality therefore depends on the task that must transfer to reality.
A simulator need not reproduce every detail of the world. Human world models also leave most things unmodeled while representing what matters to the current purpose. The engineering choice is where to spend fidelity: include the objects, properties and dynamics that affect the intended action, and retain control over which parts of the world receive that detail.
The closing application is an industrial process, later identified in Q&A as a conveyor belt system. A short video becomes a 3D simulation with objects and movement represented in sufficient detail to generate training data. A robotic system can then explore the simulated environment and learn how to operate in it. This is the proposed replacement for collecting 10,000 hours of real-world teleoperation.
Manning describes obtaining 10,000 hours of simulation “for free.” The useful economic claim is that repeated simulated experience can avoid the corresponding human teleoperation effort. The presentation does not quantify simulation compute costs or establish measured real-world transfer performance; cheap, effective training remains the intended payoff of building a sufficiently accurate model.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Games, symbolic interfaces and the long tail of situations
The first questions extend the approach beyond the industrial example. Several applications and representation choices share the same need for an agent to practice actions in a modeled environment:
- Games: Millions of players provide abundant behavior data, and their goals are often fairly clear. A reinforcement learning loop can use those goals to learn effective actions. Moonlake has explored gaming, although Manning says its recent emphasis is physical infrastructure.
- Symbolic interfaces: Code brings back some of the function of older knowledge representations without requiring a return to old-style ontologies. A purely neural latent representation may be powerful, but people and other applications have difficulty connecting to it. Moonlake’s bet is that symbolic representations offer a practical interface for physical AI.
- Rare situations: Humanoid robotics, space robotics and other new forms of automation need experience beyond routine cases. Collecting that experience physically is costly, especially in the long tail of unusual scenarios. Good simulation provides a way to explore those cases.
The symbolic choice is a practical bet about integration, rather than a claim that neural representations cannot work. A code-based world gives neural systems something they can reason over and plan with, while giving humans an accessible way to change the model. The older ambition of knowledge representation returns in executable form.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Physics supplies dynamics; reality tests the model
Does the simulation learn all physical behavior from video? Manning answers that it uses physics engines and knowledge of physics to generate and control movement. Water is his example: starting from an image does not mean the system must infer all fluid behavior from that image. A physics model supplies predictions about how it will behave. The neurosymbolic approach therefore combines generated representations with established physical modeling.
The next question presses on the simulation-to-reality gap: mechanical details and tolerances can make real behavior diverge from a simulation. Moonlake’s proposed response extends the earlier coding loop from appearance to behavior. Compare simulated behavior against real-world video, then use neural optimization to improve the simulation. The aim is to shrink discrepancies automatically rather than depend entirely on people hand-writing and adjusting a physics simulation. 45:52
Physical dynamics are only one layer of a working environment. An audience question asks about modeling people, processes and the decision makers who control objects. Manning says Moonlake has not really been addressing that sociotechnical layer. He sees potential in language models as human-behavior simulators because they absorb large amounts of human behavior data, but presents that as a neighboring area of development, rather than a capability established by the physical-world examples.
The final question asks whether exploration in latent space can reveal connections humans have not already identified. A sufficiently good simulation can support surprising discoveries, Manning replies, but correctness inside the simulated world does not ensure that every relevant fact about reality has been captured. Reality can also contain relationships absent from the model. The value of simulation is wider exploration with a useful approximation; its discoveries still depend on what the approximation represents. 49:16
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
From the talk
Revisit Riley Goodside’s generative-video example and the specific distinction Manning draws between visual observations and semantic state for planning.
The company behind the simulation infrastructure described here; a starting point for investigating its work on controllable worlds for physical AI.
Further reading
Research background for the NLP perspective that shapes the talk’s history of language models and its turn toward embodied intelligence.
Read the complete timestamped transcript
- 0:01
[music]
- 0:12
Okay. Hi everyone. [applause]
- 0:16
Um, and thanks a lot for making it down
- 0:19
to the second floor and finding my
- 0:21
session here. Okay. So I'm Chris Manning
- 0:24
and what I want to do today is present
- 0:26
something about Moonlakes's approach to
- 0:28
producing a simulation infrastructure
- 0:30
for practical physical AI. The overall
- 0:33
goal here is that a north star for AI
- 0:37
and actually cognitive science as well
- 0:39
has always been to understand and work
- 0:42
out how to build embodied intelligence.
- 0:45
So today I'm going to tell us about some
- 0:47
of the recent work that we've been doing
- 0:49
at approaching a practical form of
- 0:52
embodied artificial general
- 0:53
intelligence.
- 0:55
But you know I'm not actually going to
- 0:58
start there because uh you know um Swick
- 1:03
said no you shouldn't just do that. Um,
- 1:07
you should have a double slot. And first
- 1:10
of all, since I've got you here today, I
- 1:13
should have people tell you about the
- 1:15
history. You should tell people about
- 1:16
the history of AI and your journey
- 1:19
through it and how you ended up here.
- 1:21
Um, so, uh, sit back, um, get ready for
- 1:26
story time and we're going to have the
- 1:29
really slow build and we're going to
- 1:31
hear about all about the history of AI
- 1:33
for 20 minutes and then we're going to
- 1:35
hear the details about what Moon Lake is
- 1:38
doing. Um, he looks very persuasive
- 1:41
there, doesn't he? Whereas at the time
- 1:43
we were having this conversation, I was
- 1:45
having a very bad hair day. So I had
- 1:48
very little um choice um to agree to
- 1:51
this task. Um so here I am with some
- 1:54
long-term remarks on how did I and all
- 1:57
of us get to this moment from the far
- 2:00
ago days um when AI didn't work. Um so
- 2:04
the very beginning of AI was the
- 2:07
Dartmouth summer research project in
- 2:10
1956. It was for the holding of this um
- 2:13
kind of summer group project was when um
- 2:18
John McCarthy coined the term AI and got
- 2:22
together um this group of people. Um so
- 2:25
that one on the back right there, that's
- 2:28
John McCarthy, there's Minsky and both
- 2:31
most of them look like a real bunch of
- 2:33
geeks as you can see. Um but the
- 2:35
attractive one on right on the right end
- 2:38
is you know my personal hero as more of
- 2:40
a language guy um Claude Shannon.
- 2:43
So this is sort of the start of AI 1956.
- 2:48
It's where the the term AI came from.
- 2:50
But really there's other stuff that came
- 2:53
earlier than 1956.
- 2:56
So really starting in the 40s and early
- 2:58
in the 50s there was work on
- 3:00
cybernetics. So cybernetics sought to
- 3:03
tie together communications control and
- 3:06
feedback in living things and computers.
- 3:08
Um so uh you know kubernetes is really
- 3:12
the same word as cybernetics just less
- 3:15
anglicized um but has very different
- 3:18
meaning. Um yeah so this led to the
- 3:20
earliest work on neural networks and the
- 3:23
perceptron network. So here's Frank
- 3:25
Rosenlat's perceptron um that was you
- 3:29
know predated the term artificial
- 3:31
intelligence and as you can see like in
- 3:34
those days you actually used to wire
- 3:36
neural networks now this modern matrix
- 3:38
multiplication but another thing that
- 3:41
actually started before the term
- 3:44
artificial intelligence and has been my
- 3:47
long-term area of interest is doing
- 3:50
things in natural language processing
- 3:52
and natural language processing
- 3:54
began as machine translation. And so
- 3:57
here we are again in 1954,
- 4:01
two years before the term artificial
- 4:03
intelligence was coined. And here's the
- 4:06
front page of the New York Times. Now,
- 4:09
there's a way in which the front page of
- 4:10
the New York Times in 1954 is
- 4:14
surprisingly reminiscent of the issues
- 4:16
that you see in the front page of the
- 4:18
New York Times um in 2026 because it's
- 4:23
exactly the same issue. President
- 4:25
proposing ending citizenship for people.
- 4:28
Um not much has changed in the um
- 4:31
intervening 75 years. Um allegations of
- 4:35
various kinds of behavior akin to
- 4:37
treason. All sounds very familiar. Um,
- 4:40
but this was the top half of the page.
- 4:41
And if you went down to the bottom half
- 4:43
of the page, um, you found this article.
- 4:46
Russian is turned into English by a fast
- 4:49
electronic, um, translator. A public
- 4:52
demonstration of what's believed to be
- 4:54
the first successful use of a machine to
- 4:56
translate meaningful text from one
- 4:58
language to another took place here
- 5:00
yesterday afternoon. Um, and here's a
- 5:03
little bit of video of showing that
- 5:05
system.
- 5:11
into
- 5:16
one of the first nonmerica
- 5:26
were made that the computer would
- 5:27
replace most human translators.
- 5:32
>> Okay. Um, and we'll come back to that
- 5:35
again in a little bit. Um, but coming
- 5:38
off of the founding of AI, um, by John
- 5:42
McCarthy, fairly soon after that, John
- 5:45
McCarthy moved to Stanford and founded
- 5:48
Stanford Artificial Intelligence, the
- 5:50
Stanford Artificial Intelligence Lab,
- 5:53
starting from 1963.
- 5:56
And the original Stanford AI lab was in
- 5:59
this um, building up in the foothills.
- 6:01
Um if you down if you know the South Bay
- 6:04
well um if you've ever been to a
- 6:06
restraero and you look over next to a
- 6:09
restraero um where there's the Portola
- 6:12
pastures um horse area or if you ride a
- 6:15
horse um that's where this AI lab used
- 6:18
to be and that was the site of many of
- 6:20
the um founding um work in artificial
- 6:24
intelligence and also a lot of other
- 6:26
stuff. Um the Stanford AI lab was the
- 6:29
where the very first video game
- 6:31
tournament was played in 1972.
- 6:35
Um now McCarthy himself um was um
- 6:40
mathematical logician and nearly all of
- 6:43
his work was sort of building out these
- 6:45
not notions of mathematical logic but in
- 6:48
his thinking he was very wide ranging
- 6:52
and so he
- 6:55
he liked the idea of how could we build
- 6:58
an embodied artificial general
- 7:00
intelligence and was very happy to sort
- 7:03
of provide the facilities and the money
- 7:06
for other people to start to explore
- 7:08
this. Um so this was the Stanford AI
- 7:11
Labs um very first robot um the Stanford
- 7:15
cart not very fancy um the old um sale
- 7:19
building um but more relevant to what
- 7:22
we're going to talk about today for
- 7:25
simulation infrastructures for AI um was
- 7:28
another robot um Stanford SRRI shaky
- 7:32
robot
- 7:33
>> mach
- 7:37
audio Um, so this was a robot that could
- 7:41
um perceive the environment, move
- 7:44
around, move boxes and so on in the
- 7:48
environment and do things like that.
- 7:52
Okay. And so this sort of started to
- 7:56
explore this idea of having an embodied
- 8:00
in artificial intelligence that had an
- 8:03
understanding of a world and its
- 8:05
environment.
- 8:07
Um, just one little distraction that's
- 8:10
not really about AI. Um, who here has a
- 8:13
connection to Stanford? Stanford
- 8:16
connections. Who here has a connection
- 8:19
to Berkeley? Berkeley connections. Less
- 8:22
Berkeley people. Anybody with a
- 8:23
connection to both places? There are
- 8:25
some people who've been to both of them.
- 8:27
No. Okay. Um, so I just have to put in
- 8:31
my advertisement for early Stanford. Um
- 8:34
if you're an artificial intelligence
- 8:37
person you know the story in the early
- 8:40
days is all Stanford. I mean there you
- 8:43
know there are two different stories I
- 8:45
can tell right you know if you're in the
- 8:47
history of west coast universities
- 8:50
really in the first half of the 20th
- 8:52
century all the prominent stuff was
- 8:55
Berkeley and that Stanford was um this
- 8:58
sort of pokey regional school but once
- 9:01
you get to the second half of the 20th
- 9:04
century which is the era of computers
- 9:06
and AI um it all happened at Stanford um
- 9:12
So that you know if you look at things
- 9:14
like the beginnings of the internet, the
- 9:16
Arpanet, um here it is in 1972
- 9:20
and it connects up some UC campuses,
- 9:23
Berke um UCLA and Santa Barbara and
- 9:27
connects to Stanford, MIT, Harvard, but
- 9:30
no Berkeley. Go ahead to 1980. Um
- 9:34
Berkeley still isn't part of the
- 9:36
internet. you know, it's made it as far
- 9:38
as Hawaii and London and Berkeley is
- 9:40
still not there. Um, and really in all
- 9:43
of this period, there just wasn't AI at
- 9:47
Berkeley. I mean, there was other stuff
- 9:49
that went on at Berkeley, I should be
- 9:51
fair. There was important systems work.
- 9:53
There was ingress database and BSD Unix,
- 9:56
which some of you probably remember if
- 9:58
you have gray hair like me. Um, but you
- 10:01
know, essentially through the 60s and
- 10:04
70s and the first half of the 80s, there
- 10:07
was just no AI at Berkeley. And so AI at
- 10:10
Berkeley really only got underway in
- 10:12
1986
- 10:14
um when two fresh Stanford PhD grads um
- 10:17
Jendra Malik and Stuart Russell um moved
- 10:21
from Stanford um to Berkeley to take up
- 10:24
professorships. Now of course that's 40
- 10:26
years ago now. So, they've had they've
- 10:28
had um AI for a while at Berkeley now,
- 10:30
but not in the old days. Um, if you'd
- 10:33
like to know more about this history of
- 10:35
the old days, um, a couple of years ago,
- 10:38
we actually put together, um, a story of
- 10:41
the first 60 years of, um, AI at
- 10:43
Stanford. Conveniently, um, the 60 years
- 10:46
stopped just before the arrival of chat
- 10:49
GPT and large language models. Um, and
- 10:53
you can, um, find it on YouTube.
- 10:56
Um, okay. So, that was my very brief
- 10:59
potted history of AI. Um, what did I do
- 11:03
for my life? Well, what I did for my
- 11:06
life was to be a natural language
- 11:09
processing person. And so, a natural
- 11:11
language processing that was where we
- 11:14
developed language models. And language
- 11:18
models actually in some form go back a
- 11:23
very very long way. So the first
- 11:26
language model was proposed by Andre
- 11:29
Marov who invented Markoff models right.
- 11:32
So in developing the idea of Markoff
- 11:35
models he actually did it with language
- 11:38
um taking a novel of Pushkin's Eugene
- 11:41
Anagen and developed a character level
- 11:44
language model over that. Um but the
- 11:47
famous thing is then back to Claude
- 11:49
Shannon who we saw earlier who sort of
- 11:51
formalized information theory and
- 11:54
started to build um word and character
- 11:57
engram language models in the late 1940s
- 12:01
which is the dominant stuff that we went
- 12:03
on and using and in particular um then
- 12:06
in 1975
- 12:08
again so back in the sort of early part
- 12:11
of AI um a famous group at IBM Fred
- 12:15
Gelan's group at IBM sort of defined
- 12:18
well they came up with the term language
- 12:20
model that we still use today and they
- 12:23
defined this idea of probabilistic
- 12:26
models of text that could be used as a
- 12:29
basis for all kinds of speech and
- 12:31
natural language processing
- 12:33
applications. So really from very early
- 12:36
on in the speech and NLP tradition,
- 12:39
language models were seen as a central
- 12:41
technology that other people in other
- 12:43
areas of AI and um machine learning just
- 12:48
didn't really know about. And it was
- 12:50
what enabled
- 12:53
um interesting good things to happen
- 12:55
early in speech and various areas of
- 12:59
NLP. So both early speech recognition
- 13:02
systems but also the kind of um machine
- 13:06
translation that you got at Google
- 13:09
starting about 2007. It was considerably
- 13:12
powered by language models. And so I
- 13:16
think there's actually a kind of an
- 13:17
interesting story here of the
- 13:19
development of our modern AI
- 13:23
um sorry which is um slightly different
- 13:26
to the story most people remember which
- 13:30
um is a story that's dominated by the
- 13:33
vision story because the vision sort of
- 13:35
came in most people's head to be seen as
- 13:39
the sort of entry place of neural
- 13:41
networks in a large scale. Um but
- 13:44
nevertheless I'll point out the
- 13:45
limitation of this early work in
- 13:47
language models. It was used for
- 13:49
spelling correction, machine
- 13:50
translation, all of these things. But at
- 13:53
this time, nobody thought of language
- 13:56
models as these were going to solve
- 13:57
artificial intelligence. We still all
- 14:00
bought the old AI story of that we're
- 14:02
going to need memories, knowledge
- 14:04
representations, planning systems,
- 14:06
reasoning systems, all of these classic
- 14:08
AI things.
- 14:11
Yeah. So here's the sort of history of
- 14:13
large language models from an NLP
- 14:16
suspect perspective. So the very first
- 14:18
mention of the term large language
- 14:20
models was in 1998 as far as I can tell.
- 14:24
But you know really that was sort of a
- 14:26
how are we going to store all this text
- 14:29
um to build a model. So the first
- 14:31
interesting um connection of large
- 14:33
language models was in 2000 when Joshua
- 14:36
Benjio and colleagues defined new
- 14:39
probabilistic language models. the first
- 14:41
neural language model. Um, but as you
- 14:44
can see from those stats of 32 million
- 14:47
um token corpus, 31,000word vocabulary,
- 14:50
this was a teeny model because at that
- 14:53
point they just didn't have enough
- 14:55
compute to do anything interesting. Um,
- 14:58
but interestingly as early as 2007
- 15:02
at Google they were able to solve the
- 15:05
compute problem and the data problem.
- 15:08
Um, so most people forget this now, but
- 15:11
in 2007,
- 15:13
um, Google had a language model that was
- 15:16
built on two trillion tokens of text.
- 15:19
So, you know, that's a bit smaller than
- 15:22
the state-of-the-art models now that
- 15:23
might be trained on 15 trillions of
- 15:25
text, but it's actually the same order
- 15:27
of magnitude, right? So we all already
- 15:30
in sort of two decades ago there was the
- 15:34
scale of data and compute to build a
- 15:37
form of language model. The big problem
- 15:39
was they didn't have enough model
- 15:42
flexibility that the kind of modern
- 15:44
powerful neural networks that we use
- 15:46
today hadn't yet been invented.
- 15:48
Um and so it was then only in 2018 that
- 15:52
modern large language models started to
- 15:55
appear using transformers. But the early
- 15:58
versions of those um were back to small
- 16:02
amounts of data. So the early GPT model
- 16:05
only used 3.3 billion tokens. So down
- 16:08
three orders of magnitude in size again.
- 16:11
So we reverted to not enough data. And
- 16:14
then it was eventually only in 2020 with
- 16:17
GPT3 forward that all of those things
- 16:20
came together and language models really
- 16:22
took off as we know about it now.
- 16:26
Um and so that led to this surprising
- 16:29
victory of natural language processing.
- 16:31
I mean it actually gives me a bit of a
- 16:33
laugh um that if in the sort of mass
- 16:36
media if you see a reference to AI these
- 16:39
days with pretty high probability
- 16:41
they're actually going to be talking
- 16:42
about a large language model which
- 16:45
wasn't the way it used to be where NLP
- 16:47
used to be a fairly marginal area of
- 16:49
artificial intelligence. And as we all
- 16:51
know, these large language models have
- 16:54
allowed us to do amazing amazing things.
- 16:57
And so the ability of these systems to
- 17:01
reason and solve complex problems like
- 17:04
math problems is just completely
- 17:06
stunning. I mean, you know, even for
- 17:09
someone like me who's spent my 30 years
- 17:11
working in natural language processing,
- 17:13
I kind of find it hard to believe that
- 17:16
you can pick really difficult math
- 17:19
problems of a kind I certainly could not
- 17:21
solve myself and just feed it into a
- 17:24
large language model and somehow it can
- 17:27
um chunk along doing its test time
- 17:29
thinking and come up with right answers.
- 17:31
I mean it's just been a sort of a
- 17:33
stunning breakthrough from an unexpected
- 17:36
direction. Um so that's been most of my
- 17:39
world but then here we are um now
- 17:43
getting back into the main topic of the
- 17:45
talk of well where are we um why is
- 17:50
embodied artificial general intelligence
- 17:53
a north star that we should be looking
- 17:55
at and the reason for that is even
- 17:58
though it's been so amazing what can be
- 18:02
done with large language models they're
- 18:06
still this textbased description of the
- 18:09
world. And it turns out that you can do
- 18:13
a lot of stuff with a textbased
- 18:15
description of the world. I think it's
- 18:17
true that we can do just way more than
- 18:21
almost anybody believed possible with a
- 18:23
textbased description of the world.
- 18:26
Certainly a lot of prominent people in
- 18:28
robotics and computer vision spent a
- 18:30
decade saying, "Oh, you'll never be able
- 18:32
to do that with a large language model."
- 18:35
where actually we've been able to do a
- 18:36
lot of that with a large language model.
- 18:39
But still at the end of the day um we do
- 18:42
actually want to deal with the world
- 18:44
around us and have artificial
- 18:46
intelligence that can operate in our
- 18:48
world. And so the question is then how
- 18:51
can we build these embodied artificial
- 18:53
general intelligences?
- 18:55
And so gradually I started to um get a
- 18:58
bit more interested in well how can we
- 19:01
start to incorporate the visual world
- 19:04
and I started to look at things like
- 19:06
visual question answering and how you
- 19:09
could connect between um text and then
- 19:12
generative models of visual worlds and
- 19:15
reason about that. And so today I'm
- 19:18
going to be telling you more about how
- 19:20
you can build out that line of work and
- 19:23
be using simulation as a way to start to
- 19:26
approach um embodied artificial general
- 19:29
intelligence. So why do we want
- 19:32
simulation? Um there are sort of two
- 19:35
ways that you can go about um starting
- 19:38
to build intelligent models of the
- 19:41
physical world. One way of doing it is
- 19:44
that you're actually going to directly
- 19:46
learn in the real world. So you can um
- 19:50
set up hardware or collect your YouTube
- 19:54
videos in the real world and start
- 19:57
learning intelligent policies for how to
- 20:00
act. So this was the kind of approach
- 20:02
that was used in Google X's QOP um work
- 20:05
that you've got your um rows of robot
- 20:08
arms that are doing things. you're
- 20:10
recording what they're doing and you're
- 20:12
in the physical world starting to learn
- 20:15
a policy. Um, that's a really unpleasant
- 20:20
way to try and make progress. So, when
- 20:22
people are do making progress in this
- 20:25
kind of world, you're getting about
- 20:28
10,000 hours of teleyop, that means a
- 20:31
human is moving the robot around um data
- 20:35
to start to train a model. um it's not a
- 20:38
very appealing picture. If we want to
- 20:40
start having generally good robotics,
- 20:43
we're just not going to get very far if
- 20:46
we're sort of um chugging along at the
- 20:49
speed of these robots um moving.
- 20:52
But interestingly, if we go back to
- 20:55
Shaky in the 1970s,
- 20:57
Shaky had a different answer. Shaky
- 21:00
said, "Well, we shouldn't just have the
- 21:05
real world that's surrounding us. We
- 21:08
should also have in the robot's head a
- 21:12
world model that has an internal
- 21:16
representation of what the world is
- 21:18
like. So this idea of a world model as
- 21:22
an abstracted internal representation of
- 21:24
the world which can be used for
- 21:26
simulation and planning that's an idea
- 21:29
that goes back to cognitive science. So
- 21:31
it was first proposed by Kenneth Craig
- 21:34
um in the 1940s. So he argued that
- 21:38
humans and other creatures have world
- 21:41
models inside their own heads. Um so um
- 21:45
that they can think about alternatives
- 21:48
and how they're likely to play to play
- 21:50
out and therefore they can think of good
- 21:53
plans um before acting. And so that's
- 21:57
also what we'd like our robots to do is
- 22:01
to have the same kind of abilities. Um
- 22:04
so this was Shaky's world model. So,
- 22:06
Shiki had a kind of a blocks world where
- 22:08
it was actually a grid world as you
- 22:10
might remember from your early AI
- 22:12
textbooks, but it was representing um
- 22:15
what was in different places in the
- 22:17
world and what were the attributes of
- 22:20
different things and it was all stored
- 22:22
in this kind of logical representation
- 22:25
in those days. So how now can we start
- 22:30
um building a 21st century version of an
- 22:34
embodied artificial general
- 22:36
intelligence? And I think the right way
- 22:38
to do it is to work out how to build
- 22:41
good action condition world models. And
- 22:44
today I want to talk about a practical
- 22:48
effective way to build a simulation
- 22:51
infrastructure that can be used um for
- 22:54
as a basis from embodied AGI.
- 22:58
So the idea of a world model is now
- 23:00
normally formalized in terms of
- 23:02
reinforcement learning ideas that what
- 23:05
we do is we have observations of a world
- 23:08
but we assume that underlying those we
- 23:11
have a semantic abstracted
- 23:14
representation of the state of the world
- 23:16
in our head. And what the world model
- 23:19
does is gives us an ability to try and
- 23:23
predict um when an action is taken in
- 23:26
one state, what new state is going to
- 23:29
emerge. And so the crucial thing there
- 23:32
is that the world model is abstracted
- 23:36
and has more semantics. Quite a lot of
- 23:39
the time when people have talked about
- 23:42
world models, they haven't really been
- 23:45
thinking about the abstraction and the
- 23:47
semantics, they've just been talking
- 23:49
about can you produce beautiful
- 23:52
generative AI video. Um so here's um
- 23:56
Genie 3. Um this is kind of a cute
- 23:59
example by Riley Goodside who was the
- 24:01
same person who got a lot of fame in the
- 24:04
early LLM days by being a good prompt
- 24:07
engineer. um and he's now playing around
- 24:10
um here with Genie 3. And you know the
- 24:13
the kind of things you can do with Genie
- 24:15
3, you know, it looks beautiful, it's
- 24:18
wonderful, it's pretty fluid, it seems
- 24:23
great, but you know, most of visual AI
- 24:26
for this period has been just judged by
- 24:29
the pixels. If you've got beautiful
- 24:32
pixels, you've got a beautiful piece of
- 24:34
software, but these pixels are trying to
- 24:37
simulate observations. They don't
- 24:40
actually have or represent any of the
- 24:43
semantics behind the world. And
- 24:45
therefore, they aren't good at realism.
- 24:48
They aren't good for giving a basis to
- 24:50
plan. They aren't good for actually
- 24:52
having a robust simulation
- 24:54
infrastructure of the real world. And so
- 24:57
we want this kind of simulation
- 24:59
infrastructure because with a good
- 25:01
simulator we actually have causal
- 25:03
knowledge of how the world works which
- 25:06
allows us to predict and plan how things
- 25:09
will work in any situation. And so
- 25:13
that's the kind of world um that we're
- 25:16
wanting to have and make available at
- 25:19
Moon Lake. So the starting point is a
- 25:23
real world observation. Um, so given an
- 25:27
image or a bit of video, we want to be
- 25:30
able to interpret this image as an
- 25:33
observation which is always partial of
- 25:36
what's actually in the underlying world.
- 25:39
And then what we want to do is
- 25:41
reconstruct a model of this world which
- 25:45
actually allows us to do stuff in it and
- 25:48
see how it reacts. that this will give
- 25:50
us a basis of being able to work out
- 25:54
causality, work out how to plan and
- 25:57
reason in a repres representation
- 26:00
condition way which gives us the basis
- 26:02
for intelligent robot actions inside
- 26:06
this world.
- 26:08
So what might one do here? Well, the
- 26:11
first thing you can do is take that um
- 26:14
picture and feed it into one of our
- 26:17
well-known um other companies products
- 26:20
and say, "Okay, make a simulation of
- 26:22
this world." Um and so this is what you
- 26:26
get from Marble. And it's a pretty good
- 26:28
simulation of the world. And if what you
- 26:31
want to do is just uh walk around in
- 26:34
this world and see it from different
- 26:36
angles, this works pretty well. Um the
- 26:42
question is is this a good simulation
- 26:44
and whether it's a good simulation
- 26:47
depends on
- 26:50
what problem you want to solve in the
- 26:52
real world. And then is does this give
- 26:55
you sufficient information to solve the
- 26:58
problem in the real world? And if all
- 27:01
you want to do in the real world is to
- 27:03
be able to wander around and see the
- 27:05
view from different angles, then this is
- 27:08
great. We're done. But a lot of the time
- 27:11
what you'd like to do in the real world
- 27:15
is understand the objects are here and
- 27:18
to be able to do things with them.
- 27:21
Right? So there's some objects here.
- 27:22
There's a kettle and there's a box and
- 27:25
there's a cup. But in the marble world,
- 27:29
there's nothing you can actually do with
- 27:31
these things. All you can do is sort of
- 27:33
wander around as a disembodied figure.
- 27:36
Um, and so for a lot of purposes such as
- 27:39
doing things with robotics or other
- 27:42
physically accurate worlds, um, we need
- 27:45
to be able to have more understanding in
- 27:47
our model of the world, a better, more
- 27:50
detailed simulation. And so at Moon
- 27:53
Lake, we're wanting to work out what
- 27:55
people actually want to do with their
- 27:57
simulation, and then to build the kind
- 27:59
of simulation that will power that. So
- 28:02
you might think, oh, I actually want to
- 28:04
be able to move around the objects in
- 28:06
the world. And so that's the kind of
- 28:09
starting point of the kind of thing that
- 28:11
we're trying to do at Moon Lake. So if
- 28:14
we want to um have a world in which we
- 28:18
can move things around then we're saying
- 28:20
okay kind of like um image blaster a
- 28:23
system like that we actually need to
- 28:26
take this world and understand what's in
- 28:29
it and then have objects that can be
- 28:31
manipulated. So we separate out a
- 28:34
background and then objects inside that
- 28:38
background with then being able to sort
- 28:40
of have the background world in a
- 28:43
representation that's similar to marble.
- 28:46
But then in the foreground there are
- 28:48
various kinds of objects that we can
- 28:50
manipulate and move around. Now, well,
- 28:53
that's a start, but um in the previous
- 28:57
picture, um the tea box was always
- 29:00
closed. And we might wonder if you can
- 29:03
open up the tea box and see what's
- 29:06
inside it. Um and well, this is sort of
- 29:09
a part of how observations of a world
- 29:12
are always partial. Um for what we could
- 29:15
see in the actual photo, there was a
- 29:17
closed tea box. We couldn't even see the
- 29:19
tea box very well. How could we possibly
- 29:22
know what's inside it? And well, the
- 29:26
answer to the way we can know what's
- 29:28
inside it in the 21st century is we go
- 29:32
off and do a little bit re of research.
- 29:34
We fire up um getting information um off
- 29:38
the web um in a
- 29:42
a rag style fashion um to work out
- 29:45
what's there. And then we can find
- 29:48
images of this um tea um box on the web
- 29:52
which show it in more detail. They show
- 29:54
a picture of what it's like when it's
- 29:56
open. There are descriptions of it.
- 29:58
Organic luxury tea bag collection,
- 30:00
leather gift box. It gives its size. We
- 30:03
can find all about it. So therefore, we
- 30:05
can build um in our simulation um a tea
- 30:09
box with an understanding of the
- 30:11
contents of that tea box. Well, that's
- 30:14
really good. Um, so how now we have um
- 30:17
the tea box with tea bags inside it. But
- 30:20
if we do nothing else, um, they're just
- 30:24
sort of sitting there and movable. So
- 30:26
we'd like to realize the fact, well,
- 30:28
wait a minute, tea bags you can lift up
- 30:31
and you can take out of a tea box. So
- 30:35
then we need to start having these
- 30:37
teaags also um be objects that are
- 30:40
modeled in our simulation that they have
- 30:42
a size and an ability to be manipulated
- 30:46
as well. Um so how are we doing all of
- 30:50
this? And so a distinctive part of
- 30:52
what's happening at Moon Lake is
- 30:55
building although part of this is in the
- 30:59
world of vision
- 31:01
a lot of what we're doing is actually
- 31:03
back in the world of code. So this is
- 31:06
picking up on the idea um that what's
- 31:10
normally referred to as language models
- 31:12
but these days are really symbolic
- 31:14
models which work on not only human
- 31:17
languages but also math code and things
- 31:20
like that that they that has been just a
- 31:23
very powerful substrate with which to
- 31:26
make progress. And so we are producing
- 31:29
controllable manipulable world models by
- 31:33
generating the controllable parts of
- 31:36
these worlds by putting code under them.
- 31:40
And in particular, um, we can train
- 31:43
these code models using the same kind of
- 31:46
loop engineering that many of you will
- 31:48
have seen, um, with Claude that we're
- 31:51
having a loop where we're writing code
- 31:55
to render objects,
- 31:58
um, and put on that um, diffusionbased
- 32:01
textures, etc. We can then assess how
- 32:05
good our render is against physical
- 32:08
reality and then we can work out the
- 32:11
deviations. We can revise the code and
- 32:15
make better and better um renders of
- 32:17
what goes along. And so that's the way
- 32:20
that so we're building this detailed
- 32:24
action conditional world model so that
- 32:26
we can take actions on the objects in
- 32:29
the world.
- 32:30
Um so that means that we have a kind of
- 32:33
neuros symbolic representation of the
- 32:36
world here and neuros symbolic
- 32:38
representations have a big advantage
- 32:41
that the symbolic representation can be
- 32:44
easily interfaced controlled edited and
- 32:47
maintained in our human world. Um, and
- 32:51
so the message I'd like to give as a
- 32:56
little delta of a message here, um, is
- 32:59
most of you are probably familiar with
- 33:02
Mark Andre's famous statement 15 years
- 33:04
ago, software will eat the world. And
- 33:08
you know, that was mostly right. Um but
- 33:12
I think it wasn't completely right
- 33:15
because it's just not the case that
- 33:18
software ate the physical world. Um it,
- 33:22
you know, a lot of the world went
- 33:24
virtual and um a lot of the world is
- 33:27
inside computers now and yeah it could
- 33:30
eat all of the processes of managing,
- 33:34
counting, supplying all um records of
- 33:37
suppliers, all the stuff that's in the
- 33:40
virtual world. But we hadn't had the
- 33:43
power for software to eat the physical
- 33:45
world. Whereas the hope is with this new
- 33:49
ability of being able to build
- 33:52
verifiable simulations powered by code
- 33:55
of the sort that I've just sketched for
- 33:57
a moment there that this will allow us
- 34:00
to actually have verifiable simulation
- 34:03
which will also eat the physical world.
- 34:06
And so that's the kind of thing we
- 34:08
built. And we can go on from here and um
- 34:11
you know keep on making further steps of
- 34:14
this. Right? So we maybe don't want to
- 34:15
only have tea bags um in foil
- 34:19
containers, but we'd like to be able to
- 34:21
um take them out of the foil container
- 34:24
and we don't we somehow want to actually
- 34:27
get the tea bag inside the cup because
- 34:29
that's a useful step for making tea. And
- 34:32
then of course we also want to have um
- 34:35
the jug of um water which we want to be
- 34:38
able to boil and then be able to pour
- 34:40
that into the tea. Um and once we have a
- 34:45
good world simulator like that, the idea
- 34:48
of this is that we are building in all
- 34:52
of the parts of the simulated world
- 34:56
which will allow effective transfer into
- 35:00
the real world. And it's important to
- 35:02
think about that as to sort of what's
- 35:05
necessary for different kinds of
- 35:07
transfer. And the argument um that I
- 35:12
we'd like to make is that you know
- 35:14
normally uh simulation isn't complete
- 35:19
and accurate in every detail because
- 35:22
there are sort of parts of the world
- 35:24
that are important to you and parts of
- 35:27
the world that aren't important to you.
- 35:29
I mean, this is the same with human
- 35:31
world models, right? That a lot of the
- 35:33
world we're not actually modeling at any
- 35:35
time, but we're modeling the bits of it
- 35:37
that are important to us. And it's
- 35:39
having that control of which things you
- 35:42
need to model and use is what needed
- 35:45
then to give you the actual simulations
- 35:48
that will be allowed to be effectively
- 35:50
used um in the real world. So, what kind
- 35:53
of applications
- 35:55
um is this going to allow us to do? So
- 35:58
the hope is that um rather than having
- 36:02
to collect 10,000 hours of data in by
- 36:07
teley op in the real world, we can
- 36:09
instead build accurate simulations which
- 36:12
will allow transfer to real. So let me
- 36:15
show you the kind of progress that we've
- 36:17
been making on this. So this is um not
- 36:20
highuting robotics um but is the kind of
- 36:23
physical AI that you find everywhere in
- 36:27
the real world that there are things
- 36:29
happening in physical processes um which
- 36:32
you'd like to be able to um understand
- 36:37
and automate. And so in particular, um,
- 36:40
if we're going to realize any of the
- 36:42
dreams of bringing back America as a
- 36:45
great manufacturing economy, it seems
- 36:47
like we have to work out how to be able
- 36:49
to automate much more in the physical
- 36:52
world. Well, um, using the Moonlake AI
- 36:56
technology, what we can then do is say,
- 36:59
so from this short video, we can
- 37:02
automatically produce an accurate
- 37:05
simulation of that world, turning it
- 37:07
into a 3D world with enough detail about
- 37:11
how things move and what objects are in
- 37:14
the world that we can start to build on
- 37:17
this and use it as a basis of simulated
- 37:20
data which transfers refers accurately
- 37:23
to the real world. So we then have this
- 37:26
model which allows us to get 10,000
- 37:28
hours of simulation um for free. And so
- 37:32
on the basis of that um we can then
- 37:35
train up a robotic system that can
- 37:38
operate um to actually um explore what
- 37:44
you can do in this world and learn the
- 37:47
way to operate in the world so it works.
- 37:50
and then we can have the effective and
- 37:53
cheap training of robotic systems.
- 37:57
Okay. Um so that's the story. Um thanks
- 38:00
a lot everyone.
- 38:02
[applause]
- 38:09
And I do have time for questions I
- 38:11
believe.
- 38:16
>> Thank you very much for the
- 38:17
conversation. Um, super deep. My first
- 38:20
question is, would this be applied to
- 38:23
something like the gaming industry and
- 38:26
how could they use it for we're seeing
- 38:28
GTA 6? Everyone is talking about this
- 38:31
new game. Um, I could see a very
- 38:34
application there where you could have
- 38:36
endless interaction with the world model
- 38:39
that they're building. U, have you been
- 38:41
working with the gaming industry at all
- 38:42
or this is more applied to robotics?
- 38:44
>> Um, so yeah, absolutely. Um, another
- 38:49
hugely good area for this is in the
- 38:51
gaming industry. And you know, the real
- 38:55
world is then kind of a virtual world,
- 38:57
but you can have a simulation of your
- 38:59
gaming world um, and then be building
- 39:02
the same kind of embodied intelligence.
- 39:06
And the gaming industry has some really
- 39:08
appealing attributes. Well, you know,
- 39:10
there are millions and millions of game
- 39:12
players, so there's lots of easy data to
- 39:14
collect. Um the people who are gaming
- 39:18
have goals which are fairly clearly
- 39:20
known. So you can sort of learn in a
- 39:22
reinforcement learning loop good ways to
- 39:25
act. And so absolutely one of the
- 39:27
applications that Moon Lake has explored
- 39:30
is in the gaming industry and there's
- 39:32
likely to be more of that. Um but
- 39:34
recently we've been actually
- 39:37
particularly emphasizing physical
- 39:39
infrastructure and doing simulations in
- 39:40
the physical infrastructure world.
- 39:42
>> Thank you so much.
- 39:55
Uh so it's fascinating to me that that
- 39:58
you're turning back to neurosy symbolic
- 40:00
representations and um the part of the
- 40:04
original kind of old school AI involved
- 40:07
a lot of ontology development and
- 40:09
knowledge representation and logic and
- 40:11
representation of that. Do you see a
- 40:13
role for uh deepening the neuros
- 40:16
symbolic representation using kind of
- 40:19
like what we how ontologies have been
- 40:21
developed previously to model um a
- 40:24
representation of the real world? Um and
- 40:27
do you see that as being a kind of an
- 40:28
area of expansion and development for
- 40:31
the type of tools that you're working
- 40:32
in? Yeah, I mean I don't think we're
- 40:36
quite going to go back to old style
- 40:38
ontologies and knowledge bases, but I
- 40:41
mean effectively
- 40:43
that is what we're doing, you know,
- 40:46
apart from it's in new clothing of um
- 40:49
having it being um code that I mean I I
- 40:55
do you know this is an interesting space
- 40:58
and we can talk about it in more general
- 41:00
right there is you know there's sort a
- 41:03
purely neural approach in which your
- 41:05
latent representation of the world is
- 41:07
purely neural and so that's the kind of
- 41:09
thing that the jeoper architecture is
- 41:12
after I mean in some ways that's a good
- 41:15
pure neural approach but on the other
- 41:18
hand that's very hard for humans to
- 41:24
connect to in any way or for other
- 41:26
applications to connect to in any way
- 41:28
and so our bet is that the sort of
- 41:31
practical ical way to have a physical AI
- 41:34
simulation infrastructure for the
- 41:36
foreseeable future is to base it on
- 41:39
symbolic representations.
- 41:41
Taking advantage of sort of the huge
- 41:43
power of code that we see everywhere
- 41:47
around us in the sort of um codeex cla
- 41:50
code era that those kind of symbolic
- 41:53
representations allow um neural eye
- 41:56
systems to reason, plan and do all of
- 41:59
these things excellently well. And so to
- 42:01
some extent yes that will be bringing
- 42:03
back notions like older fashion
- 42:06
knowledge representation.
- 42:14
Mic keeps on.
- 42:21
Hello. Hello. Oh hi. Um so a lot of our
- 42:25
existing um it seems like industrial
- 42:28
robotics use cases or to sort of
- 42:30
automate existing sort of human manual
- 42:34
intensive processes. Um would investment
- 42:38
in
- 42:40
development of like world models would
- 42:42
that allow us to go explore novel new uh
- 42:46
processes or activities where they're
- 42:49
currently not accessible by human labor
- 42:52
or by even the kind of existing um
- 42:55
industrial manufacturing processes that
- 42:57
we have.
- 42:58
>> I mean sure absolutely. Yeah. So for the
- 43:01
example at the end I showed a very old
- 43:04
school um conveyor belt system but I
- 43:08
mean you know we're also in this world
- 43:11
of amazing things happening in humanoid
- 43:14
robotics and for any of the new forms of
- 43:20
robotics automation
- 43:23
other cases you could think about doing
- 43:25
things in space well you'll have exactly
- 43:27
the same problem that You want to train
- 43:31
up AI agents to be able to act in
- 43:35
different scenarios and it's extremely
- 43:40
costly and difficult to do that training
- 43:43
in the real world. Um and it's very hard
- 43:47
when training in the real world to get
- 43:49
out into the sort of tale of rare cases.
- 43:53
I mean you've seen that um for
- 43:55
autonomous driving, right? There's a
- 43:57
reason why autonomous driving kind of
- 43:59
took 20 years to arrive. Um from you
- 44:02
know Stanley breakthroughs of yay
- 44:05
autonomous car wins the race to actually
- 44:07
having Whimos um widely deployed is
- 44:10
because you have this enormous tail and
- 44:12
the way to get out to that tail for all
- 44:14
of these new applications with um
- 44:17
humanoid robotics, space robotics, etc.
- 44:20
is to have good simulation.
- 44:26
>> Hey, uh thank you for your talk. I love
- 44:28
the energy. Uh my question is you showed
- 44:32
a video of water. Uh so is everything
- 44:35
like learned from videos or is like
- 44:37
physics ingrained in anything like for
- 44:39
example does your model know any
- 44:41
physical properties of water like
- 44:42
viscosity, Bernali's laws or like for
- 44:45
example friction. uh for some use cases
- 44:48
like self-driving maybe those things are
- 44:50
like not important right you just need
- 44:52
to know the velocity and stuff but like
- 44:54
for other use cases I would imagine
- 44:55
these physical properties are important
- 44:57
so
- 44:57
>> and we absolutely know physics I mean
- 44:59
this is the sense in which this
- 45:02
neurosyolic approach is you know you
- 45:06
call a conservative approach if you will
- 45:09
you know that this is absolutely using
- 45:11
physics engines and knowledge of physics
- 45:15
um to in its generation and control of
- 45:19
movement, right? That if you're what you
- 45:22
know when you've just started off with
- 45:24
an image of water or and you want to
- 45:27
understand how that's behaves, you're
- 45:30
using a physics model to predict how
- 45:32
it's going to behave.
- 45:36
Okay.
- 45:40
Um I know is there time for Oh, there's
- 45:44
someone else at the mic.
- 45:47
Hi.
- 45:47
>> Hey. Uh, Professor Manning, thank you
- 45:49
for this presentation. Um, I have a
- 45:52
question about seem to row gap. So, I
- 45:55
mean, simulation is great. It's cheap,
- 45:57
is scalable. Um do you have a suggestion
- 46:01
for um how we close the simulation to
- 46:04
real gaps
- 46:06
[snorts] encounter you know like the
- 46:09
like fix part that we are not able to
- 46:12
accurately simulate and like mechanical
- 46:16
tolerancing back latches like all those
- 46:19
things that we are not able to put into
- 46:23
the simulations.
- 46:24
Um yeah so traditionally the problem has
- 46:27
always been the simtoreal gap and the
- 46:30
simtoreal gap seeming too large so that
- 46:34
a lot of the training has had to happen
- 46:38
in the real world
- 46:41
and we believe that the answer to that
- 46:44
is you know effectively claude code loop
- 46:49
we're now in this world in which we can
- 46:52
do neural optimization.
- 46:55
So we can have the simulation
- 46:59
compared to behavior in the real world
- 47:02
with video and we can automatically
- 47:06
learn to shrink the sim to real gap in a
- 47:10
way that just wasn't possible when you
- 47:12
had people trying to handr write a
- 47:16
physics simulation of something.
- 47:22
This this was a great discussion. Thank
- 47:24
you. Um tangential to the question that
- 47:27
was just asked, how do you see like what
- 47:30
do you see the biggest challenges are to
- 47:32
adding like the sociote techchnical
- 47:36
layer, the pieces, the the processes,
- 47:38
the people, the authorities that make
- 47:40
decisions on the objects that you're
- 47:42
looking to simulate. What are the
- 47:44
biggest challenges to developing a
- 47:46
system that has accurate representation?
- 47:49
So for example, if we want to model uh
- 47:53
EV tolls or um doing you know
- 47:56
forecasting of energy at airports uh you
- 47:59
know things like that not necessarily
- 48:00
extending beyond the robotics into these
- 48:04
other um use cases areas.
- 48:06
>> Um yeah that's a great area. Um in all
- 48:09
honesty it's not something we've been um
- 48:11
really dealing with at Moon Lake. Um,
- 48:14
you know, at the there are still some
- 48:17
limitations, but you know, large
- 48:20
language models with their slurping up
- 48:23
of enormous amounts of human behavior
- 48:26
data that they're actually getting
- 48:30
better and better at being able to
- 48:32
simulate how human beings, different
- 48:35
kinds of human beings are going to
- 48:38
behave and react in different
- 48:40
circumstances.
- 48:42
So I think we are starting to approach
- 48:45
the point in which we can have fairly
- 48:49
good human behavior simulators that are
- 48:53
being powered by the knowledge of large
- 48:55
language models and there are a couple
- 48:57
of companies that are now starting to
- 48:58
look at that.
- 49:02
>> Cool. Last question. Yeah.
- 49:09
person over here has one for a long
- 49:11
time. [laughter]
- 49:16
>> Thank you. Um so if you look at um
- 49:20
analog chips for instance right the
- 49:22
physics is understood but as you go
- 49:24
higher up the chemical operations then
- 49:27
the heat and those things are still
- 49:29
people are working on it. So that's not
- 49:31
what you call completely
- 49:33
um understood. Now I'll give you another
- 49:36
example of say I'm looking at kidney
- 49:38
related uh literature and stuff and then
- 49:41
liver related stuff
- 49:43
these things it gives a pretty good
- 49:45
answer the kidney stuff it gives answer
- 49:47
but when you look at the relations
- 49:49
between the two people have somewhere
- 49:52
u doctors have made a link and that's
- 49:54
why you are able to see my take is if we
- 49:58
operated completely in the
- 50:02
latent space and we did not know the
- 50:04
connection connections at a human level.
- 50:06
Can we explore in the latent space
- 50:08
completely and be able to find hidden
- 50:10
connections over there that could
- 50:12
translate into the real world or am I
- 50:15
thinking something crazy?
- 50:16
>> Yeah, I mean absolutely. I mean to the
- 50:19
extent that we have a pretty good
- 50:21
simulation, we can hope to find
- 50:26
surprising discoveries that turn out to
- 50:28
be correct in the simulated world. I
- 50:31
mean, you know, things go both ways,
- 50:35
right? There are also likely to be
- 50:37
things that turn out to be true in the
- 50:39
real world that weren't in our
- 50:41
simulation, right? There's this famous
- 50:42
statement about all models are wrong,
- 50:44
but some models are useful, right? And
- 50:47
that so anybody's world model,
- 50:51
regardless of whether it's the one we're
- 50:52
building or the world model in a human
- 50:55
head, right, they're not always right.
- 50:57
Sometimes we think a person's going to
- 50:59
react in one way and they react in
- 51:00
another way. But nevertheless, a lot of
- 51:03
the time it can let us explore much more
- 51:06
widely and discover new facts and new
- 51:08
connections that we weren't aware of in
- 51:10
the real world.
- 51:13
Okay,
- 51:16
>> awesome. That's a wrap everyone. Thank
- 51:17
you so much, Chris.