AI Engineer World's Fair 2026
From VLM/VLA's to Embodied Agents — Armen Aghajanyan, Perceptron AI
Read the talk
From VLM/VLAs to Embodied Agents
Armen Aghajanyan explains Perceptron AI’s approach to combining perception, reasoning, and control: learn useful visual targets, spend compute on relevant tokens, and use video pretraining to reduce the need for expensive robot demonstrations.
From a talk by Armen Aghajanyan
At a glance
Ideas worth remembering
Visual supervision needs useful targets as well as density. Predicting every pixel can spend learning effort on background details; Perceptron proposes automatically learning future percepts, without disclosing the objective here.
Learned token routing makes compute allocation task-dependent: a general question spreads attention across an image, while fruit segmentation concentrates more tokens on likely fruit.
Perception can itself involve actions. Tiling, zooming, changing contrast, and revisiting video intervals let a model gather better evidence before producing a box or annotation.
Joint training reportedly lets 10× more video pretraining substitute for 10× less teleoperation data within Perceptron’s tested compute range, offering a way to reduce dependence on demonstrations costing around $100 per hour.
Combining reasoning and control enables tasks such as sorting books by their titles, but temporal reliability and resistance to severe visual disruption remain open problems.
Perception, reasoning, and action in one model
A camera keeps producing observations; a robot has to turn those observations into decisions and movements. Perceptron AI’s co-founder and CEO, Armen Aghajanyan, starts with that combined problem. The goal is physical intelligence that can perceive, understand, and interact with the world in real time, whether attached to a robot, an instrument, or a sensor.
The architectural starting point is early fusion: bring modalities together as early as possible. Aghajanyan connects this to his work scaling multimodal recipes at FAIR. The difficult part is choosing representations that work together on both sides of the model—what it receives and what it produces. Joining inputs early does not make text, images, video, and actions interchangeable; their representations still have to preserve what matters about each.
The familiar model categories describe different input and output capabilities:
- Vision-language models (VLMs): receive images or video alongside text and produce text. Embodied reasoning variants can also produce grounding points and handle spatial questions.
- Vision-language-action models (VLAs): extend the output space to actions, commonly using a VLM backbone.
- World models: produce video conditioned on inputs that may include images, video, and actions.
- Semantic world models: learn representations intended to be useful later, rather than directly producing a visible output.
An embodied foundation model is Perceptron’s framing for putting standard perception, embodied reasoning, and control within one model. It should reason across sensory inputs and support the relevant outputs together. That ambition immediately raises two training problems: how to learn enough from visual data, and how to process a stream that never stops.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A million visual tokens, very little supervision
Depending on the representation, one hour of video can bring roughly 1 million visual tokens into a model. The training target may be only a transcript or answers attached to a few synthetically labeled frames. In Aghajanyan’s example, the loss is calculated on something like 2% of the incoming token count. The model processes a large visual sequence while receiving a much smaller amount of explicit target information.
Predicting every pixel supplies a dense target, but density alone does not allocate learning effort well. A background pixel receives the same importance as a gripper tip, a contact point, or evidence of a physical failure. For manipulation, those details have different consequences: the gripper and contact geometry can determine whether an action succeeds, while much of the background contributes little to that decision.
Perceptron’s proposed perceptive objective predicts percepts expected to matter in the future. A hand-chosen version might predict a robotic gripper’s future tip position: a useful target, but one whose importance the designer has already specified. The research question is whether the model can learn which semantic percepts deserve prediction automatically. Aghajanyan says Perceptron has a method, but does not disclose its construction here; the talk explains the target-selection problem rather than a reproducible objective.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Let the task decide where compute goes
Always-on cameras create the second problem: context bloat. Robots do not pause their observations while a model catches up. Video supplies many tokens, but useful information can be sparse relative to text. The first practical response is compression. Perceptron began by averaging patch representations across images or video, which Aghajanyan says can achieve up to 10× compression. It works, but he regards that fixed compression strategy as a hack rather than a sufficient architectural answer.
Data sparse mixture of experts gives the model a learned choice about which tokens to process. A router predicts which tokens to admit and which to skip across the model’s layers. This addresses a different question from the perceptive objective: the objective concerns what the model should learn to predict; routing concerns where the model spends computation while processing its input.
The described compute visualizations show two useful behaviors:
- Information-driven allocation: within a figure, the model concentrates on the graph, a region containing dense information.
- Task-driven allocation: a general question spreads computation across an image; asking to segment all the fruit shifts more tokens toward regions the model considers fruit.
What changes when the question becomes specific? The diagram follows the fruit example: the visual input stays the same, while the requested task changes the allocation. A general question gives the model little reason to favor a particular region. Segmentation supplies that reason, and computation concentrates on likely fruit. The architectural choice enables selective processing without hardcoding fruit locations or a fixed region of interest.
No specific object class is requested.
The router selects or skips tokens across layers. Task-specific prompting changes which image regions receive more computation.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Finding a bird by changing how the image is examined
Perceptron combines these ideas in a model trained on a reported 1 petabyte of text, images, videos, and trajectories. The trajectories include desktop use and video games; the collection combines internet crawls, custom training recipes, and synthetic pipelines. Aghajanyan reports performance exceeding Gemini’s embodied reasoning model at roughly 15× lower cost. These are his reported comparisons; the talk does not specify the benchmark conditions or cost basis needed to generalize them.
One resulting capability turns object detection into an agentic task. The model can write code, request a closer view of part of an image, and change contrast. A difficult image therefore becomes something it can investigate through successive operations, rather than something it must resolve with one bounding-box prediction.
The bird-finding example makes that investigation concrete. The model decides to tile the image, adjusts contrast, and proposes a candidate box. Aghajanyan then describes a contrast increase that lets it find the bird. The observable change is in the evidence available for detection: subdividing the image makes regions easier to examine, and changing contrast makes the target easier to distinguish. The final box follows those inspection steps.
The important connection is between perception and action even before a robot moves. Here, actions alter the model’s view of the input. Detection becomes a small investigation in which the model chooses how to look, then uses the resulting view to locate the object.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Long tasks need decomposition—and ways to check the evidence
Physical tasks add another design choice: how much work belongs to a single action policy, and how much belongs to an orchestrator. Making coffee might take three minutes. One approach asks a VLA to handle the entire task. Another uses an embodied reasoning model to break it into subtasks, with a tactile control policy handling execution. Aghajanyan presents a spectrum between these arrangements and treats embodied reasoning itself as an unsolved problem.
Robotic video annotation shows a related use of reasoning without physical execution. The model moves among different portions of a video, clips relevant intervals, checks whether captions match them, and verifies its own annotations. This extends the bird example’s inspection behavior into time: the model chooses which evidence to revisit instead of accepting an initial description unchecked. Aghajanyan connects the practicality of repeated inspection to the model’s speed and lower cost.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Trading video pretraining for expensive teleoperation
The largest research result Aghajanyan reports concerns how training sources substitute for one another. The unified model mixes control-based training, trajectories, perceptive training, and embodied reasoning. With suitable objectives and mixing, Perceptron finds a lever that a pure policy-training setup does not expose as strongly: general video pretraining can reduce the amount of teleoperation data needed.
Teleoperation data costs on the order of $100 per hour in his account. The same budget can collect substantially more video pretraining data. Pure policies still benefit from additional teleoperation, but joint training opens a different spending decision: acquire more general visual experience instead of buying every increment of capability through robot demonstrations.
The reported trade is 10× more video pretraining data for 10× less teleoperation data. That relationship has held within the compute range Perceptron tested; the talk does not establish the task metric, experimental setup, or behavior beyond that range. Its practical significance is the possibility of reducing dependence on costly demonstrations while retaining control training, rather than eliminating demonstrations altogether.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Reading a book title becomes a control decision
The control demonstrations return to the original ambition: one model performs embodied reasoning and emits control tokens. The book-sorting task requires more than moving an object. It reads a book’s title, uses knowledge of what kind of book it is, and assigns it to an appropriate bin. Perception supplies the title; reasoning determines the category; control carries out the placement.
Where does understanding enter the movement? The diagram separates the dependencies within the task, while the demonstration uses a single model. Choosing a bin depends on understanding the book, so a control model unaware of the necessary perceptual work has a harder problem than a policy that only needs to move an already classified object.
The motion is visibly jittery in Aghajanyan’s description; improving it through scale remains a hope. He also reports relatively good zero-shot behavior and anticipates opening a smaller model in July. That is an announcement made during the talk, not confirmation of a subsequent release. Larger model weights are offered to a limited set of partners.
Perception extracts information from the book.
These are explanatory stages of the task, not separate deployed models. A title must inform a category before control places the book in a bin.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Temporal understanding still needs context management
The first Q&A answer narrows the deployment claim. Temporal understanding is not yet solved to a reliably deployable level. Even a context of roughly 1 million tokens can fill readily with high-frame-rate video. Aghajanyan describes keyframes versus delta frames as an earlier context-management approach Perceptron used and moved beyond: the decision is how much complete visual state to retain versus how much change to represent.
Pretraining also needs data that teaches relationships useful to robots. Left, right, and below are simple examples, yet internet crawls do not necessarily label them explicitly. Aghajanyan connects early attention to the data distribution with learning these capabilities quickly. Large quantities of general data do not automatically supply every spatial distinction a robot needs; the training mix must make those distinctions learnable.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Background changes, missing cameras, and structured extraction
Background robustness provides a concrete test of what the policy has learned. Aghajanyan describes fine-tuned action policies that fail when the table’s background changes. Joint perceptive and control modeling has made Perceptron’s models more tolerant of background changes and modest lighting differences in his observations. That tolerance has limits: he expects a flashlight directed into an arm camera could still break the behavior.
Training adds explicit variation alongside joint modeling:
- Camera loss: online augmentations simulate an arm camera being off.
- Lighting changes: augmentations simulate sunlight arriving from a particular direction.
These examples expose the model to altered observations during training. Aghajanyan credits the larger gains to bringing early fusion into robotics, with augmentation contributing additional robustness.
The final question asks about knowledge bases. The useful capability identified here is captioning and deep structured extraction from images and video, including complex egocentric annotation for robotics partners. Building an ontology is a separate undertaking. The model’s role is to turn difficult visual evidence into detailed structured information—the same kind of inspection and annotation work demonstrated earlier. Aghajanyan closes by saying the demonstrated capabilities are available through public APIs and that benchmarks are public.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
From the talk
A contact route for the limited partner access to larger model weights offered near the end of the talk.
Read the complete timestamped transcript
- 0:13
Yeah, I guess uh let's get started. I'm
- 0:15
I'm Armen. I'm I'm the co-founder and
- 0:18
CEO of Perceptron. Um, we'll get a
- 0:20
little bit into into what we do, but uh,
- 0:23
primarily what I want to talk about
- 0:24
today is kind of our research stance
- 0:27
that we want to move away from
- 0:29
distinctions between VLMs, VAS, world
- 0:32
models, whatever you want to call it, to
- 0:34
something that we call embodied
- 0:35
foundation models.
- 0:42
So specifically what we what we do at
- 0:45
Perceptron kind of our northstar really
- 0:46
is we want to be able to build physical
- 0:49
AI foundations that give us the ability
- 0:51
to perceive, understand and interact
- 0:53
with the physical world in real time.
- 0:55
And so kind of the the northstar mission
- 0:57
is really to bridge the the physical and
- 0:59
digital world uh worlds fundamentally
- 1:02
meaning that our our goal is kind of
- 1:04
wherever there is a um a device, an
- 1:07
instrument, a robot, a camera, a sensor,
- 1:10
uh we're essentially there providing
- 1:11
intelligence to it.
- 1:14
And specifically when we talk about this
- 1:16
paradigm of being able to perceive,
- 1:18
being able to reason, being able to act,
- 1:20
we view this as kind of a a unification
- 1:23
of uh traditional multimodal modeling.
- 1:25
So I came from fair, I was there for six
- 1:28
years and and and my target there was
- 1:30
really to try to figure out how to scale
- 1:32
up recipes for multimodal models. And so
- 1:35
one of the early things that we we
- 1:36
started working on and we've published a
- 1:38
lot in this domain was around early
- 1:39
fusion. So this concept that you want to
- 1:41
bring in all these modalities as early
- 1:43
as you can and and the real complexity
- 1:46
there is trying to figure out what is
- 1:47
the correct way to actually properly
- 1:48
represent all the different modalities
- 1:51
uh both on the input and on on the
- 1:53
output that you want to be able to
- 1:55
represent holistically and so VLMs have
- 1:57
kind of became the when I talk about
- 1:59
multimodal models um you probably think
- 2:01
of VLM. So this is the ability to take
- 2:03
in some image video and some text and be
- 2:05
able to output some some text
- 2:07
essentially. And there's there's
- 2:09
variations of this. There's there's
- 2:10
models like the ER models, the embodied
- 2:12
reasoning models that are able to output
- 2:14
maybe some grounding points um that are
- 2:17
able to do spatial understanding or
- 2:18
reasoning a little bit better. Then we
- 2:20
have things like VALAS that extend kind
- 2:22
of the output domain apart from just
- 2:24
text to now actions. And these are
- 2:26
traditionally of course continue to be
- 2:28
built on by standard VLM backbones.
- 2:30
Although there's been some efforts to
- 2:32
try to migrate away from uh from VLMs to
- 2:36
things like world models or world action
- 2:38
models, although nothing that's been
- 2:39
super fruitful just yet. And then we get
- 2:42
into kind of more interesting and
- 2:44
complex uh variants of multimodal models
- 2:46
like world models where you essentially
- 2:48
are outputting video uh from some inputs
- 2:51
and the inputs can be either image or
- 2:52
video or honestly image video and and
- 2:55
actions. Um and the last point that I'll
- 2:57
talk about is something recent which we
- 2:59
call semantic world models which is you
- 3:01
don't really output anything but you
- 3:03
learn some type of representation that
- 3:04
you think is useful um in the future.
- 3:07
What we kind of view as an embodied
- 3:08
foundation model is actually uh a
- 3:11
framing that allows you to both to do
- 3:13
the standard perception the embodied
- 3:15
reasoning and then the northstar target
- 3:17
of control all within one model. Um, so
- 3:20
being able to reason across different
- 3:22
sense of modalities on the input and
- 3:24
being able to do kind of most of what I
- 3:26
mentioned on the output all within one
- 3:28
single unified model.
- 3:31
And so very quickly I'm going to talk
- 3:33
about two challenges and these are kind
- 3:34
of fundamental research challenges that
- 3:36
we face and I'll talk about kind of how
- 3:38
our company has approached this um and
- 3:40
and and what other folks are doing in
- 3:42
here as well. So the very first thing is
- 3:45
think about purely if you're going to
- 3:46
try to model something like video like
- 3:48
one hour of video depending on the
- 3:50
representation you might have something
- 3:52
like 1 million visual tokens that are
- 3:53
coming in. The truth is is that there's
- 3:55
not actually any ground truth that you
- 3:58
can use uh effectively, right? So you
- 4:00
can do things like and people have done
- 4:02
this of course like pull out the
- 4:03
transcripts uh predict the transcripts
- 4:06
from the video or or maybe synthetically
- 4:08
label some frames, ask it some questions
- 4:10
and it turns out that this is kind of a
- 4:12
humongous waste, right? So if you think
- 4:14
about what's going into your model, you
- 4:15
have 1 million tokens going in and
- 4:17
you're essentially calculating the loss
- 4:18
on something like you know 2% of all the
- 4:21
tokens that are going in. And so this is
- 4:24
very fundamentally problematic and and
- 4:25
and the truth is is any way you try to
- 4:27
figure out how to fix this, you're
- 4:29
essentially injecting a wrong training
- 4:30
signal. Either the signal is too sparse,
- 4:33
it's too synthetic, or it's too
- 4:35
indiscriminate. And so approaches beyond
- 4:37
just synthetic enrichment have been,
- 4:39
well, let's predict every single pixel.
- 4:41
I mean, true, this is a very dense
- 4:43
signal, but it actually very poorly
- 4:45
allocates attention. there's, you know,
- 4:47
you're comparing a, you're treating a
- 4:48
background pixel with the same degree of
- 4:50
importance as you're treating a gripper
- 4:52
tip or the contact points or the
- 4:54
specific failures or the physics. And so
- 4:56
what we do kind of at Perceptron is, and
- 4:59
this is kind of some of our core IP is
- 5:01
we think about what does a natural
- 5:02
perceptive objective look like? So
- 5:04
specifically, how can I predict the
- 5:06
percepts that we think will matter in in
- 5:08
the future in a very very automatic way.
- 5:11
So, as an example, you might hardcode
- 5:13
something like, well, if you have a
- 5:14
robotic arm, well, the the the tip of
- 5:16
the grippers turns out to be a very
- 5:17
useful percept that you can predict into
- 5:19
the future. And folks have started doing
- 5:21
this. Um, like the Momo app folks from
- 5:23
AI2 have done this. There are other VAS
- 5:25
that have done this. But this is still a
- 5:27
hard-coded percept. So, the question is,
- 5:29
can you figure out an automatic way that
- 5:31
the model semantically is able to learn
- 5:33
this very unique objective? And the
- 5:36
truth is, we have figured out a way. uh
- 5:38
we're not going to share how we do it
- 5:39
here, but this is just kind of hinting
- 5:41
at how we approach the the problem of of
- 5:43
of sparity.
- 5:46
The the second core problem that we've
- 5:48
spent a lot of time focusing on is is
- 5:50
context bloat. So, if you have always on
- 5:52
cameras, if you have robots that don't
- 5:54
necessarily wait, they don't stop,
- 5:55
you're essentially having to reason over
- 5:57
a very very long um very long amount of
- 6:00
tokens. And and so text is relatively
- 6:02
dense and video is relatively sparse.
- 6:05
And so the question is, are there
- 6:07
architectural breakthroughs that are
- 6:09
there in order for you to be able to
- 6:10
deal with this problem problem natively
- 6:13
rather than just trying to on some
- 6:15
synthetic level figure out how to fix
- 6:17
this imbalance?
- 6:19
Um, and so there's a couple of things
- 6:21
that you can do. One kind of core first
- 6:24
principle is you need to start treating
- 6:26
different uh modalities as completely
- 6:28
different. So you can't treat text
- 6:30
tokens as the same as image tokens as
- 6:32
the same as audio or video tokens. So,
- 6:35
one thing that you can start thinking
- 6:36
about doing is focusing on spatial
- 6:38
compression or token compression. And
- 6:40
people do really dumb things and we
- 6:41
started off doing the dumb things and it
- 6:43
does work. You can start thinking about
- 6:44
averaging, you know, patch-wise
- 6:47
representations across a a video or an
- 6:49
image. Um, and you can start getting
- 6:51
some interesting compression rates, you
- 6:53
know, up up up to 10x.
- 6:55
But still, this is relatively of a hack
- 6:57
and there's there's not, you know,
- 6:59
wellused architectural methods to
- 7:01
actually solve this. Um, so what we've
- 7:04
done in in in the last couple of months,
- 7:05
we've released what we think is our
- 7:07
approach to dealing with um varying
- 7:10
degrees of sparity, which is just let
- 7:12
the model figure out what tokens it
- 7:13
should look at and what tokens it
- 7:15
shouldn't look like. And so we released
- 7:17
our data sparse mixture of experts
- 7:18
paper, which essentially allows you to
- 7:20
do this. Um, it allows the model, well,
- 7:22
there's a router in the model allows it
- 7:24
to actually predict what token I should
- 7:27
input, what token I should skip. Um, and
- 7:29
it allows us to do it kind of for free
- 7:31
through all the different layers.
- 7:34
And it turns out that if you just let
- 7:36
the model learn, if you're not actually
- 7:37
hard coding any significant um
- 7:40
architectural priors, um, the model
- 7:42
actually does learn. So if you end up
- 7:44
visualizing the data sparse compute uh
- 7:47
that our models use, you'll actually see
- 7:48
that our models innately learn to start
- 7:51
focusing on very um high density
- 7:53
information or task relevant
- 7:55
information. Um so in in in the upper
- 7:57
right you can kind of see that the model
- 7:58
decides to focus in on the graph which
- 8:00
actually you as a human would also kind
- 8:02
of zoom into because this is likely uh
- 8:05
what's interesting within the figure. Um
- 8:07
and it turns out that even task
- 8:09
dependent um uh task dependent
- 8:12
allocation ends up happening as well. So
- 8:14
if you look at the bottom left if you
- 8:15
just ask you know a very very general
- 8:18
question you're kind of going to see an
- 8:20
attention graph that is throughout the
- 8:21
whole image. So the model just doesn't
- 8:23
know what the proper way to allocate
- 8:24
compute is. At the same time, if you ask
- 8:26
it to do something like, you know,
- 8:27
segment out all the fruit, you can see
- 8:29
that it's going to allocate more tokens
- 8:31
to what it thinks are fruit tokens. And
- 8:33
so this is a very nice and clever trick
- 8:35
that that we use and we've published and
- 8:37
and other folks are starting to use
- 8:38
around embedding priors into the
- 8:41
architecture that are useful to deal
- 8:44
with the sparsity imbalances of your
- 8:46
modalities but not too harsh to the
- 8:48
point that the models aren't actually
- 8:50
learning uh natively.
- 8:54
And so kind of we we put all this
- 8:55
together and you guys might have seen
- 8:56
the release, but we essentially released
- 8:58
our our our model which was uh what we
- 9:01
considered to be the first embodied
- 9:03
foundation model a couple weeks ago. And
- 9:05
this model is essentially frontier um
- 9:08
with respect to Gemini 3.1 Pro. It's
- 9:10
actually better than Gemini embodied
- 9:12
reasoning. Um and it's something like 15
- 9:14
times cheaper. And it's essentially
- 9:15
trained on this one pedibyte data set
- 9:18
that we've collected across uh literally
- 9:20
everything. It's uh from from internet
- 9:23
crawls to to our own custommade training
- 9:27
recipes or synthetic data pipelines. We
- 9:29
have this one one pabyte of data across
- 9:31
text, images, videos, trajectories. And
- 9:33
these trajectories can be very very
- 9:35
general. They can be desktop use
- 9:37
trajectories. It can be playing a video
- 9:39
game trajectory.
- 9:41
And it turns out that once you start
- 9:43
doing these things, very interesting
- 9:45
properties end up emerging. And so the
- 9:46
the biggest property that we kind of saw
- 9:48
which is kind of obvious in retrospect
- 9:50
is that you can actually start thinking
- 9:52
of doing classical CV tasks as being an
- 9:56
agentic task. And so in this case we
- 9:58
essentially reframe detection as an
- 10:00
agentic task. So our model can you know
- 10:02
write code it can it can ask to zoom in
- 10:04
into specific portions. Uh it can change
- 10:07
the contrast. Um and you can actually
- 10:10
see here this is a very hard problem. I
- 10:12
think there's a whole Reddit subreddit
- 10:14
of these problems of trying to find very
- 10:16
hard objects and images. And our models
- 10:18
essentially do this very well, but they
- 10:19
do this in an agentic sense. So this
- 10:21
isn't the this isn't a classical, you
- 10:23
know, detect this one box. This is the
- 10:25
model actually deciding that it needs to
- 10:27
tile things up. It needs to change the
- 10:29
contrast. It proposes a box here. I
- 10:31
think here it yeah uh increases contrast
- 10:35
and it can find the bird. So again, very
- 10:37
hard to do if if like even for a human
- 10:40
and humans are very good at perceptive
- 10:41
tasks. This is a relatively tough thing
- 10:43
to do. And this all comes from just
- 10:45
having natively embodied models that
- 10:48
actually understand how to look at
- 10:49
different modalities.
- 10:51
And so following up kind of how does
- 10:54
this relate to the general physical AI
- 10:56
stance around robotics? So I kind of
- 10:58
stole this slide from from GDM folks.
- 11:00
And so one thing we've we're starting to
- 11:02
see from kind of robotics um u agentic
- 11:06
systems is this separation between what
- 11:09
we call kind of embodied reasoning
- 11:11
models or orchestrators and tactile
- 11:13
policy models. You can think of problems
- 11:15
as like you know if I have a if if I'm
- 11:19
making coffee and that takes me three
- 11:20
minutes to do I mean one thing I can do
- 11:21
is try to force my whole you know VLA to
- 11:24
try to figure out how to do this
- 11:25
individual task or what I can do is I
- 11:27
can have an orchestrator model that
- 11:28
breaks up this tasks into subtasks and
- 11:31
there's a tactile control policy that's
- 11:32
running on top and there's kind of a
- 11:34
full spectrum between you know full VA
- 11:36
only all the way to this kind of agentic
- 11:38
system. Uh but the main thing I'm trying
- 11:40
to highlight is that embodied reasoning
- 11:42
is actually very interesting and complex
- 11:44
problem that is yet to be solved. Um
- 11:46
that being said, our models continue to
- 11:48
be frontier on embodied reasoning. And
- 11:51
because they're frontier, we start
- 11:52
seeing really cool things that we
- 11:53
haven't seen before. Here's a concrete
- 11:56
example of doing very complex egocentric
- 11:59
or not egocentric, but this is robotic
- 12:01
data annotation. And you can kind of see
- 12:03
here the model is jumping around looking
- 12:05
at different portions of the video, uh,
- 12:07
you know, clipping it, figuring out
- 12:08
whether or not the captions are correct,
- 12:10
self-ver uh, self-verifying. And we can
- 12:13
all do this because AR models are fast.
- 12:15
They're significantly cheaper than
- 12:17
anything else that's out there. So, if
- 12:19
you try to do this with Gemini, this
- 12:20
video would probably cost you a couple
- 12:21
of dollars where for us, it's probably
- 12:23
in the sense. Um, and these all kind of
- 12:25
emerged from being able to have these
- 12:27
frontier embodied reasoning capabilities
- 12:29
that we just previously have not seen
- 12:31
from other models.
- 12:35
Um,
- 12:37
okay. Probably going to share the the
- 12:39
biggest research breakthrough that we've
- 12:40
had and I think we'll we'll share more
- 12:42
of this uh in the upcoming weeks
- 12:45
probably on Twitter. Uh, but one thing
- 12:46
that we found is we've we've we've
- 12:48
discovered new scaling laws for embodied
- 12:50
foundation models. So these are models
- 12:52
again that you can jointly do
- 12:54
control-based training, you can do
- 12:56
trajectory training, you can do
- 12:58
perceptive training, you can do embodied
- 13:00
reasoning training. If you just figure
- 13:02
out what the right way to mix this all
- 13:03
together is and the right objectives to
- 13:05
use, you actually start seeing very
- 13:07
interesting levers that you maybe
- 13:08
previously haven't been able to see
- 13:10
before. So the concrete lever that I'll
- 13:13
talk about is this ability to trade um
- 13:16
this ability to trade general video uh
- 13:20
pre-training data for teleop data. So
- 13:22
kind of as we know teleop data is very
- 13:24
expensive. It's on the orders of you
- 13:26
know $100 per hour of data for a hundred
- 13:29
you know dollars I can collect
- 13:30
significantly more video pre-training
- 13:32
data. And so what this graph is showing
- 13:34
is that um if you're just training pure
- 13:37
VA's pure policies there's this band
- 13:39
that you have. So you do actually still
- 13:41
have scaling loss. So you do get
- 13:43
benefits from more and more teop data.
- 13:45
That being said, the benefits are not as
- 13:47
substantial as if you are really
- 13:48
training these unified embodied
- 13:50
foundation models. And so this is what
- 13:52
kind of the bottom bottom half of the
- 13:54
graph is. Uh and the really cool kind of
- 13:58
uh lever that we get is you can
- 14:00
essentially trade 10x uh less teleop
- 14:03
data if you have 10x more video
- 14:05
pre-training data. And so far this is
- 14:07
kind of held for the amount of compute
- 14:09
that our company has. Um and it will be
- 14:11
continuing to to kind of push the fold
- 14:14
on uh how far you can push these
- 14:16
embodied foundation models.
- 14:22
So here's a couple of videos of of a
- 14:25
policy that hopefully will will um we'll
- 14:28
open source one of the smaller models in
- 14:29
a couple of weeks. Uh but this is all
- 14:31
running natively within a single model
- 14:33
that is capable of doing the embodied
- 14:36
reasoning in order to figure out the
- 14:37
task actually is outputting control
- 14:40
tokens. Um you can see it's a little bit
- 14:42
jittery but that's okay. Hopefully it'll
- 14:44
be figured out at scale and you can
- 14:46
actually see very complex tasks that
- 14:48
previously I think would be really tough
- 14:49
for pure VA to do. So I think if you
- 14:52
look at the right hand video, this is
- 14:54
requiring them all to actually read the
- 14:56
title of the book, have the knowledge
- 14:57
about what type of book this is and then
- 15:00
properly allocate it within one of the
- 15:02
bins. Um so this is actually a
- 15:03
multi-step task between perception and
- 15:06
control that is really really tough to
- 15:07
do if you have a kind of a pure
- 15:09
endto-end control model that is not
- 15:11
aware of the different perceptive tasks
- 15:14
that it needs to accomplish in order to
- 15:15
do this task.
- 15:23
Yeah,
- 15:25
it's pretty cool. It also works zero
- 15:27
shot relatively well out of the box. So,
- 15:29
we're excited to get this in the hands
- 15:31
of folks in in a couple of weeks
- 15:33
sometime in July.
- 15:38
Um, yeah, going to leave a couple of
- 15:40
minutes for for for general questions,
- 15:42
but if you guys are interested, let's uh
- 15:45
let's connect. So, one cool thing that
- 15:47
we do with our company is we actually
- 15:49
for a limited set of partners give
- 15:50
access to our Mark1 weights. We give
- 15:52
access to our larger embodied foundation
- 15:54
models weights. So, email me, DM me on
- 15:57
Twitter, uh, whatever is easier. Um, and
- 15:59
then yeah, I'll open up. There's a
- 16:01
couple minutes left for for questions.
- 16:17
these models
- 16:22
and like when I saw
- 16:31
a lot of contss
- 16:39
So can you talk a little bit on that?
- 16:41
>> Yeah, I think it's it's so yeah. So the
- 16:43
question is how are we able to nail
- 16:44
temporal and uh understanding to this
- 16:47
degree? Um it's a good question. I mean
- 16:49
to be honest it's not nailed. So there's
- 16:51
still a lot of work to actually get it
- 16:52
to a place where you can reliably deploy
- 16:54
it. The the core thing is how do you
- 16:56
think about context management? So you
- 16:58
have relatively limited context. So I
- 17:01
think the the the models here have 1
- 17:03
million context, but that's relatively
- 17:04
easy to fit in with the, you know, high
- 17:07
FPS video. And so you have to start
- 17:09
thinking about, are there interesting
- 17:10
things that you can do? I'll throw
- 17:12
something out there. We we used to do
- 17:14
this, but we got past this, but like how
- 17:15
do I think about like uh uh key frames
- 17:18
versus delta frames? How can I manage my
- 17:20
context by training these two off? Um,
- 17:23
and then you start thinking about during
- 17:25
your pre-training objective, how can I
- 17:26
start kind of natively ingesting things
- 17:30
that I think will be useful for the
- 17:31
robotics tasks. Um, so for example, a
- 17:34
very basic thing that even kind of the
- 17:35
the Gemini models used to struggle at, I
- 17:38
think the new ones are pretty good, but
- 17:39
being able to tell cardalities. So like
- 17:41
left and right is very hard to tell if
- 17:43
you do internet scale crawls because no
- 17:45
one on the internet is necessarily
- 17:46
labeling things as, you know, this
- 17:48
object is to the left of this object,
- 17:50
it's below this object. So really
- 17:52
thinking about data distributions early
- 17:53
on gives you this ability relatively
- 17:55
quickly. And it's also like er this type
- 17:58
of embodied reasoning was a very
- 17:59
concrete focus with this which is why I
- 18:02
think we were able to kind of surpass
- 18:03
Gemini ER with relatively less compute.
- 18:12
>> One issue with
- 18:16
sometimes I was wondering whether
- 18:20
you looked
- 18:22
your models are more robust than
- 18:28
>> Yeah. So, probably the coolest
- 18:29
robustness that we've seen is that uh
- 18:32
for VA models specifically, like if you
- 18:34
go and you take one of the Chinese ones
- 18:37
um and you try to fine-tune it for a
- 18:38
specific policy, if you just change the
- 18:40
background of the I don't know, you can
- 18:43
like in the table uh if you change the
- 18:45
the background of the table, the policy
- 18:47
will actually fail. What's really
- 18:49
interesting, if you do this type of
- 18:50
joint perceptive and control modeling,
- 18:52
you're much more robust to these types
- 18:53
of errors or if the light is hitting it
- 18:56
a slightly different way. And I think we
- 18:58
primarily view this as robustness to
- 18:59
background uh in a way that I think
- 19:02
traditional models don't necessarily
- 19:03
have. That being said, I'm not going to
- 19:05
overclaim like uh I mean like it's still
- 19:07
relatively hard. I think if I was going
- 19:08
to go and shine a flashlight into one of
- 19:10
the one of the arms, it's probably not
- 19:12
going to work. Uh but being able to
- 19:14
jointly model these things helps a
- 19:17
significant amount. We also do a lot of
- 19:18
online augmentations. So, so we do
- 19:20
actually, you know, fake um I don't know
- 19:23
um one of the one of the arm cameras
- 19:26
being off, right? We fake uh you know,
- 19:29
sunlight coming in from a certain
- 19:31
direction, right? So, we do these things
- 19:33
during training to improve robustness.
- 19:35
Uh but the big gains come from taking
- 19:37
this early fusion paradigm and then
- 19:39
moving it into the robotics domain.
- 19:44
>> Cool.
- 19:47
uh create knowledge base using the K1
- 19:50
model.
- 19:51
>> Uh oh, knowledge bases. Um it's useful
- 19:53
for like I I mean if you want to caption
- 19:55
images, videos, it's relatively well. We
- 19:58
work with robotics partners for like
- 19:59
very complex egocentric annotation like
- 20:01
this. Um so it's not necessarily
- 20:03
building out an ontology, but being able
- 20:05
to do kind of very deep structured
- 20:07
extraction, I think our models are are
- 20:09
very good at. Yeah. By the way,
- 20:11
everything that I kind of showed here is
- 20:13
kind of public APIs, so you can go play
- 20:15
around with it. The benchmarks are
- 20:16
public. Um, I think I'm out of time.
- 20:19
They're cutting me off. So, uh, I can
- 20:22
talk with folks outside, but thank you
- 20:23
guys.