AI Engineer World's Fair 2026

From VLM/VLA's to Embodied Agents — Armen Aghajanyan, Perceptron AI

Read the talk

From VLM/VLAs to Embodied Agents

Armen Aghajanyan explains Perceptron AI’s approach to combining perception, reasoning, and control: learn useful visual targets, spend compute on relevant tokens, and use video pretraining to reduce the need for expensive robot demonstrations.

From a talk by Armen Aghajanyan

At a glance

Ideas worth remembering

  • Visual supervision needs useful targets as well as density. Predicting every pixel can spend learning effort on background details; Perceptron proposes automatically learning future percepts, without disclosing the objective here.

  • Learned token routing makes compute allocation task-dependent: a general question spreads attention across an image, while fruit segmentation concentrates more tokens on likely fruit.

  • Perception can itself involve actions. Tiling, zooming, changing contrast, and revisiting video intervals let a model gather better evidence before producing a box or annotation.

  • Joint training reportedly lets 10× more video pretraining substitute for 10× less teleoperation data within Perceptron’s tested compute range, offering a way to reduce dependence on demonstrations costing around $100 per hour.

  • Combining reasoning and control enables tasks such as sorting books by their titles, but temporal reliability and resistance to severe visual disruption remain open problems.

Perception, reasoning, and action in one model

A camera keeps producing observations; a robot has to turn those observations into decisions and movements. Perceptron AI’s co-founder and CEO, Armen Aghajanyan, starts with that combined problem. The goal is physical intelligence that can perceive, understand, and interact with the world in real time, whether attached to a robot, an instrument, or a sensor.

Source frame: Perception, reasoning, and action in one model
Source frame: Perception, reasoning, and action in one model

The architectural starting point is early fusion: bring modalities together as early as possible. Aghajanyan connects this to his work scaling multimodal recipes at FAIR. The difficult part is choosing representations that work together on both sides of the model—what it receives and what it produces. Joining inputs early does not make text, images, video, and actions interchangeable; their representations still have to preserve what matters about each.

The familiar model categories describe different input and output capabilities:

  • Vision-language models (VLMs): receive images or video alongside text and produce text. Embodied reasoning variants can also produce grounding points and handle spatial questions.
  • Vision-language-action models (VLAs): extend the output space to actions, commonly using a VLM backbone.
  • World models: produce video conditioned on inputs that may include images, video, and actions.
  • Semantic world models: learn representations intended to be useful later, rather than directly producing a visible output.

An embodied foundation model is Perceptron’s framing for putting standard perception, embodied reasoning, and control within one model. It should reason across sensory inputs and support the relevant outputs together. That ambition immediately raises two training problems: how to learn enough from visual data, and how to process a stream that never stops.

0:150:18
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:13 · section reference included

A million visual tokens, very little supervision

Depending on the representation, one hour of video can bring roughly 1 million visual tokens into a model. The training target may be only a transcript or answers attached to a few synthetically labeled frames. In Aghajanyan’s example, the loss is calculated on something like 2% of the incoming token count. The model processes a large visual sequence while receiving a much smaller amount of explicit target information.

Source frame: A million visual tokens, very little supervision
Source frame: A million visual tokens, very little supervision

Predicting every pixel supplies a dense target, but density alone does not allocate learning effort well. A background pixel receives the same importance as a gripper tip, a contact point, or evidence of a physical failure. For manipulation, those details have different consequences: the gripper and contact geometry can determine whether an action succeeds, while much of the background contributes little to that decision.

Perceptron’s proposed perceptive objective predicts percepts expected to matter in the future. A hand-chosen version might predict a robotic gripper’s future tip position: a useful target, but one whose importance the designer has already specified. The research question is whether the model can learn which semantic percepts deserve prediction automatically. Aghajanyan says Perceptron has a method, but does not disclose its construction here; the talk explains the target-selection problem rather than a reproducible objective.

3:453:46
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

3:45 · section reference included

Let the task decide where compute goes

Always-on cameras create the second problem: context bloat. Robots do not pause their observations while a model catches up. Video supplies many tokens, but useful information can be sparse relative to text. The first practical response is compression. Perceptron began by averaging patch representations across images or video, which Aghajanyan says can achieve up to 10× compression. It works, but he regards that fixed compression strategy as a hack rather than a sufficient architectural answer.

Source frame: Let the task decide where compute goes
Source frame: Let the task decide where compute goes

Data sparse mixture of experts gives the model a learned choice about which tokens to process. A router predicts which tokens to admit and which to skip across the model’s layers. This addresses a different question from the perceptive objective: the objective concerns what the model should learn to predict; routing concerns where the model spends computation while processing its input.

The described compute visualizations show two useful behaviors:

  • Information-driven allocation: within a figure, the model concentrates on the graph, a region containing dense information.
  • Task-driven allocation: a general question spreads computation across an image; asking to segment all the fruit shifts more tokens toward regions the model considers fruit.

What changes when the question becomes specific? The diagram follows the fruit example: the visual input stays the same, while the requested task changes the allocation. A general question gives the model little reason to favor a particular region. Segmentation supplies that reason, and computation concentrates on likely fruit. The architectural choice enables selective processing without hardcoding fruit locations or a fixed region of interest.

Compare the ideasSame image, different compute allocation

No specific object class is requested.

The router selects or skips tokens across layers. Task-specific prompting changes which image regions receive more computation.

5:465:50
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

5:46 · section reference included

Finding a bird by changing how the image is examined

Perceptron combines these ideas in a model trained on a reported 1 petabyte of text, images, videos, and trajectories. The trajectories include desktop use and video games; the collection combines internet crawls, custom training recipes, and synthetic pipelines. Aghajanyan reports performance exceeding Gemini’s embodied reasoning model at roughly 15× lower cost. These are his reported comparisons; the talk does not specify the benchmark conditions or cost basis needed to generalize them.

Source frame: Finding a bird by changing how the image is examined
Source frame: Finding a bird by changing how the image is examined

One resulting capability turns object detection into an agentic task. The model can write code, request a closer view of part of an image, and change contrast. A difficult image therefore becomes something it can investigate through successive operations, rather than something it must resolve with one bounding-box prediction.

The bird-finding example makes that investigation concrete. The model decides to tile the image, adjusts contrast, and proposes a candidate box. Aghajanyan then describes a contrast increase that lets it find the bird. The observable change is in the evidence available for detection: subdividing the image makes regions easier to examine, and changing contrast makes the target easier to distinguish. The final box follows those inspection steps.

The important connection is between perception and action even before a robot moves. Here, actions alter the model’s view of the input. Detection becomes a small investigation in which the model chooses how to look, then uses the resulting view to locate the object.

8:548:56
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

8:54 · section reference included

Long tasks need decomposition—and ways to check the evidence

Physical tasks add another design choice: how much work belongs to a single action policy, and how much belongs to an orchestrator. Making coffee might take three minutes. One approach asks a VLA to handle the entire task. Another uses an embodied reasoning model to break it into subtasks, with a tactile control policy handling execution. Aghajanyan presents a spectrum between these arrangements and treats embodied reasoning itself as an unsolved problem.

Source frame: Long tasks need decomposition—and ways to check the evidence
Source frame: Long tasks need decomposition—and ways to check the evidence

Robotic video annotation shows a related use of reasoning without physical execution. The model moves among different portions of a video, clips relevant intervals, checks whether captions match them, and verifies its own annotations. This extends the bird example’s inspection behavior into time: the model chooses which evidence to revisit instead of accepting an initial description unchecked. Aghajanyan connects the practicality of repeated inspection to the model’s speed and lower cost.

10:5110:54
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

10:51 · section reference included

Trading video pretraining for expensive teleoperation

The largest research result Aghajanyan reports concerns how training sources substitute for one another. The unified model mixes control-based training, trajectories, perceptive training, and embodied reasoning. With suitable objectives and mixing, Perceptron finds a lever that a pure policy-training setup does not expose as strongly: general video pretraining can reduce the amount of teleoperation data needed.

Source frame: Trading video pretraining for expensive teleoperation
Source frame: Trading video pretraining for expensive teleoperation

Teleoperation data costs on the order of $100 per hour in his account. The same budget can collect substantially more video pretraining data. Pure policies still benefit from additional teleoperation, but joint training opens a different spending decision: acquire more general visual experience instead of buying every increment of capability through robot demonstrations.

The reported trade is 10× more video pretraining data for 10× less teleoperation data. That relationship has held within the compute range Perceptron tested; the talk does not establish the task metric, experimental setup, or behavior beyond that range. Its practical significance is the possibility of reducing dependence on costly demonstrations while retaining control training, rather than eliminating demonstrations altogether.

12:3712:39
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

12:35 · section reference included

Reading a book title becomes a control decision

The control demonstrations return to the original ambition: one model performs embodied reasoning and emits control tokens. The book-sorting task requires more than moving an object. It reads a book’s title, uses knowledge of what kind of book it is, and assigns it to an appropriate bin. Perception supplies the title; reasoning determines the category; control carries out the placement.

Source frame: Reading a book title becomes a control decision
Source frame: Reading a book title becomes a control decision

Where does understanding enter the movement? The diagram separates the dependencies within the task, while the demonstration uses a single model. Choosing a bin depends on understanding the book, so a control model unaware of the necessary perceptual work has a harder problem than a policy that only needs to move an already classified object.

The motion is visibly jittery in Aghajanyan’s description; improving it through scale remains a hope. He also reports relatively good zero-shot behavior and anticipates opening a smaller model in July. That is an announcement made during the talk, not confirmation of a subsequent release. Larger model weights are offered to a limited set of partners.

How it fits togetherThe dependencies inside book sorting

Perception extracts information from the book.

These are explanatory stages of the task, not separate deployed models. A title must inform a category before control places the book in a bin.

14:2214:31
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

14:22 · section reference included

Temporal understanding still needs context management

The first Q&A answer narrows the deployment claim. Temporal understanding is not yet solved to a reliably deployable level. Even a context of roughly 1 million tokens can fill readily with high-frame-rate video. Aghajanyan describes keyframes versus delta frames as an earlier context-management approach Perceptron used and moved beyond: the decision is how much complete visual state to retain versus how much change to represent.

Pretraining also needs data that teaches relationships useful to robots. Left, right, and below are simple examples, yet internet crawls do not necessarily label them explicitly. Aghajanyan connects early attention to the data distribution with learning these capabilities quickly. Large quantities of general data do not automatically supply every spatial distinction a robot needs; the training mix must make those distinctions learnable.

16:4316:44
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

16:17 · section reference included

Background changes, missing cameras, and structured extraction

Background robustness provides a concrete test of what the policy has learned. Aghajanyan describes fine-tuned action policies that fail when the table’s background changes. Joint perceptive and control modeling has made Perceptron’s models more tolerant of background changes and modest lighting differences in his observations. That tolerance has limits: he expects a flashlight directed into an arm camera could still break the behavior.

Training adds explicit variation alongside joint modeling:

  • Camera loss: online augmentations simulate an arm camera being off.
  • Lighting changes: augmentations simulate sunlight arriving from a particular direction.

These examples expose the model to altered observations during training. Aghajanyan credits the larger gains to bringing early fusion into robotics, with augmentation contributing additional robustness.

The final question asks about knowledge bases. The useful capability identified here is captioning and deep structured extraction from images and video, including complex egocentric annotation for robotics partners. Building an ontology is a separate undertaking. The model’s role is to turn difficult visual evidence into detailed structured information—the same kind of inspection and annotation work demonstrated earlier. Aghajanyan closes by saying the demonstrated capabilities are available through public APIs and that benchmarks are public.

18:2818:29
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

18:12 · section reference included

Resources

From the talk

  • A contact route for the limited partner access to larger model weights offered near the end of the talk.

Read the complete timestamped transcript
  1. 0:13

    Yeah, I guess uh let's get started. I'm

  2. 0:15

    I'm Armen. I'm I'm the co-founder and

  3. 0:18

    CEO of Perceptron. Um, we'll get a

  4. 0:20

    little bit into into what we do, but uh,

  5. 0:23

    primarily what I want to talk about

  6. 0:24

    today is kind of our research stance

  7. 0:27

    that we want to move away from

  8. 0:29

    distinctions between VLMs, VAS, world

  9. 0:32

    models, whatever you want to call it, to

  10. 0:34

    something that we call embodied

  11. 0:35

    foundation models.

  12. 0:42

    So specifically what we what we do at

  13. 0:45

    Perceptron kind of our northstar really

  14. 0:46

    is we want to be able to build physical

  15. 0:49

    AI foundations that give us the ability

  16. 0:51

    to perceive, understand and interact

  17. 0:53

    with the physical world in real time.

  18. 0:55

    And so kind of the the northstar mission

  19. 0:57

    is really to bridge the the physical and

  20. 0:59

    digital world uh worlds fundamentally

  21. 1:02

    meaning that our our goal is kind of

  22. 1:04

    wherever there is a um a device, an

  23. 1:07

    instrument, a robot, a camera, a sensor,

  24. 1:10

    uh we're essentially there providing

  25. 1:11

    intelligence to it.

  26. 1:14

    And specifically when we talk about this

  27. 1:16

    paradigm of being able to perceive,

  28. 1:18

    being able to reason, being able to act,

  29. 1:20

    we view this as kind of a a unification

  30. 1:23

    of uh traditional multimodal modeling.

  31. 1:25

    So I came from fair, I was there for six

  32. 1:28

    years and and and my target there was

  33. 1:30

    really to try to figure out how to scale

  34. 1:32

    up recipes for multimodal models. And so

  35. 1:35

    one of the early things that we we

  36. 1:36

    started working on and we've published a

  37. 1:38

    lot in this domain was around early

  38. 1:39

    fusion. So this concept that you want to

  39. 1:41

    bring in all these modalities as early

  40. 1:43

    as you can and and the real complexity

  41. 1:46

    there is trying to figure out what is

  42. 1:47

    the correct way to actually properly

  43. 1:48

    represent all the different modalities

  44. 1:51

    uh both on the input and on on the

  45. 1:53

    output that you want to be able to

  46. 1:55

    represent holistically and so VLMs have

  47. 1:57

    kind of became the when I talk about

  48. 1:59

    multimodal models um you probably think

  49. 2:01

    of VLM. So this is the ability to take

  50. 2:03

    in some image video and some text and be

  51. 2:05

    able to output some some text

  52. 2:07

    essentially. And there's there's

  53. 2:09

    variations of this. There's there's

  54. 2:10

    models like the ER models, the embodied

  55. 2:12

    reasoning models that are able to output

  56. 2:14

    maybe some grounding points um that are

  57. 2:17

    able to do spatial understanding or

  58. 2:18

    reasoning a little bit better. Then we

  59. 2:20

    have things like VALAS that extend kind

  60. 2:22

    of the output domain apart from just

  61. 2:24

    text to now actions. And these are

  62. 2:26

    traditionally of course continue to be

  63. 2:28

    built on by standard VLM backbones.

  64. 2:30

    Although there's been some efforts to

  65. 2:32

    try to migrate away from uh from VLMs to

  66. 2:36

    things like world models or world action

  67. 2:38

    models, although nothing that's been

  68. 2:39

    super fruitful just yet. And then we get

  69. 2:42

    into kind of more interesting and

  70. 2:44

    complex uh variants of multimodal models

  71. 2:46

    like world models where you essentially

  72. 2:48

    are outputting video uh from some inputs

  73. 2:51

    and the inputs can be either image or

  74. 2:52

    video or honestly image video and and

  75. 2:55

    actions. Um and the last point that I'll

  76. 2:57

    talk about is something recent which we

  77. 2:59

    call semantic world models which is you

  78. 3:01

    don't really output anything but you

  79. 3:03

    learn some type of representation that

  80. 3:04

    you think is useful um in the future.

  81. 3:07

    What we kind of view as an embodied

  82. 3:08

    foundation model is actually uh a

  83. 3:11

    framing that allows you to both to do

  84. 3:13

    the standard perception the embodied

  85. 3:15

    reasoning and then the northstar target

  86. 3:17

    of control all within one model. Um, so

  87. 3:20

    being able to reason across different

  88. 3:22

    sense of modalities on the input and

  89. 3:24

    being able to do kind of most of what I

  90. 3:26

    mentioned on the output all within one

  91. 3:28

    single unified model.

  92. 3:31

    And so very quickly I'm going to talk

  93. 3:33

    about two challenges and these are kind

  94. 3:34

    of fundamental research challenges that

  95. 3:36

    we face and I'll talk about kind of how

  96. 3:38

    our company has approached this um and

  97. 3:40

    and and what other folks are doing in

  98. 3:42

    here as well. So the very first thing is

  99. 3:45

    think about purely if you're going to

  100. 3:46

    try to model something like video like

  101. 3:48

    one hour of video depending on the

  102. 3:50

    representation you might have something

  103. 3:52

    like 1 million visual tokens that are

  104. 3:53

    coming in. The truth is is that there's

  105. 3:55

    not actually any ground truth that you

  106. 3:58

    can use uh effectively, right? So you

  107. 4:00

    can do things like and people have done

  108. 4:02

    this of course like pull out the

  109. 4:03

    transcripts uh predict the transcripts

  110. 4:06

    from the video or or maybe synthetically

  111. 4:08

    label some frames, ask it some questions

  112. 4:10

    and it turns out that this is kind of a

  113. 4:12

    humongous waste, right? So if you think

  114. 4:14

    about what's going into your model, you

  115. 4:15

    have 1 million tokens going in and

  116. 4:17

    you're essentially calculating the loss

  117. 4:18

    on something like you know 2% of all the

  118. 4:21

    tokens that are going in. And so this is

  119. 4:24

    very fundamentally problematic and and

  120. 4:25

    and the truth is is any way you try to

  121. 4:27

    figure out how to fix this, you're

  122. 4:29

    essentially injecting a wrong training

  123. 4:30

    signal. Either the signal is too sparse,

  124. 4:33

    it's too synthetic, or it's too

  125. 4:35

    indiscriminate. And so approaches beyond

  126. 4:37

    just synthetic enrichment have been,

  127. 4:39

    well, let's predict every single pixel.

  128. 4:41

    I mean, true, this is a very dense

  129. 4:43

    signal, but it actually very poorly

  130. 4:45

    allocates attention. there's, you know,

  131. 4:47

    you're comparing a, you're treating a

  132. 4:48

    background pixel with the same degree of

  133. 4:50

    importance as you're treating a gripper

  134. 4:52

    tip or the contact points or the

  135. 4:54

    specific failures or the physics. And so

  136. 4:56

    what we do kind of at Perceptron is, and

  137. 4:59

    this is kind of some of our core IP is

  138. 5:01

    we think about what does a natural

  139. 5:02

    perceptive objective look like? So

  140. 5:04

    specifically, how can I predict the

  141. 5:06

    percepts that we think will matter in in

  142. 5:08

    the future in a very very automatic way.

  143. 5:11

    So, as an example, you might hardcode

  144. 5:13

    something like, well, if you have a

  145. 5:14

    robotic arm, well, the the the tip of

  146. 5:16

    the grippers turns out to be a very

  147. 5:17

    useful percept that you can predict into

  148. 5:19

    the future. And folks have started doing

  149. 5:21

    this. Um, like the Momo app folks from

  150. 5:23

    AI2 have done this. There are other VAS

  151. 5:25

    that have done this. But this is still a

  152. 5:27

    hard-coded percept. So, the question is,

  153. 5:29

    can you figure out an automatic way that

  154. 5:31

    the model semantically is able to learn

  155. 5:33

    this very unique objective? And the

  156. 5:36

    truth is, we have figured out a way. uh

  157. 5:38

    we're not going to share how we do it

  158. 5:39

    here, but this is just kind of hinting

  159. 5:41

    at how we approach the the problem of of

  160. 5:43

    of sparity.

  161. 5:46

    The the second core problem that we've

  162. 5:48

    spent a lot of time focusing on is is

  163. 5:50

    context bloat. So, if you have always on

  164. 5:52

    cameras, if you have robots that don't

  165. 5:54

    necessarily wait, they don't stop,

  166. 5:55

    you're essentially having to reason over

  167. 5:57

    a very very long um very long amount of

  168. 6:00

    tokens. And and so text is relatively

  169. 6:02

    dense and video is relatively sparse.

  170. 6:05

    And so the question is, are there

  171. 6:07

    architectural breakthroughs that are

  172. 6:09

    there in order for you to be able to

  173. 6:10

    deal with this problem problem natively

  174. 6:13

    rather than just trying to on some

  175. 6:15

    synthetic level figure out how to fix

  176. 6:17

    this imbalance?

  177. 6:19

    Um, and so there's a couple of things

  178. 6:21

    that you can do. One kind of core first

  179. 6:24

    principle is you need to start treating

  180. 6:26

    different uh modalities as completely

  181. 6:28

    different. So you can't treat text

  182. 6:30

    tokens as the same as image tokens as

  183. 6:32

    the same as audio or video tokens. So,

  184. 6:35

    one thing that you can start thinking

  185. 6:36

    about doing is focusing on spatial

  186. 6:38

    compression or token compression. And

  187. 6:40

    people do really dumb things and we

  188. 6:41

    started off doing the dumb things and it

  189. 6:43

    does work. You can start thinking about

  190. 6:44

    averaging, you know, patch-wise

  191. 6:47

    representations across a a video or an

  192. 6:49

    image. Um, and you can start getting

  193. 6:51

    some interesting compression rates, you

  194. 6:53

    know, up up up to 10x.

  195. 6:55

    But still, this is relatively of a hack

  196. 6:57

    and there's there's not, you know,

  197. 6:59

    wellused architectural methods to

  198. 7:01

    actually solve this. Um, so what we've

  199. 7:04

    done in in in the last couple of months,

  200. 7:05

    we've released what we think is our

  201. 7:07

    approach to dealing with um varying

  202. 7:10

    degrees of sparity, which is just let

  203. 7:12

    the model figure out what tokens it

  204. 7:13

    should look at and what tokens it

  205. 7:15

    shouldn't look like. And so we released

  206. 7:17

    our data sparse mixture of experts

  207. 7:18

    paper, which essentially allows you to

  208. 7:20

    do this. Um, it allows the model, well,

  209. 7:22

    there's a router in the model allows it

  210. 7:24

    to actually predict what token I should

  211. 7:27

    input, what token I should skip. Um, and

  212. 7:29

    it allows us to do it kind of for free

  213. 7:31

    through all the different layers.

  214. 7:34

    And it turns out that if you just let

  215. 7:36

    the model learn, if you're not actually

  216. 7:37

    hard coding any significant um

  217. 7:40

    architectural priors, um, the model

  218. 7:42

    actually does learn. So if you end up

  219. 7:44

    visualizing the data sparse compute uh

  220. 7:47

    that our models use, you'll actually see

  221. 7:48

    that our models innately learn to start

  222. 7:51

    focusing on very um high density

  223. 7:53

    information or task relevant

  224. 7:55

    information. Um so in in in the upper

  225. 7:57

    right you can kind of see that the model

  226. 7:58

    decides to focus in on the graph which

  227. 8:00

    actually you as a human would also kind

  228. 8:02

    of zoom into because this is likely uh

  229. 8:05

    what's interesting within the figure. Um

  230. 8:07

    and it turns out that even task

  231. 8:09

    dependent um uh task dependent

  232. 8:12

    allocation ends up happening as well. So

  233. 8:14

    if you look at the bottom left if you

  234. 8:15

    just ask you know a very very general

  235. 8:18

    question you're kind of going to see an

  236. 8:20

    attention graph that is throughout the

  237. 8:21

    whole image. So the model just doesn't

  238. 8:23

    know what the proper way to allocate

  239. 8:24

    compute is. At the same time, if you ask

  240. 8:26

    it to do something like, you know,

  241. 8:27

    segment out all the fruit, you can see

  242. 8:29

    that it's going to allocate more tokens

  243. 8:31

    to what it thinks are fruit tokens. And

  244. 8:33

    so this is a very nice and clever trick

  245. 8:35

    that that we use and we've published and

  246. 8:37

    and other folks are starting to use

  247. 8:38

    around embedding priors into the

  248. 8:41

    architecture that are useful to deal

  249. 8:44

    with the sparsity imbalances of your

  250. 8:46

    modalities but not too harsh to the

  251. 8:48

    point that the models aren't actually

  252. 8:50

    learning uh natively.

  253. 8:54

    And so kind of we we put all this

  254. 8:55

    together and you guys might have seen

  255. 8:56

    the release, but we essentially released

  256. 8:58

    our our our model which was uh what we

  257. 9:01

    considered to be the first embodied

  258. 9:03

    foundation model a couple weeks ago. And

  259. 9:05

    this model is essentially frontier um

  260. 9:08

    with respect to Gemini 3.1 Pro. It's

  261. 9:10

    actually better than Gemini embodied

  262. 9:12

    reasoning. Um and it's something like 15

  263. 9:14

    times cheaper. And it's essentially

  264. 9:15

    trained on this one pedibyte data set

  265. 9:18

    that we've collected across uh literally

  266. 9:20

    everything. It's uh from from internet

  267. 9:23

    crawls to to our own custommade training

  268. 9:27

    recipes or synthetic data pipelines. We

  269. 9:29

    have this one one pabyte of data across

  270. 9:31

    text, images, videos, trajectories. And

  271. 9:33

    these trajectories can be very very

  272. 9:35

    general. They can be desktop use

  273. 9:37

    trajectories. It can be playing a video

  274. 9:39

    game trajectory.

  275. 9:41

    And it turns out that once you start

  276. 9:43

    doing these things, very interesting

  277. 9:45

    properties end up emerging. And so the

  278. 9:46

    the biggest property that we kind of saw

  279. 9:48

    which is kind of obvious in retrospect

  280. 9:50

    is that you can actually start thinking

  281. 9:52

    of doing classical CV tasks as being an

  282. 9:56

    agentic task. And so in this case we

  283. 9:58

    essentially reframe detection as an

  284. 10:00

    agentic task. So our model can you know

  285. 10:02

    write code it can it can ask to zoom in

  286. 10:04

    into specific portions. Uh it can change

  287. 10:07

    the contrast. Um and you can actually

  288. 10:10

    see here this is a very hard problem. I

  289. 10:12

    think there's a whole Reddit subreddit

  290. 10:14

    of these problems of trying to find very

  291. 10:16

    hard objects and images. And our models

  292. 10:18

    essentially do this very well, but they

  293. 10:19

    do this in an agentic sense. So this

  294. 10:21

    isn't the this isn't a classical, you

  295. 10:23

    know, detect this one box. This is the

  296. 10:25

    model actually deciding that it needs to

  297. 10:27

    tile things up. It needs to change the

  298. 10:29

    contrast. It proposes a box here. I

  299. 10:31

    think here it yeah uh increases contrast

  300. 10:35

    and it can find the bird. So again, very

  301. 10:37

    hard to do if if like even for a human

  302. 10:40

    and humans are very good at perceptive

  303. 10:41

    tasks. This is a relatively tough thing

  304. 10:43

    to do. And this all comes from just

  305. 10:45

    having natively embodied models that

  306. 10:48

    actually understand how to look at

  307. 10:49

    different modalities.

  308. 10:51

    And so following up kind of how does

  309. 10:54

    this relate to the general physical AI

  310. 10:56

    stance around robotics? So I kind of

  311. 10:58

    stole this slide from from GDM folks.

  312. 11:00

    And so one thing we've we're starting to

  313. 11:02

    see from kind of robotics um u agentic

  314. 11:06

    systems is this separation between what

  315. 11:09

    we call kind of embodied reasoning

  316. 11:11

    models or orchestrators and tactile

  317. 11:13

    policy models. You can think of problems

  318. 11:15

    as like you know if I have a if if I'm

  319. 11:19

    making coffee and that takes me three

  320. 11:20

    minutes to do I mean one thing I can do

  321. 11:21

    is try to force my whole you know VLA to

  322. 11:24

    try to figure out how to do this

  323. 11:25

    individual task or what I can do is I

  324. 11:27

    can have an orchestrator model that

  325. 11:28

    breaks up this tasks into subtasks and

  326. 11:31

    there's a tactile control policy that's

  327. 11:32

    running on top and there's kind of a

  328. 11:34

    full spectrum between you know full VA

  329. 11:36

    only all the way to this kind of agentic

  330. 11:38

    system. Uh but the main thing I'm trying

  331. 11:40

    to highlight is that embodied reasoning

  332. 11:42

    is actually very interesting and complex

  333. 11:44

    problem that is yet to be solved. Um

  334. 11:46

    that being said, our models continue to

  335. 11:48

    be frontier on embodied reasoning. And

  336. 11:51

    because they're frontier, we start

  337. 11:52

    seeing really cool things that we

  338. 11:53

    haven't seen before. Here's a concrete

  339. 11:56

    example of doing very complex egocentric

  340. 11:59

    or not egocentric, but this is robotic

  341. 12:01

    data annotation. And you can kind of see

  342. 12:03

    here the model is jumping around looking

  343. 12:05

    at different portions of the video, uh,

  344. 12:07

    you know, clipping it, figuring out

  345. 12:08

    whether or not the captions are correct,

  346. 12:10

    self-ver uh, self-verifying. And we can

  347. 12:13

    all do this because AR models are fast.

  348. 12:15

    They're significantly cheaper than

  349. 12:17

    anything else that's out there. So, if

  350. 12:19

    you try to do this with Gemini, this

  351. 12:20

    video would probably cost you a couple

  352. 12:21

    of dollars where for us, it's probably

  353. 12:23

    in the sense. Um, and these all kind of

  354. 12:25

    emerged from being able to have these

  355. 12:27

    frontier embodied reasoning capabilities

  356. 12:29

    that we just previously have not seen

  357. 12:31

    from other models.

  358. 12:35

    Um,

  359. 12:37

    okay. Probably going to share the the

  360. 12:39

    biggest research breakthrough that we've

  361. 12:40

    had and I think we'll we'll share more

  362. 12:42

    of this uh in the upcoming weeks

  363. 12:45

    probably on Twitter. Uh, but one thing

  364. 12:46

    that we found is we've we've we've

  365. 12:48

    discovered new scaling laws for embodied

  366. 12:50

    foundation models. So these are models

  367. 12:52

    again that you can jointly do

  368. 12:54

    control-based training, you can do

  369. 12:56

    trajectory training, you can do

  370. 12:58

    perceptive training, you can do embodied

  371. 13:00

    reasoning training. If you just figure

  372. 13:02

    out what the right way to mix this all

  373. 13:03

    together is and the right objectives to

  374. 13:05

    use, you actually start seeing very

  375. 13:07

    interesting levers that you maybe

  376. 13:08

    previously haven't been able to see

  377. 13:10

    before. So the concrete lever that I'll

  378. 13:13

    talk about is this ability to trade um

  379. 13:16

    this ability to trade general video uh

  380. 13:20

    pre-training data for teleop data. So

  381. 13:22

    kind of as we know teleop data is very

  382. 13:24

    expensive. It's on the orders of you

  383. 13:26

    know $100 per hour of data for a hundred

  384. 13:29

    you know dollars I can collect

  385. 13:30

    significantly more video pre-training

  386. 13:32

    data. And so what this graph is showing

  387. 13:34

    is that um if you're just training pure

  388. 13:37

    VA's pure policies there's this band

  389. 13:39

    that you have. So you do actually still

  390. 13:41

    have scaling loss. So you do get

  391. 13:43

    benefits from more and more teop data.

  392. 13:45

    That being said, the benefits are not as

  393. 13:47

    substantial as if you are really

  394. 13:48

    training these unified embodied

  395. 13:50

    foundation models. And so this is what

  396. 13:52

    kind of the bottom bottom half of the

  397. 13:54

    graph is. Uh and the really cool kind of

  398. 13:58

    uh lever that we get is you can

  399. 14:00

    essentially trade 10x uh less teleop

  400. 14:03

    data if you have 10x more video

  401. 14:05

    pre-training data. And so far this is

  402. 14:07

    kind of held for the amount of compute

  403. 14:09

    that our company has. Um and it will be

  404. 14:11

    continuing to to kind of push the fold

  405. 14:14

    on uh how far you can push these

  406. 14:16

    embodied foundation models.

  407. 14:22

    So here's a couple of videos of of a

  408. 14:25

    policy that hopefully will will um we'll

  409. 14:28

    open source one of the smaller models in

  410. 14:29

    a couple of weeks. Uh but this is all

  411. 14:31

    running natively within a single model

  412. 14:33

    that is capable of doing the embodied

  413. 14:36

    reasoning in order to figure out the

  414. 14:37

    task actually is outputting control

  415. 14:40

    tokens. Um you can see it's a little bit

  416. 14:42

    jittery but that's okay. Hopefully it'll

  417. 14:44

    be figured out at scale and you can

  418. 14:46

    actually see very complex tasks that

  419. 14:48

    previously I think would be really tough

  420. 14:49

    for pure VA to do. So I think if you

  421. 14:52

    look at the right hand video, this is

  422. 14:54

    requiring them all to actually read the

  423. 14:56

    title of the book, have the knowledge

  424. 14:57

    about what type of book this is and then

  425. 15:00

    properly allocate it within one of the

  426. 15:02

    bins. Um so this is actually a

  427. 15:03

    multi-step task between perception and

  428. 15:06

    control that is really really tough to

  429. 15:07

    do if you have a kind of a pure

  430. 15:09

    endto-end control model that is not

  431. 15:11

    aware of the different perceptive tasks

  432. 15:14

    that it needs to accomplish in order to

  433. 15:15

    do this task.

  434. 15:23

    Yeah,

  435. 15:25

    it's pretty cool. It also works zero

  436. 15:27

    shot relatively well out of the box. So,

  437. 15:29

    we're excited to get this in the hands

  438. 15:31

    of folks in in a couple of weeks

  439. 15:33

    sometime in July.

  440. 15:38

    Um, yeah, going to leave a couple of

  441. 15:40

    minutes for for for general questions,

  442. 15:42

    but if you guys are interested, let's uh

  443. 15:45

    let's connect. So, one cool thing that

  444. 15:47

    we do with our company is we actually

  445. 15:49

    for a limited set of partners give

  446. 15:50

    access to our Mark1 weights. We give

  447. 15:52

    access to our larger embodied foundation

  448. 15:54

    models weights. So, email me, DM me on

  449. 15:57

    Twitter, uh, whatever is easier. Um, and

  450. 15:59

    then yeah, I'll open up. There's a

  451. 16:01

    couple minutes left for for questions.

  452. 16:17

    these models

  453. 16:22

    and like when I saw

  454. 16:31

    a lot of contss

  455. 16:39

    So can you talk a little bit on that?

  456. 16:41

    >> Yeah, I think it's it's so yeah. So the

  457. 16:43

    question is how are we able to nail

  458. 16:44

    temporal and uh understanding to this

  459. 16:47

    degree? Um it's a good question. I mean

  460. 16:49

    to be honest it's not nailed. So there's

  461. 16:51

    still a lot of work to actually get it

  462. 16:52

    to a place where you can reliably deploy

  463. 16:54

    it. The the core thing is how do you

  464. 16:56

    think about context management? So you

  465. 16:58

    have relatively limited context. So I

  466. 17:01

    think the the the models here have 1

  467. 17:03

    million context, but that's relatively

  468. 17:04

    easy to fit in with the, you know, high

  469. 17:07

    FPS video. And so you have to start

  470. 17:09

    thinking about, are there interesting

  471. 17:10

    things that you can do? I'll throw

  472. 17:12

    something out there. We we used to do

  473. 17:14

    this, but we got past this, but like how

  474. 17:15

    do I think about like uh uh key frames

  475. 17:18

    versus delta frames? How can I manage my

  476. 17:20

    context by training these two off? Um,

  477. 17:23

    and then you start thinking about during

  478. 17:25

    your pre-training objective, how can I

  479. 17:26

    start kind of natively ingesting things

  480. 17:30

    that I think will be useful for the

  481. 17:31

    robotics tasks. Um, so for example, a

  482. 17:34

    very basic thing that even kind of the

  483. 17:35

    the Gemini models used to struggle at, I

  484. 17:38

    think the new ones are pretty good, but

  485. 17:39

    being able to tell cardalities. So like

  486. 17:41

    left and right is very hard to tell if

  487. 17:43

    you do internet scale crawls because no

  488. 17:45

    one on the internet is necessarily

  489. 17:46

    labeling things as, you know, this

  490. 17:48

    object is to the left of this object,

  491. 17:50

    it's below this object. So really

  492. 17:52

    thinking about data distributions early

  493. 17:53

    on gives you this ability relatively

  494. 17:55

    quickly. And it's also like er this type

  495. 17:58

    of embodied reasoning was a very

  496. 17:59

    concrete focus with this which is why I

  497. 18:02

    think we were able to kind of surpass

  498. 18:03

    Gemini ER with relatively less compute.

  499. 18:12

    >> One issue with

  500. 18:16

    sometimes I was wondering whether

  501. 18:20

    you looked

  502. 18:22

    your models are more robust than

  503. 18:28

    >> Yeah. So, probably the coolest

  504. 18:29

    robustness that we've seen is that uh

  505. 18:32

    for VA models specifically, like if you

  506. 18:34

    go and you take one of the Chinese ones

  507. 18:37

    um and you try to fine-tune it for a

  508. 18:38

    specific policy, if you just change the

  509. 18:40

    background of the I don't know, you can

  510. 18:43

    like in the table uh if you change the

  511. 18:45

    the background of the table, the policy

  512. 18:47

    will actually fail. What's really

  513. 18:49

    interesting, if you do this type of

  514. 18:50

    joint perceptive and control modeling,

  515. 18:52

    you're much more robust to these types

  516. 18:53

    of errors or if the light is hitting it

  517. 18:56

    a slightly different way. And I think we

  518. 18:58

    primarily view this as robustness to

  519. 18:59

    background uh in a way that I think

  520. 19:02

    traditional models don't necessarily

  521. 19:03

    have. That being said, I'm not going to

  522. 19:05

    overclaim like uh I mean like it's still

  523. 19:07

    relatively hard. I think if I was going

  524. 19:08

    to go and shine a flashlight into one of

  525. 19:10

    the one of the arms, it's probably not

  526. 19:12

    going to work. Uh but being able to

  527. 19:14

    jointly model these things helps a

  528. 19:17

    significant amount. We also do a lot of

  529. 19:18

    online augmentations. So, so we do

  530. 19:20

    actually, you know, fake um I don't know

  531. 19:23

    um one of the one of the arm cameras

  532. 19:26

    being off, right? We fake uh you know,

  533. 19:29

    sunlight coming in from a certain

  534. 19:31

    direction, right? So, we do these things

  535. 19:33

    during training to improve robustness.

  536. 19:35

    Uh but the big gains come from taking

  537. 19:37

    this early fusion paradigm and then

  538. 19:39

    moving it into the robotics domain.

  539. 19:44

    >> Cool.

  540. 19:47

    uh create knowledge base using the K1

  541. 19:50

    model.

  542. 19:51

    >> Uh oh, knowledge bases. Um it's useful

  543. 19:53

    for like I I mean if you want to caption

  544. 19:55

    images, videos, it's relatively well. We

  545. 19:58

    work with robotics partners for like

  546. 19:59

    very complex egocentric annotation like

  547. 20:01

    this. Um so it's not necessarily

  548. 20:03

    building out an ontology, but being able

  549. 20:05

    to do kind of very deep structured

  550. 20:07

    extraction, I think our models are are

  551. 20:09

    very good at. Yeah. By the way,

  552. 20:11

    everything that I kind of showed here is

  553. 20:13

    kind of public APIs, so you can go play

  554. 20:15

    around with it. The benchmarks are

  555. 20:16

    public. Um, I think I'm out of time.

  556. 20:19

    They're cutting me off. So, uh, I can

  557. 20:22

    talk with folks outside, but thank you

  558. 20:23

    guys.