▶ Watch ↗AI Engineer World's Fair 202620:41
From VLM/VLA's to Embodied Agents — Armen Aghajanyan, Perceptron AI
Read the full talk →Key ideas
Scroll to read ↓Armen Aghajanyan explains Perceptron AI’s approach to combining perception, reasoning, and control: learn useful visual targets, spend compute on relevant tokens, and use video pretraining to reduce the need for expensive robot demonstrations.
- Visual supervision needs useful targets as well as density. Predicting every pixel can spend learning effort on background details; Perceptron proposes automatically learning future percepts, without disclosing the objective here.4:39 ↗
- Learned token routing makes compute allocation task-dependent: a general question spreads attention across an image, while fruit segmentation concentrates more tokens on likely fruit.7:22 ↗
- Perception can itself involve actions. Tiling, zooming, changing contrast, and revisiting video intervals let a model gather better evidence before producing a box or annotation.10:02 ↗
- Joint training reportedly lets 10× more video pretraining substitute for 10× less teleoperation data within Perceptron’s tested compute range, offering a way to reduce dependence on demonstrations costing around $100 per hour.12:52 ↗
- Combining reasoning and control enables tasks such as sorting books by their titles, but temporal reliability and resistance to severe visual disruption remain open problems.14:31 ↗