The Best Models Still Reason Like Toddlers — Andrew Dai, Elorian
AI Engineer World's Fair 2026 · 18:16
Organization in the AI Engineer archive
Conference talks featuring speakers affiliated with Elorian when their sessions were recorded.
Elorian describes its work as building a foundation for visual thinking. Its official team page lists Andrew Dai as co-founder and CEO, Yinfei Yang as co-founder and chief multimodal architect, and Dustin Tran as chief reasoning architect, alongside staff working on multimodal learning, data, infrastructure and physical autonomy. Dai’s recorded presentation offers a technical entry point: it separates recognizing a scene from counting its contents, tracking changes and checking physical constraints. The company page documents the organization; the demonstrations and proposed capabilities below are Dai’s account, without independent performance validation in the supplied material.
Start with Dai’s board-game examples. A model answers 32 white squares for a partial chessboard, substituting a fact about a complete board for a count of the visible image. In Catan, a response uses 10 blue roads beside the board to infer five placed roads; Dai reports seven when counted directly. These examples expose a tradeoff: pattern recognition supports useful identification of plants and objects, but familiar categories can override scene-specific evidence. They are reported failures, not measured failure rates, and the supplied text does not include the images needed to verify the counts.
A robot-arm sequence extends the problem from space to time: Dai reports models missing a lifted lid and a later stove activation. His diagnosis is loss of state across video. He distinguishes quick understanding from deliberate reasoning with a practical heuristic: could a person answer after one second of inspection? This is guidance for task design, not a universal capability threshold. His benchmark criticism asks whether success requires the relevant visual detail. He describes reasoning tasks at 32 × 32 or 64 × 64 pixels and multimodal science questions answerable without inspecting the image. Those are his characterizations; the science benchmark’s exact name is unclear in the unreviewed captions. The useful evaluation question is whether geometry, spatial relationships and temporal changes are necessary to reach the answer.
Dai distinguishes visual thinking from high-fidelity generation and passive detection: convincing imagery need not establish physical causality, while object labels alone do not provide planning logic. He describes four parts of Elorian’s approach: collected and generated visual-reasoning data, a synthetic-data flywheel using evaluations, agents, supervised fine-tuning and reinforcement learning, changes to transformer-based architecture, and native visual chain of thought. The hotel example makes the last part concrete: identify candidates with boxes, then narrow them to red hotels. Intermediate selections remain attached to the image rather than relying on its general appearance. The presentation supplies neither a final hotel count nor comparative accuracy, and leaves architectural details and the training loop’s operation unspecified.
For robotics, Dai proposes action-relevant scene understanding feeding existing planning and control systems. He presents a model API release by the end of the talk’s year as a plan; the supplied sources do not establish availability. Construction adds a language-and-vision requirement: relate camera streams to changing, zone-specific safety policies, count workers wearing helmets, or compare progress with plans. Dai contrasts repeatedly training separate models with interpreting written rules alongside video using existing cameras. These are proposed applications and capabilities being built, rather than documented deployments or demonstrated safety outcomes.
The closing technical discussion connects counting and spatial grounding to blueprints, 3D CAD and CAM files. Dai reports a mechanical engineering company describing 100–200 hours to design a small part of a robot testing platform; he estimates 2,000–3,000 human hours for the whole platform. These are an anecdote and an estimate, not industry-wide productivity evidence. His proposed workflow extracts geometric logic through multimodal reasoning, then checks the design programmatically or in simulation, analogous to testing code. CAD/CAM quality control is another potential use. He ends with ambitions for faster cars, more efficient rockets and better batteries, emphasizing their visual and physical constraints, then directs listeners to company channels and offers to take questions. The ending sets a research direction without establishing achieved engineering gains.
AI Engineer World's Fair 2026 · 18:16
Affiliations reflect their AIE appearances, not necessarily current employment.