← All organizations

Organization in the AI Engineer archive

Dyna Robotics

Conference talks featuring speakers affiliated with Dyna Robotics when their sessions were recorded.

Explore the recordings

Dyna Robotics develops robot models and hardware for physical workflows. Its supplied official website describes laundry, hospitality, and factory operations, and presents a stack spanning reasoning, task dexterity, whole-body control, and sensing. Jason Ma’s recorded presentation offers a technical account of a narrower question: how a generalist manipulation policy learns to keep working after mistakes. The recording is useful for understanding recovery-data collection, customer-defined quality, and the distinction between mastering a task and transferring it to a new site. Its demonstrations and performance figures are speaker-reported, rather than independently verified results.

Read the company source and recording as separate accounts

The official website names DYNA-VLM for high-level reasoning, DYNA-2 for mid-level task dexterity, DYNA-System0 for whole-body control, and DYNA-SAUR for embodiment and sensing. It describes DYNA-2 as trained on a million hours of human activity video. Ma’s recording instead discusses Dyna-1, a reasoning model paired with a world action model, and more than 200,000 hours across a training pipeline. These quantities describe different supplied accounts and should not be combined. The website’s workflow throughput and delivery-quality claims also do not establish the performance of the recording’s napkin-folding model.

Broad training supplies knowledge; deployment exposes gaps

Ma describes three data sources: off-robot human recordings and public datasets for breadth, robot demonstrations for precise actions, and deployment experience to reduce the gap between laboratory and customer environments. Simulation is under consideration in his account. High-level reasoning supplies semantic understanding, while the world action model produces fine-grained movements; the talk does not specify their interface or control frequency. Two tool-use demonstrations reportedly required less than one hour of task-specific data. Such adaptation motivates a general foundation, but does not by itself establish sustained commercial reliability.

Napkin folding makes reliability measurable

The robot must select exactly one napkin, fold it to an acceptable geometry, and place it in a bin. Parallel-jaw grippers can pull out extra cloth, and Ma contrasts acceptable grade-five and unacceptable grade-three folds with roughly one inch of seam-position difference. He reports that initial pre-training and post-training reached about 80% success, with errors leaving the policy stuck in unfamiliar states. Dyna-1 reportedly achieved 99.4% napkin-folding success over 24 hours; Ma also presents four distinct 24-hour office trials. The supplied material does not give evaluation sample sizes or independent verification, so this figure remains specific to the reported napkin evaluation.

Progress scoring selects recovery lessons

A reward model watches robot video and estimates progress from zero toward one. Dips or reversals flag possible mistakes, directing operators to difficult cases instead of requiring continuous manual observation. Humans collect targeted recovery demonstrations, fine-tune the policy, and repeat the active-learning cycle. The reward model helps select what to teach; it is not itself the recovery controller. Ma’s examples include separating extra napkins and recovering after pulling over the entire stack. Because deformable cloth has too many configurations for exhaustive demonstrations, the important claim is transfer to recovery situations beyond those explicitly collected.

Separate site adaptation from transfer

Ma describes restaurant napkin folding and towel folding at a Sacramento laundromat, explicitly qualifying those deployments as using data from the customer sites. A worker replenishes the laundromat bin, so the example includes human material handling. Separately, he reports three days of T-shirt folding at CoRL in Korea without additional site data, including attendee interference with the camera. That account does not carry the napkin evaluation’s 99.4% figure. His Red Bull can-opening example rests on his description because the demonstration video did not play properly.

The closing Q&A defines the remaining limits

At the time of the recording, Ma says enterprise deployments are the focus, a developer kit is not being developed, and education use cases have not been explored. Consumer deployment remains an ambition requiring greater polish and attention to privacy and safety. He sketches speech-to-text feeding text instructions into the robot models, while acknowledging that arbitrary commands exceed their manipulation capabilities. The final technical answer returns to broad pre-training: recovery can draw on experience from other tasks, followed by task-specific post-training and targeted active learning. This is his argument for a reusable training recipe rather than specializing each model from the outset.

1 talk

Newest first

1 speaker at AIE

Affiliations reflect their AIE appearances, not necessarily current employment.