← All organizations

Organization in the AI Engineer archive

Skild AI

Conference talks featuring speakers affiliated with Skild AI when their sessions were recorded.

Explore the recordings

Skild AI’s supplied official website describes an organization building a unified, “omni-bodied” robot brain, with applications in security and inspection, mobile manipulation, and autonomous packing. It also describes learning from human videos and exposing manipulation and navigation skills through API calls. Deepak Pathak’s recorded presentation explains the technical rationale behind that thesis: how complementary data sources could support control across different bodies, and what manipulation, locomotion, and hardware failures demand of a shared model. The recording is useful for its mechanisms and concrete constraints; its demonstrations and deployment claims remain speaker reports rather than independently established performance results.

One brain across bodies: a response to scarce data

Pathak uses historical robotics demonstrations to question whether impressive isolated behaviors establish general capability. His diagnosis—that robotics has emphasized hardware over a general brain—is the talk’s argument, not a settled history of the field. He estimates roughly one minute to collect a teleoperated example, illustrating the labor and physical time required. Supporting multiple robot bodies is therefore a data strategy: experience from arms, quadrupeds, humanoids, and different tasks could contribute to one model. The website documents this as Skild’s thesis; the recording does not establish universal compatibility with arbitrary robots.

Combine data sources with different weaknesses

The central framework evaluates data by scalability, diversity, and closeness to the robot’s own joint-angle ground truth. Autonomous exploration supplies direct physical experience but remains constrained by hardware and time. Teleoperation supplies high-quality robot actions but requires operators and can repeat one environment. Simulation scales trials, while scene engineering limits diversity. Human video offers breadth but requires translating human movement into robot action. Pathak assigns simulation and video to pre-training, smaller teleoperation datasets to post-training, and deployment experience to subsequent training. The proposed flywheel depends on achieving deployment scale. His retrospective explicitly includes research predating Skild and should not be treated wholesale as company-authored work.

Precision depends on the grasp and the task’s tolerance

Pathak contrasts forgiving laundry folds with inserting imitation AirPods into a case. A parallel-jaw gripper cannot manipulate an object like fingers, so the arm must choose a grasp that preserves the orientation required for insertion. He specifies approximately $5–$10 imitation earbuds without magnets, removing mechanical assistance from the fit. A separate human-video example uses egocentric demonstrations followed by less than one hour of additional robot data. Third-person video remains under investigation in his account. The robot’s “dream” footage represents model-generated scenarios, not physical trials; the talk does not explain how those imagined outcomes are checked against real physics.

Camera-guided cooking and the scope of end-to-end control

The omelet demonstration uses a setup costing roughly $4,000, with a camera and no force sensing. Pathak reports less than ten hours of training data, behavior without a hand-written start-and-stop state machine, and tolerance of changed objects; the pan and gas stove remain fixed. He presents this as extracting more capability from existing hardware, while explicitly allowing that better sensors would help. His broader control description runs from camera observations to motor power without explicit mapping and planning stages. A final control component bridges 100 Hz to 500 Hz, but its name is unclear in the unreviewed captions, so “end to end” should not imply that lower-level control is absent.

Stairs and deployments test interaction with surroundings

Pathak argues that stairs can be harder than backflips because unfamiliar step geometry and disturbances require perception to guide action, whereas a body-centered maneuver primarily concerns known robot dynamics. He describes torso-camera locomotion without an explicit 3D map. Later examples include GPU assembly work for NVIDIA’s Houston factory and package delivery from truck to front door. These are his deployment reports, not confirmations supplied by the company homepage. Front-door identification adds scene understanding to mobility, while assembly combines precision with a noisy setup. Relative deployment dates cannot be anchored to the upload date, and disturbance tolerance alone does not establish safe operation alongside people without guarding.

The ending: adaptation when the body changes

The closing examples extend the shared-body thesis to damage: a broken or disabled limb effectively creates a different robot. Pathak reports walking on two legs after three trials and switching from wheels to walking when a wheel is jammed. He describes adaptation across examples ranging from milliseconds to roughly 30 seconds, rather than one universal recovery time. He ends by reporting a further experiment in which two robot halves operate separately, noting that this footage was omitted and directing viewers elsewhere. These examples illustrate the proposed value of adapting to changed embodiment; they do not establish that continued operation after damage provides a general safety guarantee.

1 talk

Newest first

1 speaker at AIE

Affiliations reflect their AIE appearances, not necessarily current employment.