Emulated: The data for fully autonomous software engineers and companies
AI Engineer World's Fair 2026 · 16:33
Reinforcement learning environments and agent evaluation
Emulated builds reinforcement learning environments that simulate production systems to train AI models for software engineering work. Its environments address tasks beyond generating code: debugging distributed failures, migrating infrastructure, and building serverless runtimes or container orchestrators. For teams training and evaluating models, they provide running services, source code, operational tools, and realistic traffic so agents can practice changing systems while those systems remain in use.
Founders Joseph Wang and Sid Patllollu lead the company. Its Software Development Automation Benchmark, introduced in 2026, evaluates agents across 80 tasks involving deployment, debugging, migration, and maintenance. Tasks can run for up to 12 hours. The evaluation combines behavioral tests with engineering-quality rubrics, checking both whether a system works and how the agent deploys, monitors, and verifies its changes.
Emulated offers Docker environments alongside cloudboxes, sandboxed cloud environments that reproduce production infrastructure on AWS, Azure, and GCP. Cloudbox tasks use real cloud APIs, rather than mocked interfaces, and support systems such as distributed databases and serverless runtimes. This gives model developers a way to evaluate agents on evolving infrastructure state and cloud operations, with benchmark scenarios drawn from production incidents, optimization projects, and feature development.
AI Engineer World's Fair 2026 · 16:33
Affiliations reflect their AIE appearances, not necessarily current employment.
The talk explains why single-container sandboxes cannot faithfully reproduce cloud resource provisioning, network partitions, and operational telemetry needed for realistic infrastructure training.
Affiliations reflect each recorded session, not necessarily current employment.