← All speakers

Bio, Work & Ideas

Sid Patllollu

Conference affiliation: Emulated · 2026

On this page

Sid Patllollu co-founded Emulated with Joseph Wang to build training environments for more reliable, autonomous AI agents. His work brings infrastructure operations into the learning problem: agents must learn to change running systems, manage failures, and take responsibility for what happens after deployment.

From infrastructure to autonomous engineering

Patllollu and Wang brought backgrounds spanning network infrastructure, distributed databases, and sandbox infrastructure to Emulated. Experience with mission-critical systems shaped their shared concern that proficiency at writing application code does not necessarily prepare an agent to operate the infrastructure beneath it. Their approach expands software-engineering tasks to include customer problems, operational history, performance testing, and the consequences of successive changes.

Emulated’s etcd consensus-cluster environment, explained by Patllollu, makes that distinction concrete. An agent receives tickets, projects, and postmortems alongside the source code; some of that organizational context may be outdated. Making a change then requires navigating rolling deployments, failing or obsolete nodes, and unexpected migration problems while keeping the service available. Monitoring supplies feedback about consequences that a code diff alone cannot reveal. This example captures his interest in operational responsibility: an agent’s job continues as its decisions affect a running system. The infrastructure approach connects organizational context with deployments, live traffic, and distributed-system failures.

Emulated initially used single-node sandboxes to simulate distributed clusters. The founders also set out a direction toward multi-node environments that provision real cloud resources, where agents would confront resource management and service operations more directly. That ambition brings practical constraints of its own: infrastructure can take substantial time to start, costs must remain manageable, and even real resources do not automatically reproduce customer workloads or failures that emerge at scale.

Measuring work beyond the first solution

Two team-built benchmarks develop complementary parts of that ambition:

  • Software Development Automation Benchmark: Emulated’s software-engineering benchmark places agents in environments with running services and realistic traffic. Behavioral outcome tests assess what the resulting system actually does, while engineering rubrics cover deployment, verification, and monitoring. Together, these measures examine whether an agent can carry a change through the operational work needed to make it useful. They give concrete evaluation criteria to the responsibility embodied in Patllollu’s infrastructure example.
  • Autoresearch Bench: Patllollu helped build Autoresearch Bench, published in September 2026, with colleagues at Emulated. It tests iterative experimentation across optimization, model training, inference, and scientific machine learning. Agents receive an objective, a metric, and a time budget, then run experiments, measure results, and choose their next attempt. Many tasks pair visible grading feedback with a held-out evaluation score, making it possible to distinguish improvements that generalize from changes that merely exploit the feedback an agent can see.

Emulated supplies post-training datasets for software engineering, machine learning engineering, and autonomous research. Patllollu’s work connects the maintenance of complex services with the experimental discipline of research: both ask agents to interpret feedback, revise an approach, and keep working after an initial answer.

Read the topics behind these talks

1 conference talk

Key ideas

Scroll to read ↓

Reliable infrastructure agents need environments that expose deployment failures, customer context, and cloud operations—not just the source code they must change.

  • Why can an agent write the application but struggle with its infrastructure?
    0:15 ↗
  • Expand the task from a repository to a company
    2:45 ↗
  • Follow an etcd change into a running cluster
    4:35 ↗
  • How far can one sandbox go?
    6:24 ↗
  • Build outward from software to a cloud service
    8:00 ↗
  • Move the environment onto real infrastructure
    10:36 ↗
  • Real resources still leave training constraints
    12:09 ↗
  • From infrastructure ownership to broader company workflows
    13:22 ↗

References