← All speakers

Bio, Work & Ideas

Viren Baraiya

Conference affiliation: Orkes

On this page

Viren Baraiya co-founded Orkes and is an original creator of Netflix Conductor, the open-source engine for coordinating work across distributed services. He was Orkes’s CTO at the AI Engineer World’s Fair 2026. His work addresses a persistent problem in software: keeping a process understandable and recoverable when it crosses systems, encounters failures, or waits days for its next step.

From developer platforms to Conductor

Baraiya’s career has centered on infrastructure that other engineers build upon. He led engineering teams at Goldman Sachs, Netflix, and Google; at Google, his work included Firebase and Google Play. At Netflix, he originated the idea for Conductor and helped architect and build it with fellow engineers.

Conductor grew out of Netflix’s content operations: integrating studio partners, ingesting and encoding video, and preparing titles for distribution. These processes could span days and involve many independently running services. Coordination scattered across messages, service calls, and database records made an operational question difficult to answer: what still needed to happen before a title was ready?

Netflix open-sourced Conductor in December 2016, with Baraiya and Vikram Singh co-authoring its introduction. The engine separated a process’s blueprint from the workers performing individual tasks. It tracked progress centrally, scheduled subsequent work, and exposed failures and retries through a visual interface. Teams could change and scale their services while retaining an intelligible account of the larger process.

After working at Google, Baraiya reunited with former colleagues to build Orkes. Orkes Cloud launched in February 2022, extending Conductor into a managed service that took on cluster operation, tuning, patching, and availability while adding enterprise access controls. His subsequent work applies orchestration to processes involving services, people, and AI agents.

Making consequential actions recoverable

Baraiya’s technical writing connects established distributed-systems practices with the uncertainty introduced by language models. His approach distinguishes deciding what to do from reliably carrying out the decision:

  • Distributed transaction recovery: His work on transactional backends treats a business process as a sequence of smaller transactions, with explicit recovery actions when something fails. If an order cannot be delivered, recovery may require a refund, a customer notification, and a database update. He also separates the workflow’s progress from the business entity’s state, allowing multiple processes to operate on the same order without forcing all coordination into one database.
  • Late-bound sagas: Baraiya uses this term for workflows whose structure develops during execution. A model proposes the next action; the runtime records that intention, performs the work, and records the result before asking for another decision. His recruiting example searches for candidates, enriches their records in parallel, then waits days for interview responses. The process must retain its place and accept new instructions even when the worker that started it has exited.
  • Separating reasoning from execution: In his architecture for production agents, a model’s proposed tool call passes through validation before it can affect an external system. Unknown tools are rejected, and required guardrails belong to the execution graph. He favors small agents with narrow responsibilities and limited capabilities, composed into larger applications whose behavior engineers can inspect. His “Brains vs Hands” talk makes this division concrete with an SRE agent: the model’s plan is compiled into a Conductor workflow, executed, checked, and revised in a subsequent loop.
  • The limits of orchestration: Baraiya argues against rebuilding execution infrastructure inside agent harnesses, while distinguishing reliable execution from sound reasoning. Persistent state, safe retries, approvals, and recovery do not establish that an agent’s diagnosis is correct or its plan sensible. Context selection, planning, and evaluation remain separate engineering problems. His Kubernetes example illustrates the distinction: an agent can recommend a restart, while a workflow handles notification, approval, draining, health checks, and recovery.

Read the topics behind these talks

1 conference talk

Key ideas

Scroll to read ↓

Viren Baraiya explains how an LLM can assemble a plan at runtime while a deterministic harness controls execution, approvals and side effects. A two-loop SRE example shows why the next plan must account for what has already happened.

  • Treat the harness as the application that coordinates agents, systems, humans and tools around a goal.
    3:08 ↗
  • Let the LLM choose what should happen next; put execution procedures, required approvals and side-effect handling in the harness.
    8:08 ↗
  • A late-bound saga assembles a finite set of defined tasks at runtime, allowing the plan to change while task behavior remains controlled.
    10:08 ↗
  • Execution history must feed the next plan. In the SRE demo, the second iteration skips the rollback already performed in the first and moves to downstream checks.
    7:31 ↗