AI Engineer World's Fair 2026

Brains vs Hands: How to Run AI Agents Safely in Production — Viren Baraiya

Read the talk

Brains vs Hands: How to Run AI Agents Safely in Production

Viren Baraiya explains how an LLM can assemble a plan at runtime while a deterministic harness controls execution, approvals and side effects. A two-loop SRE example shows why the next plan must account for what has already happened.

From a talk by Viren Baraiya

At a glance

Ideas worth remembering

  • Treat the harness as the application that coordinates agents, systems, humans and tools around a goal.

  • Let the LLM choose what should happen next; put execution procedures, required approvals and side-effect handling in the harness.

  • A late-bound saga assembles a finite set of defined tasks at runtime, allowing the plan to change while task behavior remains controlled.

  • Execution history must feed the next plan. In the SRE demo, the second iteration skips the rollback already performed in the first and moves to downstream checks.

Production agents need more than a tool-call loop

The hello-world agent makes a tool call, uses some context or memory, and plans its next move. That example leaves out much of the work involved in running an agent in production. Viren Baraiya opens with the larger picture: agents may work without a chat session, wake up on a schedule, respond to infrastructure events or coordinate other agents over a long period.

These operating modes create different responsibilities for the software around the model:

  • Scheduled work. An agent might run every hour to check a calendar and identify something new that needs attention. Its trigger comes from time rather than a user message.
  • Event-driven work. A production agent might listen for alerts and incoming logs, then investigate what is happening in the system.
  • Long-running coordination. One agent might monitor another, decide whether it needs a nudge and suggest a different action.
  • Multi-agent work. Specialized agents communicate with one another instead of assigning every responsibility to one agent. Baraiya connects that narrower scope to reducing the risk of hallucinations.
Selected presentation frame from Brains vs Hands: How to Run AI Agents Safely in Production — Viren Baraiya at 122 secondsOpen full source frame
Slide listing background, scheduled, and event-driven agents.

Once these responsibilities accumulate, the agent starts to look like an application. The question becomes how to coordinate reasoning, tools, events and ongoing work so that the whole system achieves a goal.

0:121:07
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:12 · section reference included

The harness is the application

Microservices offer a useful precedent. A production application typically combines several services, each with a focused responsibility, and connects them through choreography or orchestration. An SRE harness has a similar shape: it can coordinate agents that examine logs, interact with a Kubernetes cluster, query metrics or inspect customer-service dashboards. The business goal belongs to the assembled system.

“Your harness is the application” is Baraiya’s framing for this assembly. The harness controls agent execution while connecting databases, internal systems, humans and tools exposed through APIs or MCP. Its job extends beyond repeatedly calling a model: it must make all those participants work together.

Selected presentation frame from Brains vs Hands: How to Run AI Agents Safely in Production — Viren Baraiya at 275 secondsOpen full source frame
Diagram placing the agent within a larger application.

That application combines two kinds of behavior. LLM reasoning supplies a non-deterministic way to interpret a situation and decide what should happen. Operations such as payments or restarting a production cluster need a defined execution procedure. For a cluster restart, the harness should know the sequence of steps and run that procedure consistently whenever the operation is selected.

The useful design decision is where to put flexibility. Interpreting an unhealthy deployment can benefit from model reasoning. The procedure that carries out a restart should live in code whose behavior the application controls. Deterministic execution makes the procedure predictable; it does not make the model’s diagnosis correct.

3:083:38
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

3:08 · section reference included

Waiting and recovering are part of the work

A harness can run for a few seconds, days or months. Much of that lifetime may involve waiting: an order-management process waits for a third-party shipment, an approval step waits for a human, and an event-driven agent waits for its next trigger. The application must continue to make sense across those pauses.

Long lifetimes make infrastructure failures unavoidable design inputs. The machine or sandbox running the harness can disappear; networks can fail or partition. Durability means the harness can recover its work when those failures occur. Baraiya calls it the “cost of admission”: a production harness needs recovery before its more interesting agent behavior can be useful.

Selected presentation frame from Brains vs Hands: How to Run AI Agents Safely in Production — Viren Baraiya at 427 secondsOpen full source frame
Slide describing durability as a requirement for long-running harnesses.

Recovery also needs a meaningful record of the world. The harness tracks completed work, successes, failures and side effects. A cluster restart or a sent email must become part of that record. The next model call then reasons from the current state and the goal, including what execution has already changed.

This gives state two jobs. It lets the application retain knowledge of completed work, and it gives the planner the context needed to choose the next action. The second job matters even when nothing crashes: a remediation loop that forgets its previous actions can propose work that has already been done.

6:036:33
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

6:03 · section reference included

The brain proposes; the hands enforce the procedure

The central split is between deciding what should happen next and executing that decision. The LLM writes a plan. The harness performs the work. If the planner concludes that an unhealthy cluster should restart, the harness supplies the restart procedure, including the steps required to carry it out.

Selected presentation frame from Brains vs Hands: How to Run AI Agents Safely in Production — Viren Baraiya at 519 secondsOpen full source frame
Diagram contrasting the agent’s planning role with the harness’s execution role.

Three execution responsibilities make that separation consequential:

  • A defined procedure. The harness controls how a restart happens, so the operation follows the same prescribed process each time.
  • Required approval. A production restart might require sending someone a Slack message and waiting for approval. If that gate is required, the workflow must enforce it. The model should not be able to omit it because it decides approval is unnecessary.
  • Side-effect handling. Ideally, an operation is idempotent: repeating it does not introduce an additional effect. Where it is not, the harness records what happened so the application can handle those effects separately.

Recording a non-idempotent action does not itself make repeating the action safe. It preserves information needed to deal with the consequences. That is why execution policy and execution history both belong in the harness: the policy controls required steps, while the history tells later work what has already occurred.

“Harness is the hands. The brain is the LLM.” The metaphor names a concrete division of labor. The model can adapt the plan to a changing situation, while application code retains responsibility for the procedures and gates attached to the chosen actions.

8:088:38
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

8:08 · section reference included

Assemble known tasks at runtime

Baraiya calls this arrangement a “late bound saga.” The comparison starts with workflows whose steps developers define ahead of time. An agentic harness still has a finite set of tools and deterministic tasks, but the agent proposes how to put those tasks together at runtime.

The late binding happens in the composition. A tool’s prescribed behavior remains defined; its place in the plan can change with the situation. The resulting workflow gives the application visibility into what is happening and a way to control execution. The planner can propose one step or several steps before the harness runs them and obtains fresh information.

Selected presentation frame from Brains vs Hands: How to Run AI Agents Safely in Production — Viren Baraiya at 672 secondsOpen full source frame
Diagram of tasks combined into a branching workflow.

This addresses the difficulty of authoring every possible branch in advance. A hand-written workflow would need to anticipate many combinations of tools and circumstances. Runtime planning chooses a relevant combination from the available tasks. The tradeoff is that the application now needs a way to turn the proposed composition into executable work while preserving the procedures that make those tasks predictable.

10:0810:38
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

10:08 · section reference included

Two SRE loops, one rollback

The SRE example makes the composition visible through a completed remediation run with two iterations. Baraiya explicitly describes it as a pre-planned, canned demo. It illustrates the planning and execution structure rather than establishing how reliably the system diagnoses unfamiliar production incidents.

The first LLM call proposes a sequence: gather evidence, analyze logs, roll back a deployment if the evidence calls for it, and verify recovery. A specialized tool named Plan and Compile receives that plan and compiles it into a runnable workflow for Conductor, the workflow execution engine used in the example.

The first workflow analyzes logs, queries metrics, performs a rollback and verifies recovery. Those results then become the starting point for another planning loop. The second iteration recognizes that the earlier work has happened, checks downstream systems and calls for recovery verification if needed before completing the task.

What changes between the two plans, and where does that change enter the system? The diagram follows the rollback through execution and back into planning. The second plan omits another rollback because the first iteration already performed it; the loop continues with checks appropriate to the updated state.

The observable change is in the work selected for the second iteration. The system goes from diagnosing and rolling back to checking what remains downstream. Its causal sequence is straightforward: execute the first plan, observe the output, reason from the resulting state, then compile and execute the next plan. Completion history changes the plan rather than merely documenting it afterward.

Each iteration can plan several steps at once, execute them and verify the result before the next iteration. The same arrangement can respond to events, run on a schedule or continue over a longer period. The two-loop demonstration is a short instance of that operating model.

The ending identifies Conductor as an open-source workflow orchestration platform that supports agentic loops and systems, with an Enterprise Edition provided by Orkes. Baraiya also describes support for running LangChain and OpenAI agents. In this architecture, the workflow engine supplies the execution layer around agent reasoning: the agent framework can produce plans, while the harness manages the work those plans select.

How it fits togetherExecution changes the next SRE plan

Gather evidence, analyze logs, conditionally roll back, verify recovery.

The first iteration performs a rollback. Its recorded result becomes input to the second iteration, which checks downstream systems without repeating the rollback.

11:5012:02
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

11:50 · section reference included

Resources

Read the complete timestamped transcript
  1. 0:12

    Cool. Um, thank you all for joining this morning. Um, I know it's been an exciting, uh, week. A lot of talk about agents, uh, building agents, coding agents and everything. And, uh, what I want to talk, uh, in the next fifteen, twenty minutes is, is about running agents in production. Um, and things that is kind of on the top of pretty much everybody's mind when it comes to building agents and running agents is, you know, harnesses.

  2. 0:43

    Before I get started, like, I would like to ask in the audience, right, um, how many of are-- how many of you are running agents in production today? Awesome. So a lot of hands up. Um, I was in another conference a couple of weeks ago, and I asked the same question. I think I had, like, one hand, so this is pretty good. It seems like, you know, things are moving forward. Um,

  3. 1:07

    all right. So when we think about agents, right, like, uh, if you go back to kind of the Hello World equivalent examples of agents, uh, it's always starts with, you know, an agent with a tool call and how an LLM kind of decides to call the tool, maybe use some sort of context or a memory and, and plans things ahead, right? Um, what, what we start to kind of realize is, you know, that's a very small narrowed picture if you think about it. Uh, when you think about agents in production, it's not

  4. 1:37

    just always about chatbots. Um, the early examples that we had seen were all about, like, you know, hey, we are gonna build a customer service chatbot or some sort of chatbot. Um, but when we start thinking about agents in a broader terms and picture, uh, think about background agents, uh, background workers, agents that are running on schedule. Um, I would like to have an agent that runs every hour to check, uh, what's happening with my schedule and if there is something new that has popped up that I should think about.

  5. 2:07

    Uh, event-driven agents, right? Agents that are listening on various events. Um, to give you an example, I have a production system running where my alerts are firing every now and then. I have, uh, logs coming through Otel. I would like my agents to listen to those events and react on it and see what's going on there, right? Um, I could have long-running coordinators, and what I mean by that is I have an agent running for a longer period of time. I would like to have an agent monitoring that agent to see, hey, what's going on there, and should

  6. 2:38

    I nudge that agent to kind of move forward, take a different action, and things like that, right? Um, also one more thing is about multi-agent systems, right? Agents talking to other agents. Uh, the last thing you want is to have an agent be responsible for multiple things and risk hallucinations. Um, so more we think about agents, right, uh, the way I would like to think about is the agent starts to feel more like an application, not just a

  7. 3:08

    component, right? Um, if you go back to the old school thought of microservices when we-- and probably we are still building microservices, right? Um, you are not going to ship a microservice in production that does everything, right? It kind of stops being a microservice at that point in time. Uh, typically, we are shipping multiple microservices. They are talking to each other through choreography or orchestration, and you have an application that is relying on

  8. 3:38

    that set of microservices, uh, which essentially are single responsibility principles, uh, to kind of deliver the business goals. There is kind of a very clear, uh, parallel to that when it comes to agent harnesses, right? If you think about a harness that is responsible for achieving a specific goal. So let's say I have an harness for doing my SRE work. Um, it's operating on multiple agents to see what's going on there, right? Maybe it's listening on events coming

  9. 4:08

    from my logging system, talking to my Kubernetes cluster as an agent, uh, to find out, you know, what's going on there. Uh, my metric systems, my customer service, uh, dashboards and, and so on and so forth. Um, and what I want to drive there is that if you think about it, right, your harness is the application and vice versa, right? An agent alone doesn't kind of deliver the whole premise, but rather when you start putting things together and you build a harness that controls the

  10. 4:37

    execution of the agent, delivers on the premise, that's where you start thinking about, um, an application, right? And which is you have systems, uh, which are your databases, your internal systems, enterprise systems. You have humans, human in the loop. Um, you have tools through either APIs or MCP that is all kind of playing together. Um, and as I mentioned, right, harness runs more than just an agent loop. And this is an important thing that I kind of learned,

  11. 5:07

    uh, while trying to build harnesses for real-world production use cases. It's a combination of both deterministic and non-deterministic part of the application, right? Non-determinism is delivered by an LLM in terms of reasoning, in terms of thinking process. Um, and then there is non-deterministic-- and there is deterministic part, right, which you do not want non-determinism sitting into it, right? Think about payments. Um, or let's say if I have an agent

  12. 5:37

    harness that is responsible for monitoring my production deployment, how I manage my Kubernetes clusters, I probably want to have a very well-defined workflow in terms of what are the sequence of steps that I execute when I want to restart my cluster. And that's a very deterministic set of processes which every time it runs, I know exactly what it does. So I want a determinism to be delivered by my harness, uh, when it matters.

  13. 6:03

    Um, and of course, the harness is a long-running process, right? It runs across the time. It can run, uh, anywhere from few seconds, if it's, uh, something very quick, like, "Hey, check what's going on here," to all the way running for days, months, uh, even longer than that, right? Think about long-running processes, order management systems where you are waiting on third-party services to deliver your shipment or waiting on humans to take actions and approve things and so on and so forth. Or harness is just waiting,

  14. 6:33

    right? It's waiting for events to happen, so when that event happens, you take an action and do something around it. So

  15. 6:41

    that brings another important point, right? That when you think about long-running systems, you need durability, right? You want to be able to recover when things fail, because at the end of the day, these harnesses are running somewhere in your entire infrastructure stack, right? Maybe in the cloud, maybe on a sandbox, but those things can go up, go down. There could be network failures, partitioning, anything that could be happening, so you need the harnesses to be durable. And one thing about durability here is it's a table stake thing, right? At the end of the day, durability is a cost of

  16. 7:11

    admission. That's not the feature that you are looking for in a harness. Um, and as I mentioned, right, the loop that the harness runs spans across agent, um, and everything else.

  17. 7:27

    So now let's think about, uh,

  18. 7:31

    how the harnesses kind of operate, right? Um, if you think about the harness, right, and, and the loop, essentially what it is doing, it, it has a state of the world. It knows what work has been completed, um, what has been recorded in terms of the side effects, right? So like if I did a cluster restart, uh, I know I have that recorded. Uh, if I sent an email, I know that has happened. Um, I know what worked, what didn't work, and then based on the current state of the world and the goal, I know

  19. 8:01

    what needs to happen next, right? And that's where the reasoning and LLM comes into the picture.

  20. 8:08

    Um, and what really happens here is that if you think about a clear distinction, and this is the most important thing. If, if one thing that I would like everyone to take away from here is this slide, which is that the responsibility of a non-deterministic agent is to plan, is to plan what should happen next, not really to do things. Um, and then harness is the one who does actually execution. And, and this is important for, uh, v-various reasons, right? One being that

  21. 8:38

    harness is a deterministic piece of code that actually understands what is involved in, you know, actually executing this piece of work. So as I mentioned, right, if I am trying to, uh, build a harness that runs my DevOps or SRE, uh, agents, uh, when the agent says that, "Hey, this cluster is unhealthy and you should restart," harness should decide how to restart the cluster, what involves in restarting the cluster, and that has to be, at least in my

  22. 9:08

    world, has to be very deterministic set of processes, right? So that every cluster restart is exactly same. There is no other, uh, category there, right? Um, also, if it needs to have an approval, for example, if I am trying to restart a production cluster, I probably want to have a human gate, probably send a Slack message to somebody and say, "Hey, I'm gonna restart this cluster. Do you think it's okay to do that or not?" Right? And I do not want this to be left to a hallucination by an LLM that, you know, it doesn't need to do it. It needs to be guaranteed

  23. 9:38

    in terms of execution. So it has to be very deterministic when it comes to those kind of things. Uh, ideally, I want this to be idempotent, and if not, I want it to be recording that, you know, it is not idempotent, and this is what has happened, so I can ca-- take care of the side effects, or I can handle the side effects, uh, separately, right? Um, so this is the most important thing, right? The clear separation between, as we would like to call, right, the brains and the hands. Harness is the hands. The brain is the LLM. Um, so

  24. 10:08

    yeah, if the agent writes the plan, uh, harness executes the plan. Um, and I'll show you in a brief, uh, a short demo in terms of how all of these things kind of meshes together. Um, but is this a new concept, right? If you think about it, this is not necessarily a new concept. This has been around for a while, right? If you think about workflows as sagas, which we used to write ourselves and define exactly what happens, it was a very-- and it is a very deterministic set of processes.

  25. 10:38

    When you think about agentic harnesses, they are essentially late bound sagas. And what I mean by that is that they have a very finite and deterministic set of tools, uh, and the tasks that they can operate on. Instead of putting them together upfront, the agent is kind of proposing and building them at runtime. Um, and therefore, they are essentially late bound sagas, right? Uh, but you get all the benefits of a, a traditional saga in a workflow out of the box in terms of visibility, what's happening, being able to

  26. 11:08

    control things, and iterating upon like, you know, how far ahead in the future, uh, the agent is able to plan. You are able to kind of plan one step at a time or multiple steps at a time.

  27. 11:21

    Um, and yeah, if you think about it, right, it's, it's more like a branching workflow. If only you could build a workflow with every possible combination of a branch, um, then you know, it kind of builds that thing for you, right? Versus with agents, that kind of simplifies the work. Um, if you have N number of tools, it can do, um, N different combinations of executions, which otherwise is gonna be almost impossible to, you know, think ahead of time and, and do it.

  28. 11:50

    So let's take a look at it, right? So what I am going to do is, uh, I'll quickly show, uh, a completed run

  29. 12:02

    of what I talked about earlier, right? Um, here is one of my example agent that essentially is- Acting like an SRE agent, right? What its job is to do is understand what's going on in the current system, um, and try to plan what should happen next and execute on those things. So it's essentially a remediation loop that runs twice. Uh, in the first iteration, it tries to understand the root cause of the problem, tries to, uh, uh,

  30. 12:32

    react to that, uh, observes the output of it, and then runs another loop and plans another set of, uh, steps to see what should happen next, right? So as an input to my agent, let's see what was the input given right here. Um,

  31. 12:49

    so step number one is, you know, it makes a call to LLM. So as you can see, right? Like, it's an SRE agent. Um, of course, it's a demo, so, you know, everything is kind of pre-planned and, uh, a canned demo. The output of LLM, as you can see, right? Are the steps, uh, what it should do. It's talking about... So as you can see, one more thing here, right, is that instead of just doing one step at a time, essentially it is proposing a sequence of steps to do. Um,

  32. 13:19

    you should gather evidences, you should analyze logs, and if required, um, based on the evidence that you gathered, you should either roll back a deployment and then verify recovery. This is then given to, uh, a specialized tool called Plan and Compile. So this is the plan that agent gave. This gets compiled into a very deterministic workflow. Um, we are relying on a Conductor as a workflow execution engine here, so the output is a fully runnable workflow.

  33. 13:49

    This gets executed, so, like, you know, what I'm gonna show you here very quickly is how that looks like. Um,

  34. 14:01

    so this is the first step of execution, right? If you look at it here, as I said, right, like, we did, uh, analyze the logs, look at the query of the metrics, decided to do a rollback and verify recovery and everything. So step number one is completed. It runs another loop. Um, same thing goes here. Um,

  35. 14:26

    in the next iteration, it decides that, okay, looks like, you know, these things have been done. Let's check the downstream, and if required, uh, verify the recovery and complete it, right? Um, and then the next one runs and completes the task. But here is an example of a loop that is self, uh, kind of, uh, planning, right? So, and to show you very clearly what's going on here, here is a loop. Um,

  36. 14:52

    at every step of the way, essentially it is looking at the current state of the world. So this is my agentic loop that is kind of understanding the world, planning and executing, uh, one st- instead of one step at a time, it is actually planning multiple steps at a time, executing that, verifying and completing the loop. Um, now this can run in production as many times as you want. This can be running on the events. It can be running on a schedule. Um, as I said, right? Like this, the whole thing could be running for a much longer period of

  37. 15:22

    time, um, as opposed to just running, uh, in a short period of time.

  38. 15:30

    All right. So that's towards the end of it. Um, this is kind of, um, overview of what we just did here. Um, as you can see, right, in the second, uh, iteration, we decided not to do a rollback because it was already done in the first iteration. Um, but yeah, that's about, uh, the harnesses and agents and how they bring determinism, uh, to your production application. The example that I showed you runs on Conductor. Uh, Conductor is an open source, uh, workflow orchestration

  39. 16:00

    platform, uh, that supports building agentic loops as well as agentic systems. It's fully open source. Um, Orkes, we provide Enterprise Edition, uh, but feel free to try it out. Here's a QR code. Uh, give it a try. Uh, join our Slack and, you know, if you have questions, happy to help. We have a booth here at Orkes. Uh, drop by if you want to see a live demo. Um, if you want to run your LangChain agents or OpenAI agents or Indica agents, Conductor can handle all of those things.

  40. 16:30

    Thank you.