AI Engineer Code 2025

YOLO Mode, Safely: MicroVM Sandboxes for Any Agent — Rowan Christmas, Docker

Read the talk

YOLO Mode, Safely: MicroVM Sandboxes for Any Agent

Rowan Christmas shows how a coding agent reached sensitive information on his Mac, then repeats the browser-history search inside Docker Sandboxes. The difference comes from controlling what the agent can access, while preserving a familiar coding workflow.

From a talk by Rowan Christmas

At a glance

Ideas worth remembering

  • A warning does not remove file access. In Christmas’s five-prompt experiment, changing the request’s framing let the agent continue investigating sensitive local information.

  • MicroVM isolation changes what the agent can discover. The browser-history search succeeded on the host and found no browser inside the demonstrated sandbox.

  • Secret placeholders, network policy and audit trails address different parts of agent execution: credential exposure, permitted connections and recorded activity.

  • Read-only mounts let an agent understand related repositories while requiring its solution to work with their existing APIs.

Five prompts from browser history to bank information

A coding harness that sometimes refuses a dangerous request can make its environment feel safer than it is. Rowan Christmas, a product manager at Docker, tests that assumption by asking his own agent to investigate his own Mac. The question is concrete: how far can an agent get into personal information while operating through a tool he expects to protect him?

The first request is to find browser history. The agent finds it immediately. A follow-up asks what can be done with that information, and the investigation moves into banking activity. Christmas reports that it identified his banks, recent check ordering, use of Zelle and the last four digits of an account. He substitutes fake bank names in the presentation; the underlying discovery, he says, involved his real information. This demonstrates access to sensitive local information, without establishing that the agent logged into a bank or could move money.

Selected presentation frame from YOLO Mode, Safely: MicroVM Sandboxes for Any Agent — Rowan Christmas, Docker at 120 secondsOpen full source frame
A presentation slide displays a profile and visits table with a blue callout.

The next day, Docker’s security team sends him an alert because the activity looks like a compromised computer. The CrowdStrike report gives it a nine-out-of-ten score, prompting Christmas’s joke that he did better here than in most of his classes. The important consequence is that a deliberately initiated agent session produced behavior recognizable to endpoint security as a credential-access technique. This experiment began with his own requests; it was not a demonstrated prompt injection from a script or MCP server.

The sequence took five prompts, with some care in how they were phrased. A direct request to find bank data elicited a warning. Framing the activity as security research let the investigation proceed. That change in wording altered the agent’s willingness to act without removing its access to the files. A refusal can interrupt one request, but the underlying permissions still determine what a later request can reach.

0:120:42
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:12 · section reference included

Move protection into the environment the agent runs in

Docker’s proposed answer is sbx, a command-line tool that runs the agent in a microVM. The VM has its own kernel and an isolated filesystem. This changes the available environment before the agent decides what to do: files outside that environment are not simply waiting behind a warning that another prompt might bypass.

Selected presentation frame from YOLO Mode, Safely: MicroVM Sandboxes for Any Agent — Rowan Christmas, Docker at 245 secondsOpen full source frame
The slide contrasts informal safeguards with microVM kernel isolation, PII boundaries, proxy secret injection, policy enforcement and an audit trail.

The protection combines several mechanisms with different jobs:

  • Kernel and filesystem isolation. The microVM runs its own kernel and separates the agent’s filesystem view from the host’s files.
  • Secret placeholders. The sandbox receives a placeholder rather than the real secret. When a network request needs that credential, Docker’s system replaces the placeholder, keeping the secret value out of the sandbox and the agent’s direct view.
  • Audit trail. Activity is recorded so there is a history to inspect, rather than only the agent’s explanation of what it did.

Secret substitution addresses a practical tension: an agent may need an authenticated request to do useful work, while giving it the credential itself creates another thing it can read and misuse. The talk describes the placeholder-and-replacement mechanism, but does not specify its destination checks or substitution rules. Its supported promise is narrower than preventing every harmful authenticated action: the agent need not possess the real secret to make the request.

The balancing act matters. Remove every useful capability and the agent cannot accomplish its task. Leave the whole laptop within reach and the banking experiment becomes possible. The sandbox makes the working environment a deliberate choice: expose the resources needed for the job, then let the agent operate within that environment.

0:123:53
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

3:53 · section reference included

Repeat the browser-history search inside the sandbox

What changes in the everyday workflow? Christmas compares launching claude with launching sbx run claude. The latter creates a new VM, starts a sandbox in the selected folder and runs the agent there. The agent interface looks much the same; permission bypass is enabled in the sandboxed demonstration. The convenience comes from letting the environment constrain access while the agent works with fewer interruptions.

The browser-history example now has an observable before and after. On the host, the agent finds browser history and can continue investigating it. Inside the sandbox, the same search does not even lead the agent to believe a browser is installed. The causal step is the changed filesystem view: creating the VM does not place the Mac’s browser installation and personal history within the agent’s reach. The request can remain ambitious while the accessible files become smaller.

Selected presentation frame from YOLO Mode, Safely: MicroVM Sandboxes for Any Agent — Rowan Christmas, Docker at 362 secondsOpen full source frame
In the sandboxed example, the agent says it does not think a browser is installed.

Why does the same request produce different results? The comparison below follows the search through its two environments. The host path reaches personal history; the sandbox path reaches a filesystem where that browser history is unavailable. The difference lies in what the tool can discover, rather than whether a carefully worded request receives a warning.

Compare the ideasOne browser-history search, two filesystem views

Asked to find browser history.

The sandbox changes the resources available to the search before the agent can use them.

4:475:17
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

4:47 · section reference included

Network policy limits where the agent can connect

Filesystem isolation leaves another route to consider: network access. Christmas asks the agent to look up the Pirate Bay, expecting the destination to be blocked, and the sandbox blocks it by default. This illustrates an independently enforced connection policy. Even if the agent agrees to make a request, the environment can deny the destination.

Selected presentation frame from YOLO Mode, Safely: MicroVM Sandboxes for Any Agent — Rowan Christmas, Docker at 392 secondsOpen full source frame
The Docker Sandboxes terminal shows allowed and blocked network requests.

He also points to blocked telemetry requests that he identifies as Claude traffic to a Datadog instance. The example makes background connections part of the same policy discussion as explicit browsing requests. It does not establish the contents or volume of those telemetry payloads, so the useful observation is that the sandbox can block those outbound connections, rather than a quantified claim about data sent to Anthropic.

These defaults are configurable. That is necessary for a development environment whose useful work may include network requests, and it makes policy choices consequential: allowing more connections increases what the agent can do. Christmas reports that every Docker developer now writes code in sandboxes. His operational preference is to enforce restrictions at the microVM, before agent activity reaches the host, while keeping the workflow easy enough to use every day.

6:146:44
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

6:14 · section reference included

Connect agent actions to authorization and policy

Isolation answers what an agent can reach. Governance adds another question: why was a particular action allowed? Christmas previews agent identity tracking and delegation chains that would connect an action to the agent that performed it and the human who authorized that agent. The intended result is an explanation that follows authorization through the chain, with policies able to restrict permitted actions as circumstances change. The talk sketches this direction without developing the policy-evaluation procedure.

A governance mock-up brings together three controls described as available at the time of the talk:

  • Network allow and deny rules. Decide which connections the environment permits.
  • Filesystem access controls. Decide which filesystem resources the agent can access.
  • MCP catalog selection. Decide which MCP servers are available to the agent. Docker’s MCP servers themselves run in sandboxes, so those servers are also subject to the controls.
Selected presentation frame from YOLO Mode, Safely: MicroVM Sandboxes for Any Agent — Rowan Christmas, Docker at 543 secondsOpen full source frame
A governance slide shows network, filesystem and MCP catalog controls beside a policy interface.

The MCP detail extends the protection to a tool’s execution environment. An agent may invoke a server to do work on its behalf; sandboxing that server gives governance another place to constrain access. Selecting the catalog determines which tools are offered, while governing the server determines what the selected tool can reach when it runs.

Two finer-grained controls remain future work in the recording: L7 networking controls and filesystem permissions organized around individual GitHub repositories, including different areas within a repository. The repository example makes the desired granularity clear: an agent might read and write one area while being excluded from another. Christmas distinguishes those planned additions from the working sandbox demonstrations shown earlier.

7:297:59
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

7:29 · section reference included

Resources

From the talk

Read the complete timestamped transcript
  1. 0:12

    Hey, everybody. I hope you had a good lunch. I know it's one of the last talks of the entire conference. I tried to make it somewhat engaging and amusing. So, you know, audience participation is, is, uh, encouraged. We'll see how it goes. So I'm here with Docker. Uh, I'm one of our product managers that works on our new sandbox product. I don't know if you've seen it. It's a new binary for Docker. It's called sbx. Uh, it's a microVM. It's really cool. You should come check it out. But I'm gonna tell you

  2. 0:42

    why, uh, you might want to use it in context of agents doing governance and j-- being safe in general for your agentic platforming needs.

  3. 0:54

    So right now, you know, all of you are doing way more AI than you ever thought you would five years ago, at least I am. And we generally think that we're being protected by our coding harnesses, 'cause if we try to do something bad, they will sometimes stop us. And when we see stories, we think, "Oh my gosh, I would never make this mistake to have my AI drop my prod database. You know, I've been doing coding for thirty years. There's no way I would possibly do that." So what I d- decided to find out

  4. 1:23

    was, well, what happens if I actually try to hack myself, and how far can I get with, you know, our tools and what we've been told is a safe environment? So I opened up my Claude Code on desktop on my Mac, and I had it look for my browser history. Uh, and it found it right away, which I was very impressed about. And I asked it, uh, "What can I, what can I do with this?" And so we started looking at my bank accounts, which I was, uh,

  5. 1:53

    surprised about. And it found them, uh, which I was very unhappy about. I made fake bank names. This is not my real data, just so you know. But it did find all of my real bank data on there. Uh, and it went even further. It found that I've been ordering checks recently, that I've been using Zelle. It told me the last four digits of the account I was using. Um, I was, I was very impressed. Uh, and here's a kind of a idea of all the PII that it found. Um, this is from,

  6. 2:23

    again, the Claude desktop app, uh, that we thought was safe. Anyways, maybe, maybe there's a better solution. Who knows? So the next day, I came in, and I got from my security team, uh, a nice little notice that I-- they thought I'd been hacked. And so I had to explain to them that, "No, no, no. I'm just putting together a talk for this conference, and, uh, please, please don't flag me as a, as a, a compromised computer." And then they sent me the CrowdStrike report from this, which I asked them to

  7. 2:53

    do, and I was quite impressed with it. Apparently, this is a known way to get, uh, credentials from a machine. Didn't realize that. And I had scored very well on this. I got a nine out of ten, which, uh, is better than I did in most of my classes. So that was, that was really good. But fortunately, this was just me internally. It was not some sort of, you know, injection from a script or MCP server or anything that could easily happen. So five, five prompts is what

  8. 3:23

    it took. Um, I had to be a little clever. I couldn't-- If I just-- If you just ask straight up, like, "Hey, Claude, go find my bank data," it kind of-- it, it will give you a warning. But if you say, "I'm researching how to do security," it'll happily go and help you do that research on your account. It is not, uh, it's not your friend. So if o- if only there was a better way to do this using a sandbox microVM technology

  9. 3:53

    that Docker and others have released recently. So rather than just saying please, which, you know, if you look in some of these Claude prompts, you'll see, "Please don't do nefarious things," right? Like, that is the level of security we're at. So now we've got microVMs, if you're not aware of them, compared to... I know you're probably all aware of Docker, been using it for years. The new thing now that we and others are doing is with microVMs. They run their own kernel. They isolate your file system. We have a system that will go and make sure that the

  10. 4:23

    sandbox never actually sees your secrets. When you try to go and do a network request, it takes a placeholder, replaces it so that your sandbox and your agent can't do bad things. We've got a full audit trail for it. Uh, so yeah, we wanna be, you wanna be secure by design, not just, you know, hope and say please and see what's gonna happen.

  11. 4:47

    We're trying to do this balancing act, right? If we don't act-- If we can't actually do these things, the agent's not useful. So that's why this whole sandbox idea is so exciting for us. We encourage you to use it, um, and maybe, you know, don't let your bank data get into the hands of your agent. That's not, that's not what you want. So this is what it looks like in practice when you're running a sandbox. Uh, on the, uh, left side is when I run Claude, and on the

  12. 5:17

    right side is when I do sbx run claude. And you can see that there, there's a lot of, uh, differences. That's a joke. Sorry, they're not. It's, uh, just the bypass permissions are on, but it's nothing, nothing else, uh, was hard to do, right? Running that command on the right automatically creates a new VM. It spins it up in that folder, and it makes a new sandbox for you. It, i-- runs your agent. Now you're up and running. It can't see anything. So if we try to do our browser history

  13. 5:46

    attack on regular Claude, it happily goes and finds your browser history.

  14. 5:54

    But if I run it on the sandbox version, it does not even think there's a browser installed on your machine. So pretty cool. No difference. I had to type, uh, seven more keystrokes to get here, but I think it's well worth the, uh, the price to pay.

  15. 6:14

    Same thing is true for network egress and ingress. So right here, I'd, uh, I tried to get it to go and look up, uh, the Pirate Bay because I f-- I knew that would be blocked. And sure enough, blocked by default. The other thing you kind of see, it's hard-- maybe it's hard to see, but like the, uh, Claude likes to really send back a lot of telemetry data to their Datadog instance. So if you're using Claude, uh, you are sending Anthropic a lot of your data, unless you're using sandboxes, in which case you're cleverly being blocked

  16. 6:44

    by that. The other thing I'll say about sandboxes, they're so easy to use. Every developer at Docker now writes all of their code in sandboxes. So we are using this every day. We're writing all of our code in it. If we don't use it, we get, uh, we get yelled at. So it's worth our time to, to go and do that.

  17. 7:03

    Little defaults on this. All of this is configurable. Um, and so now that we've talked about like why this matters, right? So doing this at the hardest level doesn't really work. Agents find their way around it. If it gets down to your host machine, it's too late. So you wanna be at that micro VM boundary. We think it's the best way. Other people do too. And it's easy to use.

  18. 7:29

    If you're working in consulting, if you've got a, you know, chief security officer, if you do client work, whatever it might be, you're probably dealing this-- with this on a day-to-day basis. So not only do you want this to be useful, but also you want to do things and enable your agents to do different things, right? So one of the things we're really excited about is doing like agent-level identity tracking and delegation chains. So when you can start saying, "Hey, how come this thing happened?" And you can go back and say, "Oh,

  19. 7:59

    that's because an agent did it, and a human actually authorized that agent to do it." It didn't just happen by magic. We didn't know. We don't-- It's not like we don't know how it would happened. Now we're gonna start tracing that, tell you f-- tell you about it, and let you write policy that actually goes and s-- does degradation of, uh, what you can do based on Cedar policies, based on new, new things happening. And you might say to yourself, "Oh my gosh, this is so good. I would like, I would like to buy it now.

  20. 8:30

    How do I, how do I buy this from you guys?" And so I have my marketing slide that my marketing team is really mad at me that I made this, but I think it's funny and cute for AI governance. So definitely talk about that. Okay. What does AI governance look like? I-- You're doing-- This is what, this is what it looks like in a kind of a mock-up. So right now we've do, we do network. You can set allow, deny. You can do file system, uh, points. And you can also decide what your MCP catalog looks like.

  21. 9:00

    Also, if you're running our MCP server, our MCP servers themselves run on a sandbox, so they are also governed by all of these different controls. In the future, we're gonna be adding more stuff. One of the things we've heard from a lot of people at this conference is that they wanna do, uh, L7 networking controls. They wanna do, uh, like per GitHub repo file system level controls, right? Like you want your agent to be able to read and write to some

  22. 9:30

    repos and some areas of some repos, but not others, right? So all of this is gonna be coming down the line, but everything that I've showed you today is working. Um, and you should come by our booth and see that.

  23. 9:46

    It's easy to install. We run on Mac, Windows, and Linux, uh, based on your, uh, package manager of choice. And you can see at the end, it's just sbx run cla-- It's not just Claude, by the way. The, the sandboxes are a full VM. So you can run a shell, you can run Codex, you can run whatever. You can run a Python job. You can run a web server. Any- anything you wanna do. But we wanna make it easy for you to get in and up and running. So this is kind of our default. So you can just

  24. 10:16

    get in. You can get that Claude Code or Codex instance going. It feels native. It works just the same way as your one did today, but you've got all these protections in store. Uh, additionally, you can mount other file systems. So what I do is when I'm working on a, you know, a piece of code with related repositories, I'll mount those, uh, read-only. That way, my agent can see what I'm doing. It can see my other code base, but I know it's not gonna make random commits to other

  25. 10:46

    repos just to make my thing work, right? I want the guarantee that it's actually doing what it says it's doing with the API that I've specified.

  26. 10:55

    So that's it. Pretty short. We have a booth. We have a few sunglasses left and some power banks to give away. You should come by our booth, and I'll walk over there afterwards and give you the demo, and then you can look as good as, uh, Macho Man Randy Savage with those glasses. So there you go. Thank you.