AI Engineer World's Fair 2026
How Many Credentials Should Your AI Agent Have? Zero. — Jim Clark, Docker
Read the talk
How Many Credentials Should Your AI Agent Have? Zero.
Jim Clark explains how task-specific sandboxes, a shared MCP gateway and existing enterprise identity systems can let agents work longer while limiting what each stage can reach.
From a talk by Jim Clark
At a glance
Ideas worth remembering
Design each sandbox around the current task's information and tools. The complete workflow can have more capabilities than any individual stage.
Separate open-web research from publishing power, and separate ordinary coding from credential-backed signing operations.
A single MCP gateway endpoint centralizes capability configuration and makes the harness easier to replace.
Zero credentials in the agent's sandbox is the design goal; XAA provides a described route from agent identity to existing enterprise authorization.
Orchestrators should choose the harness, MCP capabilities, network rules and working resources for each task, rather than expose the entire workflow's tool collection.
Longer runs change what supervision can do
An agent that works for an hour changes the safety problem. Sitting at a laptop approving every action defeats much of the value of handing it a substantial task. Jim Clark, principal software engineer at Docker, starts from that tension: needing little supervision is a sign that a task is well designed, but leaving an agent alone also raises the question of what it might do.
MCP gateways enter the discussion as a way to control that exposure. The explanation separates three things that often travel together: the harness that runs the agent, the tools and context it receives through MCP, and the sandbox where it performs work. The useful safety question is what that particular task needs access to while it runs.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The harness runs the loop; its capabilities determine the reach
At the level needed for this discussion, a harness takes context and produces tool calls. That is a simplification, but it helps explain why replacing one harness with another should be routine. Harnesses evolve quickly, and their interaction styles make them worth switching. Without information or tools, however, the loop has little useful work to do.
MCP supplies ways to bring in context and invoke tools. Those additions make the agent useful, and they also determine what an incorrect action can affect. Before building the environment, ask two practical questions: what information does this task require, and what tools must it use? Giving it only those resources limits the blast radius of a mistake.
A sandbox is the place where the agent works. Early coding-agent setups might effectively expose an entire laptop: the repository, local tools and everything else available there. A smaller sandbox narrows that environment to the task. Clark's emphasis is on controlling the tools and context flowing into the harness; the simplicity of the harness loop alone does not establish that those inputs and capabilities are safe.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A newsroom separates open research from publishing
The newsroom example develops that separation into a complete workflow. Research goes out into the world, fact-checking examines what comes back, and a final stage writes or publishes the story. The workflow needs all of those abilities, but each role receives a different environment.
- Researcher: Broad web access lets it gather information. It lacks the MCP tools that could publish to the corporate site, and writes its findings to private local storage, such as a temporary directory.
- Fact-checker: A separate sandbox reads those findings and checks them against a fact database through its allowed MCP capabilities. The example gives it no general network access.
- Publishing stage: A final sandbox receives filtered research and a Notion MCP or Publisher MCP. It has no need to browse the open web.
Follow the research through those handoffs. First, material collected from the internet becomes a local research artifact. Next, the fact-checker reads and filters it using the fact database. Only then does the publishing stage receive the filtered material together with a publishing tool. The observable change is in both the input and the available action: the stage that sees arbitrary web content cannot publish, while the stage that can publish receives the output of the checking step.
Where do information and publishing power meet? The diagram traces the handoffs rather than treating the whole newsroom as one agent with one tool list. Across the workflow, internet access, fact-checking and publishing all exist. Inside any individual stage, the intended combination is narrower. This is a separation design, not a guarantee that filtered research is harmless: its protection still depends on what the checking step passes onward.
Broad external information available to the researcher.
The open-web stage produces local research. A separate stage checks it before the publishing-capable stage receives it.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
An hour of coding does not require an hour of signing access
The second example applies the same reasoning to time rather than organizational roles. Clark describes sending a coding agent on long tasks and wanting it to commit along the way. In his illustrative hour-long run, Git signing might be needed for only two minutes. Leaving signing keys available throughout the hour would give the coding phase a capability it rarely needs.
Making a commit calls for a much smaller set of resources: the Git commit tree, the ability to create the commit and signing access. Other capabilities can be left out of that stage. Conversely, a task that is not making a commit should not have signing keys in its sandbox. The task's intent becomes a specification for its capabilities: coding and signing are different operations even when they belong to one larger job.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
One MCP endpoint makes policy independent of the harness
Docker's gateway design puts a harness and a single MCP gateway endpoint into the sandbox. Resources and tool calls pass through that endpoint. Clark sketches an internal hostname to convey the arrangement; the important property is that the harness has one route for MCP traffic.
The gateway provides a control point for which tools, resources and prompts each sandbox receives. Sandbox configuration determines the allowed capabilities, while the harness connects to the common endpoint. That separates two decisions that otherwise get tangled together: which agent interface to run, and which MCP capabilities the task may use.
The practical benefit is less per-harness configuration. Switching between Codex and Claude Code need not mean rebuilding the same MCP setup in each one's format. The harness becomes a parameter of the sandbox alongside its MCP selection and network rules. As Clark puts it, “Bob's your uncle.” The simplicity comes from moving the shared configuration to one control point.
Credential handling belongs at that control point too. The headline's answer—zero credentials in the sandbox—is Clark's design maxim. Read together with the signing example, it distinguishes needing a credential-backed operation from keeping the credential available to the agent. Removing credentials reduces what a mistaken agent can expose, although permitted tools can still perform consequential actions.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Cross App Access connects agent identity to existing SSO
Cross App Access (XAA) extends the credential discussion to enterprise services. Clark describes Docker, Anthropic and Okta partnering on this approach. Companies already have identity providers, SSO and OAuth-backed services such as Notion, Atlassian, GitHub and Slack. The missing connection is allowing an agent to act through that existing identity setup.
The described exchange starts by identifying the agent and the person on whose behalf it acts. Identity claims are exchanged with the existing identity provider, which returns an authorization grant called an ID-JAG. That grant supports access through the authorization servers the company already uses. This is the conceptual flow presented in the talk; it does not specify token storage, lifetimes or the complete exchange protocol.
How does existing SSO become useful to an agent? The diagram makes the identity provider an explicit participant between the agent's identity and service authorization. The intended administrative benefit is centralized control over what agents may do with existing resource servers, instead of repeatedly asking users to approve separate consent screens.
Identify the agent and the user it acts on behalf of.
The conceptual XAA flow reuses the company's identity provider and authorization servers.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The orchestrator builds an environment for each task
Progressive disclosure applies to capabilities as well as instructions. A workflow might use fifty different MCPs over its lifetime; that does not mean every sandbox needs all fifty. Exposing only the tools and resources needed at the current stage also keeps the agent's context smaller. The aim is to leave room for problem-solving inside an environment that expresses the task's intent.
That gives the orchestrator a concrete job: inspect the task, decide what it needs and construct the sandbox before assigning the work. Its choices include several separate dimensions:
- Harness: Select the agent implementation that will run the task.
- MCP capabilities: Select the tools and resources that this stage may use.
- Network rules: Decide whether it needs a specific service such as
api.github.com, broad web browsing or no general network access. - Working resources: Provide the relevant files and worktrees.
The resulting loop has less access than the workflow as a whole. Role separation supplies the organization; containers supply separate places to work; progressive disclosure supplies capabilities as they become necessary. The sandbox is where those decisions become the environment the agent actually receives.
The closing practical step is the sbx CLI. Clark presents brew install sbx as the way to try the tool discussed in the recording and points to further demonstrations of cloud sandboxes and orchestrators. For implementation, the supplied Docker Sandboxes documentation covers local and Docker-managed cloud environments, including MCP configuration and security topics. The CLI reference provides the command reference for creating and managing those environments.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
From the talk
Start with local or cloud sandbox setup, then explore the documented MCP, architecture and security topics relevant to task-specific agent environments.
Command reference for the sandbox CLI offered as the practical next step at the end of the talk.
Read the complete timestamped transcript
- 0:13
I guess I'm ready to start. Uh, can everyone can... I, I actually, I can hear myself. Yeah. So you guys can hear me. Uh, my name's Jim. Uh, I'm a engineer at Docker. Normally, when I introduce myself in the slides these days, or at least what I used to do, was I would say, "Hey, I'm Jim. I, uh, like spaces, not tabs. I use Neovim, not Emacs." But that's, uh, all irrelevant now.
- 0:42
So, uh, I will say that I made this slide with, uh, I made this presentation with Claude Code. I used the Codex model. The MCPs that I, that helped me write these slides were Marp and, uh, Mermaid. And I'm gonna talk to you about MCP gateways, but I'm really gonna be talking to you about AI safety. So I'm gonna kinda go through sandboxes, gateways, harnesses, and sort of motivate why we're talking
- 1:12
about AI safety at Docker right now. So, like, probably the last six months has been similar for all of you as it's been for me. Agents are doing way more than I thought they were doing, uh, than I thought they were going to do, and I'm no longer going to sit in front of my laptop and say, "Yes, yes, yes, yes." They're doing longer running things, and I like that. And as I get more and more
- 1:42
accustomed to them doing longer running things, I sort of decide that my metric for me having designed a problem that an agent can sink its teeth into is that I don't really need to give it a lot of supervision. But if I'm not gonna supervise it, maybe I'm also a little bit worried about what it's going to get up to. So I think just to set context, this is what I think is sort- the sorta shape of the problem right now. We need to talk
- 2:12
about agent harnesses, what the, what the agent harnesses do for us. We need to talk about sandboxes. Like, when a agent harness delivers work, what is the, what is the thing it delivers work into? And then of course, we've got MCPs, which, which everybody knows, and we'll talk a little bit more about what we're doing there. But let's start with the agent harness. So,
- 2:38
uh, I mean, it's, it would be an oversimplification to say that an agent harness is something which just takes in context and emits tool calls. But there's also a little bit of truth to that. It's not the most complex part of it. It's a loop that takes in some context that you give it and asks you to do something on its behalf. But belying that simplicity, we get we're all using a lot of different harnesses. And the inter- the
- 3:09
ability to... Why, why do we pick up different harnesses all the time? Like, we love the new sort of interactio- interactive possibilities that they bring us. So let's just admit right from the beginning, we're gonna be using a lot of harnesses. They're gonna evolve pretty quickly. We're gonna be swapping in different harnesses for one another. But a harness by itself is just kinda in a little bit of a dead room. Like, uh, "Hello. Someone give me some context. Some- somebody give me something to do." This is
- 3:38
where, this is of course where MCPs come in. MCPs allow us to do new things. They allow us to go ou- outside of ourselves, pull in new context, pull in new tools, and, and get actual work done. So I think it's, when you're trying to think about safety, I don't think it's an oversimplification to say that one of the things that you need to do is figure out what are the resources that you need to do your job.
- 4:08
Like, uh, any time, even for us, one of the first things you'd analyze when you wanna do a new task is, "What do I need to get this task done? What information do I need? What tools do I need to actually get this done?" And if you knew that the sandbox that you built had the right resources, had the right tools, and only the r- those resources and tools, it would make you feel safer. It would make you feel like you limited the blast radius for what could go wrong when that agent is working.
- 4:39
So this is kinda where sandboxes come in. Sandbox is a place for an agent to do w- to do work.
- 4:47
So our typical sandboxes today,
- 4:52
you're, most of us are accustomed to the, uh, the, yeah, we have code sandboxes. We give it, we give it a bunch of code. We give it all the tools on our laptop. Maybe, like, when you first started using some of these agents, the sandbox was your entire laptop. As over time, we're t- starting to learn how to get that a bit smaller. We're trying to make the sandbox boundaries small enough that they can actually just do the task that we want to do, want them to execute. That helps us feel
- 5:21
safer about letting the agent run for longer periods of time if this, if this sandbox is, i- is a little, is, is, is smaller. And, you know, it's not that we're trying to h- sandbox the harness itself. The harness is kind of already sandboxed. Like, it's a, it's a pretty simple thing. We're trying to harness the tools and the context that are flowing into this harness. So in order to kinda illustrate some of these concepts, I came up with
- 5:51
two stories that I'm, I'm just gonna walk you through. The first one is, um, a newsroom analogy. So if you've got a newspaper reporter, they go out in the world, they find cool stories, find cool information, they bring it back. Maybe someone inside of that agent- agency, uh, look, does some fact-checking on that, makes sure the information is right, and then you write a story. It's totally natural that those roles are completely
- 6:21
separate. The actual reporter is not gonna be allowed to, like, publish to the website or, or, or write, write the actual article. We, we break those up. That's actually how we already build or- complicated systems. We split them up into little pieces. So on this slide what you see is I've broken down three sandboxes in blue here. So we've got a researcher, we've got a fact-checker, and we've got a reporter. So that researcher sandbox, I think, I think about that as the
- 6:51
reporter. Now, when we define that sandbox, let it access the web. Let it access as much of the, as much of the internet as we want, because we're not gonna give it the MCPs that it can, like, write to our, our corporate site. It can't publish anything. We're just gonna let it do... We're just gonna let it go out, go crazy, do research. So our sandbox is like, "Yeah, not too many tools. I'm not gonna give you too much write access, but definitely do lots of research.
- 7:21
Just write your research out to a, I don't know, to some temp directory, some private thing that is local to this agent." A fact-checker is gonna need some MCPs, but it's not gonna need any network. So then let's start the fact-checker up. Let's put that in a different sandbox. Let's say no network. Y- and you can read the research that's been done for you and you can fact-check it. And you can have access to our, our fact database. You know, filter that out. And then finally we get to the little blue
- 7:51
thing on the right, and this certainly doesn't need any, any access to networks. In fact, it shouldn't have any. But we do wanna give it our Notion MCP or our Publisher MCP because it's just looking at filtered research. So you know, look at this whole slide. Draw a box around this entire slide, and that effective agent has a bunch of MCPs. Maybe it has an MCP for publishing, it has an MCP for fact-checking. At one point in time it has
- 8:21
access to the entire internet. But we break it up into logical sandboxes, so at each individual point we have a, don't have a dangerous conce- or a dangerous combination of both being able to read in crazy con- unfiltered context and do things with our tools. So let's look at example number two, and I'm just gonna talk about, uh, building a... Actually, this is how I, I build my coding
- 8:51
agent right now. I tend to send the, the coding agent off to do pretty long-running tasks. And let, let's just say that they work for about an hour. But as the, as the work is happening, I also like it if it just does commits along the way. But the ratio of, the amount of time that it's, my agent is actually committing is almost nothing. In a, in an hour it might be, it might need, say, my s- my c- my Git
- 9:21
signing keys for like two minutes of that time. And when I'm committing, I need to see my, my Git commit tree. I need to make that commit. I need some signing keys. But I don't need anything else. So I would like it if the actual sandbox that an agent is running a, a task in is modeled after the intent of what I'm doing. So if I'm not gonna make a commit, I don't want my commit signatures, my commit signing keys inside that
- 9:51
sandbox. And I think this, this you can map across a lot of different tasks. If you know the int- if your agents know the intent of what you're trying to do, then use that intent to model the capabilities that you give to that. So this is driving in, in our, our Docker sandbox product, this is dr- defining a lot of how we're thinking about exposing things like an MCP gateway. So when you build a sandbox, of course you
- 10:21
have to put a harness in, in that to do things for you. And then we're just putting a little gateway endpoint into that, and that gateway endpoint funnels all MCP traffic. So imagine it's something like MCP, an, an internal, an in- an internal URL, mcp-gateway.docker.internal or something like this, that every single agent harness has, and all traffic, all resources that you need to
- 10:51
pull in, all tool calls move through this one, this one single gateway. So what does this buy us? Well, we end up with a control point. So the gateway manages which tools, resources, and prompts are given to each sandbox. So you're starting to see the beginnings of your ability to say, "Okay, I'm giving you a harness. I'm giving you this gateway endpoint." The gateway, the s- the configuration of the sandbox controls what you can do inside of
- 11:21
that sandbox with things like MCP. Another nice thing is suddenly all your harnesses are MCP agnostic. So you don't have to go to Codex and configure your MCPs one way and go to, uh, Claude Code and configure your MCPs one way. You just have one-- You ma- almost make the harness a parameter of the sandbox. Stick a harness in, stick the MCPs you need in, give it some network rules. Bob's your uncle. So it's, it feels,
- 11:50
the MCp- MCP is, is a good place to manage things like credentials. In fact, I would say that a maxim that you can use here is how many credentials should be in, uh, a sandbox? Well, I mean, the, the, the right answer is always zero. It just, they should just never be in there. And the blast radius of an agent going, doing something, going, going, uh, doing something incorrect, if there are absolutely no, no
- 12:20
credentials in there, is, is reduced. So one of the things that, um, Docker, Anthropic, Okta are, are all partnering on is a, is a new thing called XAA, Cross App Access. And I think this is a really great illustration of why sandboxes and, um, and managing credentials in this way is important. So today in your, uh, your, you're working at, uh, at, at your companies, you have some sort of identity management, you have
- 12:50
SSO set up, and you've got all these resource servers that, that manage, uh, that, that you work with for OAuth. That could be Notion or it could be Atlassian or GitHub or Slack. They're-- you've already configured SSO for those, but it's not yet accessible to your agent. So with the Cross App Access scenario, your harness, your a- your, your sandbox and your gateway
- 13:20
can define things like agent identity, or who are you that this agent is behaving on behalf of. And by extension, by exchanging identity claims with your identity provider, which is already there, you can get back a new thing, which is a new part of this spec that you, you'll now see in the latest version of MCP called an ID-JAG, an authorization grant. And with that authorization grant, we're granting access to this
- 13:50
act, to th- to this actor, to this agent, to the same authorization servers that you've been using. There's nothing new here. We're not-- We're leveraging all this existing investment that these corporate internet, cor- corporate networks already have. But suddenly, without any crazy consent screens, "Yeah, you can do that. Yeah, you can do that. Yeah, you can do that," we can centralize the administration of what an agent now does with all these existing resource servers. It's a great... It's a huge
- 14:20
simplification.
- 14:23
So the MCP gateways, I like to think, are-- we're starting to talk about progressive disclosure. Well, it's a term we learned from s- from Skills. I like to think we're starting to use this now we're able to progressively disclose tools and resources and MCPs into these sandboxes using very, very similar principles. And of course, it's amazing because we keep context size down. Just because you might, at some point in a workflow, use fifty different MCPs, doesn't mean that each sandbox has to
- 14:53
have all fifty of those MCPs. Give your agents room to solve problems, but let individual sandboxes represent the intent of the task. Let them represent what you're actually trying to do in this task.
- 15:09
So I like this picture because suddenly we start to think about orchestrators as one of their jobs is to go, "Well, what am I doing? What's my task?" And I need to put a task into a sandbox. So what sandbox do I put it in? What capabilities does this need? And of course, it needs a harness. Like what do you-- put in OpenCode, put in Claude, put in Gemini, put in Codex. It's gonna need some MCPs. Take a look at the task, decide what MCPs are allowed to be
- 15:39
used in this. Um, what networking rules should you apply? Does this need access to api.github.com? Does it need access to surf the web in, in, in random ways, or does it just need nothing? What resources? What work trees? We get used to thinking about modeling capabilities, building the right sandbox for the task, putting the task in, and now we're starting to see that this individual loop is safer because it has less
- 16:09
access than the entire workflow e- effectively has. So this is, um, this is about containers. This is about containerizing agents, which is why, uh, I guess is why we're at, I'm, I'm at Docker. But this ties together as a kind of like role separation. What are the roles that you, that your agents have? How do you map them and containerize them and, and make sure that you're giving them sandboxes that express that intent? Never have any, any cr-
- 16:38
creds anywhere in any, in any harnesses. Disclose capabilities progressively into agents as they need them. The sandboxes for you are, are, are what actually represent intent. And if we allow some of these things to run longer, but they're more sandboxed, we feel a little bit safer about that. That's, this is where safety comes from. So we have a demo. Our, our EVP of, of Engineering, Tushar Jain, is, uh,
- 17:08
got a talk upstairs today, I think it's at four o'clock, and he's taking a lot of these ideas and just demo, demo, demo. How do you m- move sandboxes to the cloud? How do you build orchestrators? He'll be really showing a lot of these things live. We also have a booth over here. Um, it's a great opportunity to come over. If you're interested in any of this stuff, we can show you the command line is available today. Anyone, uh, just brew install sbx. That-- Everything I've been talking about today is this tool, sbx.
- 17:38
This is, this is how we build, we build containers. Um, so yeah, please, please, uh, please come over and talk to us at the booth. We'd, uh, we'd love to hear what you're, uh, what you're doing with AI safety and how we might be able to help. But, uh, thank you very much.