AI Engineer World's Fair 2026

How I Learned to Stop Worrying and Love the Sandbox — Matt Brockman, E2B

Read the talk

How I Learned to Stop Worrying and Love the Sandbox

Matt Brockman’s E2B workshop turns broken environments into lessons in resource debugging, snapshots, workspace identity and lifecycle management—and shows why restoring state also restores the mess you left behind.

From a talk by Matt Brockman

At a glance

Ideas worth remembering

  • Snapshots preserve RAM, files and running processes. That makes resumption convenient, but also carries forgotten background work into later sessions.

  • Separate task runtime from idle lifetime: allow work to finish, then shorten the wait before releasing running resources.

  • Stable workspaces require recording the user-to-sandbox mapping. Returning a new ID without storing it cannot make the next visit consistent.

  • Resource cleanup and access control remain application responsibilities: track abandoned environments, direct writes to writable locations, and restrict outbound access when secrets are present.

Give agent code somewhere safe to make a mess

An agent with permission to run code can consume CPU, spawn processes, delete files and read environment variables. Moving that agent from a laptop to a server carries those problems into an environment that may also hold secrets. Matt Brockman, an engineer at E2B, starts with this practical reason for sandboxes: give user-specific code an environment where its actions do not interfere with other users.

Selected presentation frame from How I Learned to Stop Worrying and Love the Sandbox — Matt Brockman, E2B at 196 secondsOpen full source frame
The slide lists isolation for user-specific code, agent exploration, unsafe actions, and provisioning as sandbox concerns.

The workshop makes those failures deliberate. A capture the flag lab presents a sequence of broken environments; solving one challenge produces a flag that unlocks the next. The early exercises require Linux inspection and configuration rather than application code. Later exercises simulate management problems that become harder with hundreds or thousands of sandboxes.

E2B’s approach combines quick creation with restoration of existing state. Brockman gives a startup target of less than 100 milliseconds; this is an aim, rather than a measured result established by the workshop. The state model resembles closing and reopening a laptop: snapshot memory and the filesystem, save them, then resume where work stopped. An agent returning from another task can find the processes and files it expected to leave behind.

That changes what a template contains. In Brockman’s comparison, a Docker image supplies the built environment, then its start command launches the process. An E2B template captures a VM after processes have already started, preserving disk and RAM so a new sandbox can resume that prepared running state. The distinction will make the lab easy to reset—and later explain why forgotten background processes survive.

0:242:51
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:24 · section reference included

Find the sandbox, then separate the experiments

The first challenge asks for the current sandbox’s ID. The dashboard already connects to a sandbox running a web server for the lab, and the ID can be copied from its URL. Setup takes longer than the challenge because the room’s Wi-Fi struggles; participants pair up where necessary. Running the code remotely does not eliminate the need for a working connection to its terminal and web interface.

The URL encodes two useful pieces: the service port and the sandbox ID, in the form port-sandboxID.e2b.app. The port identifies a listening service; the ID identifies the environment to reconnect to. Keeping that ID is the first step toward returning to a particular workspace rather than simply creating another machine.

Each subsequent challenge gets a separate sandbox. A mistake in the CPU exercise can therefore leave the initial lab intact. The browser terminal connects through a PTY—a pseudoterminal session—and the prepared template supplies the challenge’s running processes. Participants receive separate copies of the prepared VM state, rather than all modifying one shared environment.

The same mechanism can fork useful work. A data scientist with a populated dataframe could copy that state into five sandboxes and run five different mutations in parallel. A template ID identifies the starting state to copy; a sandbox ID identifies one resulting instance. An agent can also connect programmatically using an API key and the instance ID, though Brockman advises reviewing its work: “Do what I say, not what I do.”

9:059:35
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

8:20 · section reference included

Inspect CPU, memory and disk before changing anything

The CPU challenge supplies a concrete inspection-and-repair loop. A process listing shows one process consuming 99.7% CPU, alongside its process ID and memory usage. Killing the identified process removes the runaway task; the lab then produces a flag and unlocks level three. The useful sequence is to locate the resource consumer, identify the process responsible and act on that process.

The next failures involve different resources, so their remedies differ:

  • Memory pressure. Inspect memory use and terminate the process consuming it. Adding resources has fleet consequences: CPU and RAM quotas limit how many instances can run at once, so a larger allocation to one sandbox reduces capacity available elsewhere.
  • Disk pressure. Find the file taking up space, then remove it through the challenge’s cleanup flow. Terminating a process and deleting its accumulated output are different operations; the disk exercise asks for a path rather than a process ID.

The disk challenge becomes a better demonstration when the first attempt goes wrong. Brockman finds the offending file and deletes it directly, then discovers that the lab expected him to submit its path before clicking cleanup. Removing the file bypasses the expected completion sequence. “I broke the demo. That’s okay, we can just start a new one.” A fresh copy restores the challenge’s original state.

The corrected flow finds a temporary disk-fill file occupying 512 MB, submits the discovered path to the checker, and lets the cleanup button run deletion and issue the flag. When a participant cannot find the expected file, the remedy is again to reload the snapshot into a new copy. This reset is distinct from resuming the participant’s modified environment: it returns to the prepared starting state.

The browser is a convenient debugging view, not a requirement for operating the sandbox. Most interactions can happen programmatically. These first exercises establish what to inspect when arbitrary code consumes resources; the next exercises ask how to avoid wasting resources even when the code behaves correctly.

18:3419:04
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

18:34 · section reference included

Give the task time; shorten the wait afterward

Giving every visitor a sandbox for 20 minutes sounds simple until thousands of visitors arrive, try one command or never interact at all. Their idle environments occupy capacity that other users could use. Yet a uniformly short lifetime creates the opposite failure: data science, scraping or video processing may need substantially longer to finish.

The lab separates active work from the wait for another command. Its simulated policy first allows an hour for the task, runs that task, then reduces the remaining lifetime to one minute. The hour is an allowed runtime, not a requirement to keep the machine busy for an hour. Task completion is the event that switches the policy to a short idle window.

Where does the lifetime change belong? The flow below places it after completion, so a long task retains its execution window while an unused sandbox stops occupying compute shortly afterward. Pausing preserves state for a later return; shutting down frees the running resources. Brockman discusses both choices, so the exercise’s central lesson is the timing of cleanup rather than a single mandatory retention policy.

A running sandbox responds immediately; a paused one must resume, but does not occupy the same active CPU capacity. That tradeoff matters when an agent uses several environments and its next visit is unpredictable. Persistent daemons and always-on web servers are a different fit: they can run in a sandbox, including for a day, but Brockman points toward potentially cheaper traditional hosting for long-running services.

Templates can also shorten development setup. Install base dependencies once, preserve that prepared environment, then fork it for separate work. This carries forward expensive preparation in much the same spirit as a dependency cache in GitHub Actions, while leaving each copy room to develop independently.

How it fits togetherChange lifetime when the task completes

Allow an hour for execution in the simulated exercise.

The simulated exercise permits an hour of task runtime, then reduces the idle window to one minute. Cleanup follows inactivity rather than interrupting the task.

31:4132:11
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

31:11 · section reference included

Restored state includes forgotten processes

Resumption preserves unfinished work, including work nobody intended to keep. A process starts, its user moves on, and later sessions accumulate more background processes. Returning to a snapshot restores those processes alongside everything useful. This is the downside of finding things exactly where you left them.

The exercise begins with three leaked workers. A search finds them by a deliberately convenient name; after one is killed, the count drops to two. Removing the remaining workers clears the challenge and produces the next flag. Real leaked processes are rarely labeled so helpfully, so the lab simplifies identification while retaining the inspect, terminate and recheck loop.

Cleanup operates at two scales. Inside a sandbox, remove processes that no longer serve a task. Across the fleet, track the sandboxes themselves so environments that fall out of application management can be found and removed. Losing an instance ID can turn a whole sandbox into an abandoned resource.

39:0539:35
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

39:05 · section reference included

Remember which workspace belongs to which user

The first Python exercise changes a router that hands out sandboxes in round robin order. That policy distributes assignments without remembering which environment a user previously received. A returning user needs a stable workspace, so the router must instead keep a mapping from user ID to sandbox ID.

The assignment function has a small but consequential sequence: check whether the user already has an entry; if so, return the saved sandbox ID. Otherwise obtain a sandbox ID, store it under that user and return it. The lab uses a dictionary and mocked sandbox operations to exercise this behavior without creating an unnecessary fleet.

The live edit initially returns the ID without saving the assignment. The test fails, and an audience member identifies the missing write. Adding the sandbox ID to the dictionary makes the test pass. The causal difference is persistence between calls: returning an ID serves the current request; recording it lets the next request find the same workspace.

For an application, Brockman recommends a database that tracks users and their active sandboxes. Persistent records also make assignment history inspectable: excessive reassignment can be investigated rather than disappearing with an in-memory dictionary. The lab demonstrates the mapping mechanism, while the database recommendation addresses keeping that knowledge beyond one process.

42:2942:59
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

42:29 · section reference included

Send runtime writes to a writable cache

The filesystem exercise separates what an agent can read from where it should write. A template cache contains information prepared while building the template and may be read-only. A runtime cache supplies a separate destination for writes made during execution. The fix changes the cache-store configuration to point at the runtime cache.

This is a configuration decision with an operational consequence: code that needs to update its cache must target a writable location, even if it can read useful files elsewhere. Brockman mentions describing intended locations in AGENTS.md, but the demonstrated change hard-codes the cache destination. The exercise does not implement a general filesystem permission system.

48:3549:02
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

48:35 · section reference included

At fleet scale, storage, coordination and network access remain

The closing discussion moves beyond the lab. One participant wants to distribute a locally prepared environment across GPUs from different providers and bring the whole sandbox back afterward. E2B did not support GPUs at the time described in the recording. Sending results home in a POST request is offered as a common output-return pattern, but it does not establish a solution for transferring the participant’s complete GPU execution state; Brockman leaves that question unresolved.

Fleet operations introduce problems beyond a single runaway process:

  • Logging and assignment. Keeping track of many environments creates management overhead even when each environment works correctly.
  • Network capacity. Hosts can communicate with only so many sandboxes at a time, making connections another resource to manage.
  • Stored state. Creating and retaining many sandboxes raises pruning, compression and storage-growth questions. Brockman describes these as active engineering challenges rather than presenting a finished retention policy.
  • Shared files. Volumes provide a place for resources that multiple environments can share, though that capability alone does not answer every question about accumulated snapshots.

The execution environment and the agent can be placed separately. Claude Code or Codex can run inside a sandbox, or a local agent can talk to a remote sandbox. Underneath, Brockman describes E2B as Firecracker-based and running Nomad, with Kubernetes integration still facing technical issues at that point. These are the deployment choices described in the recording, rather than requirements imposed on every sandbox system.

An orchestrator can also use one sandbox to create others for a team of agents. The workshop already has a simpler version of that arrangement: the participant acts as coordinator and clicks to start a challenge from its template. Automated orchestration was omitted to avoid API-key setup and the risk of displaying a key during the recording. The same creation mechanism can be driven programmatically.

Whether templates develop into a Docker-like ecosystem remains an open question. Their immediate value is more concrete: one prepared running environment can be reused by many people. Local development also remains useful inside E2B, especially for testing, where cloud execution adds friction.

The final question returns to the risk that motivated the workshop: what can code send out? E2B sandboxes offer configurable network rules. If an environment can read secrets, its outbound access deserves corresponding restriction. A sandbox gives code a separate place to run; the application still has to decide which files, credentials and network destinations belong within its reach.

51:4251:55
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

50:24 · section reference included

Read the complete timestamped transcript
  1. 0:24

    Hey, y'all. Okay, cool. All right, so I see people still trickling in, but, uh, we're gonna go ahead and get started. Uh, so nice to meet y'all. Uh, I'm Matt. I'm an engineer from E2B. We do sandboxes. Uh, so before we get started, uh, how are y'all doing this morning?

  2. 0:39

    Good. Good.

  3. 0:41

    All right. So my hearing is, like, really messed up. Uh, I couldn't hear that. So how we doing this morning?

  4. 0:46

    Good.

  5. 0:46

    All right, great. All right, so everyone's alive. Uh, and, uh, also where are y'all from? So how many people here are from San Francisco, uh, based out of San Francisco? Okay, we got a couple. Uh, coming out of state? All right. A lot of people coming from out of state, great. Uh, also, uh, we're gonna be, uh, talking sandboxes and then doing, uh, basically a capture the flag. Uh, I don't know if y'all have done a capture the flag before, but, uh, what it's gonna be is it's a series of different levels that we're gonna go through, uh, solve it, and that unlocks the next level. Um, uh, as far as programming

  6. 1:16

    languages that we use, uh, how many people here use Python? Okay, most of us use Python. Good. Uh, so our SDK is in both Python and JavaScript. Uh, good news also is that we don't need to code at all for the first couple levels. A lot of these are gonna be configs, kind of looking at how sandboxes work. Uh, JavaScript? Okay. Uh, Go? Where's my Go engineers? Great. We are hiring, by the way, uh, for people that, uh, uh, code in Go, uh, for the back end. Uh, so, uh, c-come check out our page after this or during this.

  7. 1:46

    Um, let's see. And then how many people here don't code at all?

  8. 1:51

    All right, great. All right, so let's get started. Uh, oops. So, all right, I don't know if people can go to this link, but if you go to this, uh, tinyurl.com/, y'all can see that. If you go there, uh, that'll give you a sandbox. Uh, so you have to log into E2B. Um, but what this workshop is, is everything is basically sandboxes all the way down. Uh, so what we put together for you is, uh, a bunch of different challenges that are issues that you run into ru-running a s- a sandbox,

  9. 2:20

    um, we're about visiting in a second. Uh, and then also you can just do the QR code. Uh, everyone able to get to that page? Thumbs up, thumbs down. Cool. All right. Uh, and if you can't, don't worry about it, we'll be walking through it. So at this point, feel free to, uh, start going through the instructions, figuring out how to hack away. Uh, or if you actually wanna listen to me drone on about sandboxes, you can do that instead, and then go, uh, uh, do the sandbox, uh, start the CTF, uh, with every- with me.

  10. 2:51

    Okay, so just a little bit about E2B. Uh, so what are sandboxes? Uh, sandboxes are, uh, basically let you run user-specific code, uh, without messing with other users, right? So a lot of the time, uh, with your traditional web stack, you'd be like, "Hey, here's a web page. It's a static web page. You can go do stuff." Uh, you know, maybe you had users that could code and what they wanted to be able to do was use something like Colab, uh, where you go in there and you can, you know, play around. Uh, you might wanna have like AWS EC2, right? Get into that, run your code. Uh, what we're seeing with a

  11. 3:21

    lot of the modern stuff with agents is, you know, you've got this magical being in the cloud that can do all sorts of really cool tasks, and one of those tasks is coding. Um, now there's a couple ways you can do that. One thing is you can just run things locally on your own computer, right? Who here has used like Codex? Claude Code? Right. Uh, so you can be like, "Hey, you have all the permissions you want to do." Uh, and it can go run around and then delete things, create things, whatever. Uh, some of the problems that you run into are, you know, it'll use up all your CPU, it'll spawn a bu-bunch of processes, it'll delete things

  12. 3:51

    it's not supposed to, maybe it'll start accessing environmental variables. Uh, there's a lot of things that are dangerous about that. Um, so that's where it's okay, "Hey, let's put it on the server." Well, on the server you have similar problems. You've got secrets, you've got, uh, your environment. And so you need to make sure that everything that you want it to be able to interact with, uh, is present. Uh, and that's where sandboxes come in. Uh, also, uh, with, uh, one thing that's a little bit different, uh, the way that we view sandboxes and other things is that sandboxes, you want to be able to very quickly spin these up. Uh, E2B, we sp- we aim to, uh, spin up

  13. 4:21

    sandboxes in less than a hundred milliseconds, uh, so that way you can run multiple things at the same time. Uh, do, do, do, do. So yes, uh, also, so with our sandboxes, uh, there's a couple things that we, uh, try to do. So it's, uh, when you go into a sandbox, uh, you're working and, uh, like when you close your laptop, you open your laptop again, you're back into your state. Uh, for us it's, uh, same thing. So we try to, uh, snapshot the memory, snapshot the file system, so that way, uh, whenever you stop working, you pause everything, save everything to

  14. 4:51

    storage somewhere, and then resume it. Uh, agents especially, right? So you've got an agent working away, uh, maybe it has to go do some sort of process elsewhere, uh, and then it wa-co- wants to come back, uh, to the state it was in. Uh, for the agent, it kind of doesn't know the time has passed, and so it's expecting everything to be exactly the way it was when it, uh, left off. Uh, and so by snapshotting everything and making it so that everything is still there when it comes back, uh, it makes it so you get less errors. Um, so yeah, the way, uh, our approach to these sandboxes is basically we start the

  15. 5:20

    processes first, uh, and then we, as I just said before, we're, uh, saving the disk and RAM, uh, and we're resuming sandboxes from that state. Uh, so that's a little bit different. Like if you think about Docker, when you start a Docker image, uh, you have kind of-- you build up a bunch of things, uh, and then when the process starts, uh, that's where you have your start command. Uh, for E2B ba- most of what we do is that you're starting from everything already running. Uh, so for instance, you, uh, when we have our templates, uh, which we'll talk about in a bit, uh, what you're doing

  16. 5:50

    is, uh, you start the processes and then those processes are already running when you come back to the sandbox. All right, uh, we're also open source. Uh, and so for infra, uh- You can see our links there. You can also Google, uh, that. We've got our, our backend's mostly in Go, uh, which is why I asked about the Go engineers; we're desperately hiring. Uh, and then also our SDK is in Python and JavaScript. Uh, working on extending that to other languages, but that covers quite a few of our use cases.

  17. 6:19

    All right. So normal usage of sandboxes are, uh, we have an API key. You go to e- e2b.dev, uh, and then you can also go read our docs. However, for this workshop, uh, we're not gonna need to worry about any of that 'cause we're just gonna do everything inside of our own sandboxes. Uh, we're gonna spawn the sandboxes through our dashboard, uh, that way we can hopefully get, uh, up and running very quickly. Uh, and, uh, that should eliminate some of the issues that we've heard happen in workshops. So, all right, and again, uh, I, I guess we found out most of us already know Python, uh, which is good.

  18. 6:49

    Uh, but for this workshop, we're not gonna need to code, uh, for the first several levels. Uh, once we get past, I think, level seven, that's when, uh, we start needing to code in Python or JavaScript if you want to. Uh, and we'll have tests that make things, uh, fast. Uh, we would've liked it... Uh, one of the challenges that we, uh, had when coming up with, hey, how do we demo, uh, some of the issues that we run into sandboxes, is, uh, a lot of issues happen at scale. Uh, so sometimes when you're managing just, like, one or two sandboxes, it's pretty easy. It's just like, hey, I've got one sandbox. I go into it, pause it, resume it, no issues.

  19. 7:20

    Uh, however, uh, once you start to deal with hundreds of sandboxes, thousands of sandboxes, uh, things start to take on a life of their own. Uh, for that, we have some, uh, basically simulations, uh, in, in the, uh, workshop, uh, which hopefully are fun. Uh, also, uh, even if-- I, I know it's 2026, and we have agents. Uh, so most of us don't actually code by hand anymore, or many of us don't code by hand anymore. Uh, you c- so most of the, uh, A- AI agents, uh, know the E2B SDK, and so what you'll be able to do is,

  20. 7:50

    uh, grab your sandbox ID that you get for each of the levels, give it to your agent, and, uh, let the agent just figure out how to write the code for you. Um, but this is for learning purposes, so if you wanna do it by hand, it's probably better. All right. Also, uh, I'm supposed to point this out, uh, if you're at a startup, we have a startup program, uh, where, uh, we give, uh, credits and a free plan. All right. So let's get to the first level. So let's go back to... Let's see. I don't know if there's a way to keep this up. But was everyone able to get to this,

  21. 8:20

    uh, URL here? Yep? Cool. All right, so let's go there. I already have it open, but that's fine.

  22. 8:29

    You wanna go ahead and show that one more time? Okay.

  23. 8:32

    So it is, uh, this tinyurl.com/wh46sx6. Uh, I don't know if there's an easy way to keep this up on the screen while we're going. But, uh, all of you have neighbors, and so you can just go to your neighbor and be like, "Hey, what, what's, what's that code again?" if, if, if you need it again. Uh, so I'll give you another five seconds. Can I just get a thumbs-up as people have gotten there? Okay, looks like those guys over there are good. Over here, generally good. All right,

  24. 9:02

    cool.

  25. 9:05

    All right. So what we're inside of here is already a sandbox, all right? So at the top, what we're gonna have is, uh, so this is, uh, our, uh, web UI for viewing our sandboxes. Um, and again, so as I said, we don't need to code yet. Uh, and so this is our dashboard. Uh, inside of the dashboard, if you want to, uh, make an API key, you can. Uh, you can go into your API keys here, uh, and, uh, create an A- API key, which you can then, uh, hook into this locally. However, we've got this gorgeous UI, uh, that's gonna tell us what to do. Uh, so inside of the sandbox, what we're

  26. 9:35

    doing is we're spawning a web server, and we're just gonna go up here, right? So we can go there now. Uh, and so here is our, uh, lab, all right? So, uh, our first task is very simple. Uh, it's basically, can you find out the sandbox ID that you're inside of? Uh, now we have a bunch of instructions for how to get that from the sandbox itself. Uh, I'm going to show you a cheat code. Uh, now in the URL, we actually have the, uh, ID of the sandbox, uh, and so we can just copy that, uh, ID. Uh, again, our first levels are mostly just to make sure that everyone's oriented, and they can get the

  27. 10:05

    different pieces working. Uh, and you can just paste it here, right? We check it, and, uh, so we should be good. Everyone-- Uh, so has everyone been able to get past the first level?

  28. 10:17

    Okay. Getting the ID is hard.

  29. 10:21

    Hmm?

  30. 10:24

    Do-- Well, everyone's gonna have different sandbox IDs.

  31. 10:30

    Right, but Oh, uh, that is our, uh, helper, one of our, uh, uh, go-to-market people from E2B back there. Uh, uh, just if you wanna come up here. Huh? Wi-Fi. Sorry? The Wi-Fi's not working. The Wi-Fi's not working. All right. Anyone else having Wi-Fi issues? Everybody, I guess. Great. All right, uh, so one of the-- Okay, so I tested this before I came in. Uh, so the Wi-Fi that we're supposed to be using is...

  32. 10:56

    It was, like, AI engineer something. Here we go. It's one of these things. Uh, AI engineer, uh, Wi-Fi. If you connect, you'll have one bar possibly. What you do is you dis- after you've connected, disconnect and reconnect, and it should give you five bars.

  33. 11:09

    Okay, I don't know why that worked. I got to

  34. 11:14

    Sorry? How'd you get to this screen without the sandbox? You got here already or no? Uh, to the workshop with the sandbox, but just not this page. I was one click back. Gotcha. Okay. So- All right. All right, so set up. So, okay, at this tiny URL, this should give everyone a terminal, right? You'll have to log into E2B. Make the font bigger. Oh, make the font bigger.

  35. 11:43

    Mm-hmm.

  36. 11:45

    On the URL.

  37. 11:50

    I'm just gonna keep zooming in.

  38. 11:55

    Okay. URL. Everyone had already gotten there.

  39. 12:01

    Go back. Go back? Okay, we can give you the URL one more time. Uh, after this, we have to start charging because, uh, screen time's expensive.

  40. 12:13

    Got about 10 more seconds to try to get to this URL, but, uh, I, I can also say it out loud. This is, uh, tinyurl.com/4wh6xa6. Okay.

  41. 12:32

    And that should bring us to a terminal.

  42. 12:36

    Do you need to sign up?

  43. 12:37

    Yes, you do need to sign up for E2B to use this.

  44. 12:46

    Huh. Also, our beautiful ASCII art does not look as good on the big screen as it does on a smaller screen.

  45. 12:54

    Cool. Uh, can I get a show of hands, who was able to get to this page? Okay. Who was not able to get to this page? All right. We don't wanna leave you guys behind. We're gonna make sure that we get the first, uh, setup part. But, uh, for those of you that were able to get to this page, you're good to go ahead and keep going through levels if you're, uh, able to.

  46. 13:17

    All right. I'm gonna give another...

  47. 13:22

    Let's see.

  48. 13:25

    So we're getting maybe 300 kbps.

  49. 13:29

    Okay. Yeah, yeah. It's-- Everything's running remotely in the sandbox, so hopefully we don't need much back and forth.

  50. 13:38

    Uh, famous last words.

  51. 13:44

    All right. Uh, can I get a sh- so for people that were still loading, uh, who's still waiting to try to get to the page? One, two, uh,

  52. 13:55

    uh, Ali, are you able to help these guys out? Uh, so we've got our, uh, helpers that can come through and, uh, help debug what's going on. Uh, what, what's the main issue we still have? It's just the internet connectivity? Okay. And then worst case, uh, everyone here is friendly, right? Uh, what we can do is we can partner up. Uh, and so, uh, let's see. Uh, can I get that show of hands and who's still trying to get to the page? One, two, three. Okay. Can I get a show of hands of, uh, people that are friendly and don't mind sitting next to

  53. 14:25

    somebody? All right. So if you're still waiting, look around for someone that's got their hand raised, uh, and you can, uh, kind of pair up.

  54. 14:33

    All right. So once we're here... All right. Yeah, sorry. It's, uh, we've got limited time, and we wanna keep moving. Uh, but, uh, yeah. All right. So once we're here, uh, we should have this URL now, uh, that our first sandbox is gonna spawn for us, and then we can just go straight to that. All right. And so everyone should get to a page that looks something like this, uh, that's gonna have, uh, our CTF instructions. Cool.

  55. 14:58

    Show of hands. Yeah. Also, y- you guys didn't know this, but this is, like, uh, an interactive, uh, workshop, right? Where it's raise your hand, lower your hand, raise your hand. Uh, Simon Says is next. Um, but yeah. So who got here? Show of hands. Great. All right. And then so who was able to find out their sandbox ID? All right, cool. Who was not able to find out their sandbox ID but is still here? Who is too embarrassed to say that they couldn't find their sandbox ID? Uh, okay.

  56. 15:28

    And again, so we can just pull the sandbox ID out of the URL. Uh, this ID basically is gon- is the identifier for the sandbox, right? So once you have a sandbox, you're gonna want an identifier so that you can go back to that sandbox. Uh, for E2B, the way that we construct our URLs, uh, is basically the port that you're talking to dash, uh, the sandbox ID, uh, .e2b.app. And that way, whenever you spin up a sandbox, you can expose services listening to those ports, uh, and, and then, uh, get comms open. So we'll, uh, refinish this, uh,

  57. 15:59

    first level.

  58. 16:02

    All right. So one of the things that happens, uh, uh, is once you've got users in a sandbox, you've got AI writing code, uh, a lot of the times what'll happen is the AI will write some- something that does something that you're not expecting it to do. Uh, it turned out while writing, uh, this workshop, we ran into this issue a lot, that we had processes that just took lots of CPU. So what we can do is we can start this sandbox, and you'll see up here, right, we're now inside of a new, uh, sandbox, right? So our new-- we got a new sandbox specifically for this, uh, project, which means we can break this, and we're not gonna

  59. 16:32

    break our initial stuff. Uh, if we want to restart a sandbox, we've got this, uh, fancy button over here to reload, right, and then it'll reload it. I guess, uh, while we're here, I should kind of point out some of the things that are happening, uh, that we don't need to worry about too much, but more just kind of framework on how, uh, we just manage our dashboard. Uh, we've got this terminal here that lets us, uh, connect a PTY, uh, session to a sandbox. Uh, a-and so, uh, we have this concept of templates, which I mentioned earlier. So a template, uh, for us is that

  60. 17:01

    snapshotted, uh, VM. And so what we've done is we started a VM, ran a bunch of processes to start the capture the flag, uh, and then we saved this as this template, which means all of you, uh, are basically taking that sandbox that I'd started before, uh, and then you're starting it where I'd resumed it, right? So we just made, I don't know, two hundred copies of this thing, uh, which is pretty cool, uh, just-- but-- that we can do that by itself. Uh, there's a lot of use cases for this k-kind of thing, depending on what you're doing at your business, uh, that you'll wanna be like, "Hey, I've got, uh, this work. I want other people to continue it."

  61. 17:32

    Um, and-- or even, like, if you're doing data science stuff, right? So let's say you've got a bunch of, uh, data in a data frame. You're like, "Hey, I wanna try mutating this. There's, like, five different experiments I wanna run." You can take that, fork that into five different things, uh, and then in parallel, uh, spin that up into five different sandboxes. So, uh, this up here is just the template ID that we're spawning. Uh, and then now that we have the sandbox ID, uh, when we re-reload this, we'll go back to the same page over and over again, um, and we can connect to that. Uh, one thing also is, so let's say that, uh,

  62. 18:03

    uh, I mentioned later on we're gonna have some code. Uh, for those of us that are jumping ahead, you can take the sandbox ID, you can create an E2B, uh, API key, uh, and then in your local, uh, you know, Codex, Claude Code, whatever, uh, you can be like, "Hey, stuck these in an environmental, uh, file. Use E2B to go talk to the sandbox and solve the challenge for me." Um, and that can make life easier so that way you don't have to code, 'cause again, uh, most of us are just letting the AI code for us nowadays, or in some way. Uh, don't just let it code for you. Uh, you should always review what the AI does. Um,

  63. 18:34

    do what I say, not what I do. Uh, so right. So each of these, uh, challenges is gonna have, uh, just like the first one, we can just stick this up here. Right, and so here's our CPU, uh, runaway task. Uh, for those of us that haven't coded a lot, haven't messed around with Linux recently, uh, we've got instructions on how to, uh, figure out what's going on. Right, so we can see here, uh, we've got another terminal. We're sandboxing terminals all the way down. Uh, this is opening another PTY into the local, uh, sandbox. Uh, and so basically, uh, our instructions here, hey, we've got a

  64. 19:04

    CPU, uh-- we've got some processes taking a bunch of CPU, uh, for these early levels. We've got instructions, uh, h-how to find, uh, what's taking up CPU. Right, and so we can list here, hey, what's going on? And we have this, uh, process up here, uh, that's been running for a bit, taking up ninety-nine point seven percent of our CPU that we probably, uh, should kill. Well, how do we kill, uh, a process? Uh, I forgot how. Uh, but luckily we have instructions here for how to do it, or we could also ask our

  65. 19:34

    friendly AI. Right, so what we're gonna do is we're gonna go ahead and kill that process. Uh, right, so we've got the process ID up here. Right, so this is showing our process, uh, how much CPU each of these guys is taking, uh, memory, uh, so forth. Uh, and so we're gonna go kill two fifty-- uh, twelve fifty. Right, now that he's dead, we have now identified, hey, we have solved our runaway process. We copy this amazing code here, go back to our, uh, initial capture the flag, and we've unlocked level three. All right, everyone able to

  66. 20:04

    do that? Right. Cool. Anyone stuck?

  67. 20:10

    Anyone embarrassed to say that they're stuck? Nope? Cool.

  68. 20:16

    All right, then we can open up our next level.

  69. 20:22

    Okay. And again, same thing, right? So we get a new sandbox.

  70. 20:28

    We can open this up. Right. And so this is another thing that happens a lot, right? You've got an agent running code in a sandbox, you've got you running code in a sandbox, uh, you're blowing up memory. Right? So one of the things about these sandboxes is, uh, you've got some trade-offs in the amount of resources that you're using. Uh, a lot of the times you're gonna wanna have, I don't know, uh,

  71. 20:50

    let's say, uh, you know, th-there's a trade-off between if you have multiple CPUs, you're gonna run faster, fewer CPUs, it's gonna go slower. Also, in terms of resource utilization, uh, one of the things that you start to run into, uh, with fleets is, uh, you have limited quotas on how many, uh, CPUs and how much RAM you're able to have for your instances. Uh, so there's trade-offs, uh, on can you just overuse it. So you're gonna wanna be like, "Hey, uh, how do I minimize, uh, my RAM utilization?" So again, uh, in this walkthrough, we can easily go through here, and then we can

  72. 21:20

    go into our sandbox, show what's using up our memory. Uh, do-do-do-doof.

  73. 21:28

    Right, and then we can kill it. Right, same thing. So for the first couple of levels, all these instructions are here. Uh, twelve forty-- twelve forty-eight.

  74. 21:40

    Bam. And then we get our key.

  75. 21:44

    Okay, am I going too slow, too fast? Right speed? Good? Right speed? Okay.

  76. 21:57

    Cool.

  77. 21:59

    All right. And then so finally it's, uh, hey, what, what happens once we start filling our disk? Right? So all these things have limited resources, uh, and all these things you're gonna run into at some point or another, things are gonna blow this up. So same process. Um, right, you can see we're in a new sandbox. We can LS around, right? We can see, hey, we've got our instructions in here. We've also got a README if, uh, you want to, uh, see more info on here. Right, so we can cat the README.

  78. 22:28

    Okay. And so, uh,

  79. 22:33

    yeah. All right. Uh, so do-do-do-do.

  80. 22:40

    Let's go to here. So here we've, uh, filled our disk, right? So same thing. Uh, we can check, hey, what's going on as far as what's taking up our, uh, our disk.

  81. 22:57

    All right, and then we just gotta figure out, hey, what is the path, uh, in here that's taking up all of our, uh, this? Right.

  82. 23:08

    Well...

  83. 23:12

    Huh, actually, this used to have all of our instructions in here, but right, we can see in here, uh, let's remove this guy.

  84. 23:24

    Okay. And again, now it should say that we've, uh, succeeded.

  85. 23:39

    We also have a-- Oh, we were supposed to put it

  86. 23:43

    here. Just kidding.

  87. 24:14

    All right. So the easy levels apparently were not as simple. So in theory, well, what we can do is we can actually just reload this, uh, sand-

  88. 24:26

    I broke the demo. That's okay, we can just start a new one.

  89. 24:31

    Let's try again,

  90. 24:35

    and let's pretend I didn't break that. Ah.

  91. 24:42

    Okay.

  92. 24:51

    Okay. We should've just been able to delete it,

  93. 24:56

    but let's do it here then.

  94. 25:05

    Okay. There we go. So, okay.

  95. 25:12

    All right, and then so we were-- Then it, uh... We get the, uh, we, we were supposed to paste the, uh, path in there instead of remove it, I guess. I broke that. Um, and so everyone able to get to this, uh, flag? Or... Cool. All right, great. Anyone not able to get to this flag? Okay. So, uh, real quick, uh, for how I solved that one. Uh,

  96. 25:38

    what you do is, so uh, on the disk fill, uh, challenge, what we're looking for is, uh, which, uh, w-what is taking up all of our file, uh, s-space. Uh, and so what we need to do is we need to find, uh, up here in the hints, uh, you can copy the, uh, the, uh, script, uh, to find where, uh, what's taking up all of our, uh, disk space. And then here we should be able to see that we have this /temp/self-own-diskfill.bin. Um,

  97. 26:07

    and then what we can do is just paste it, uh, down here into the path, uh, checker. All right? And then click on, uh, here to clean it up, and that'll run the delete command for us, and then, uh, we should get this, uh, task, uh, complete flag. Everyone got there?

  98. 26:23

    I don't see disk fill path.

  99. 26:25

    So-

  100. 26:25

    I don't get disk fill path.

  101. 26:27

    You don't get disk...

  102. 26:30

    Disk fill path.

  103. 26:31

    Disk fill path.

  104. 26:33

    The path that you're copying, it's not there in my sandbox.

  105. 26:36

    Okay, one second. So you've gotten to, uh, disk fill level four?

  106. 26:41

    Yeah.

  107. 26:42

    Cool.

  108. 26:45

    Third time's the charm. All right, so we can open, uh, th-the page for the challenge.

  109. 26:54

    Uh, and then so what we wanna do is, so up here we have, uh, some code that we can run in order to find, uh, what the, uh, what, what files, uh, are taking up the most disk. Okay? So are we able to get to here so far? Cool. Right, then w-when we run this, uh, duah, uh, we should be able to find the temp file, the file in temp, uh, this, uh, diskfill.bin, uh, that's taking up five hundred and twelve megs of, uh, disk space.

  110. 27:29

    I don't have that.

  111. 27:31

    You don't have that?

  112. 27:32

    I don't have that file.

  113. 27:33

    What do you have?

  114. 27:35

    Uh, just that sandbox

  115. 27:41

    Sorry?

  116. 27:41

    There's nothing inside of it.

  117. 27:41

    There's nothing inside of there? All right, so what we can do, right, and that's where, uh, the great thing about sandboxes, you can restart them. Uh, right, so if you go back into here, when you, uh, click on the Start Sandbox, uh, we can reload the snapshot by clicking on this reload button up here. Right, and so this will spawn us a new copy of it.

  118. 28:09

    Okay, everyone else good?

  119. 28:10

    Did you just hit the refresh button on the top right-

  120. 28:12

    Yep

  121. 28:13

    ... of the sandbox?

  122. 28:14

    Yeah, exactly. So it's this button here, refreshes it, and then you can see here that's saying, "Hey, uh, start here, open this browser page, uh, for this, uh, set of, uh, tutorials." Um, it'll say r-we've returned, uh, if it's the same one.

  123. 28:28

    And that resets it to like its original state?

  124. 28:30

    Exactly. Yep. And so that's where, uh...

  125. 28:36

    Yeah, it's, it's-- You can see, like, our start time is pretty fast, right? Clicking here, you're able to create a sandbox in-- relatively quickly. Uh, and so it's-- we can create as many sandboxes as we wanna create.

  126. 28:49

    Quick question.

  127. 28:50

    Yeah.

  128. 28:51

    You don't have any funny characters that if have some coding agent go and scrape this or things like that, we won't have any issues, right?

  129. 28:59

    Nope. Yeah, I mean, yeah. Uh, this is just the terminal showing the sandbox. Most of the time when people interact with this, it's all programmatic. So you have a web server that's going in programmatically and running code in the sandbox purely for demonstration purposes, uh, and also for debugging purposes. Uh, being able to quickly go into the browser and pull up a sandbox can be useful. Um, and also here it's kind of we're able to spin up a webpage really quickly to get in and see what's going on. Um-

  130. 29:23

    I can do it in programmatically?

  131. 29:24

    Yes. Yes, yes. And that's what I was saying. So, uh, yeah, um, we'll come to that in a couple levels where we actually start, uh, interacting, uh, programmatically with, uh, these sandboxes. But mostly it's right now we just wanna make sure everyone's on the same page for what is CPU, what is RAM, um, what's gonna blow up, uh, when we're starting to run a whole bunch of different, uh, things that we're getting arbitrary code execution run on. Cool. Everyone good now?

  132. 29:52

    Awesome. All right. Um, one thing I didn't do was get the code for that level, so I'm just gonna go, uh,

  133. 30:04

    restart him real quick.

  134. 30:08

    I'm just clicking around on buttons.

  135. 30:20

    Okay. Uh, let me find our pro-- our, our, uh, file that's taking up all of our disk. So let's go here.

  136. 30:30

    And then we can check our path. Okay. And we'll click on clean up.

  137. 30:37

    Great. All right, and so that's the end of our easy, uh, easy tasks.

  138. 30:45

    All right, how's everyone feeling? We, we know all about how to mess around with Linux now. Feel comfortable? Yes, no? Great. Uh, let's stand up for like a minute real quick. Uh, get the blood flowing, uh, kind of stretch out.

  139. 30:58

    Okay. We're also a calisthenics, uh, company.

  140. 31:05

    Okay.

  141. 31:08

    Cool.

  142. 31:11

    All right. So that was, hey, this is what Linux is. Uh, our next set of, uh, uh, tutorials are gonna get us into actually looking at the sandbox life cycle. Uh, so with a normal server, you spin up the server, and it lasts there until you're like, "Hey, I wanna shut it down," or Amazon is like, "You haven't paid your bill in a while, we're shutting it down for you." Uh, or actually, if they're doing... Anyway, TLDR, uh, for sandboxes, a l- a lot of what you're doing is looking around the life cycle of the sandbox. Uh, there's multiple things that we need to care about here. So one

  143. 31:41

    is, let's say you've got thousands of users or tens of thousands of users, and you're like, "I wanna give every single one of these guys a sandbox when they come to my application." Well, how long should that sandbox last? Right? A lot of, uh, you know, intuitively it's like, hey, I'll give them the sandbox. They land on my webpage, I'll give them the sandbox for 20 minutes. If they don't use it, I'll shut it down. But if you've got thousands of users and they're only gonna use it once or maybe not even interact, interact with the sandbox, you're wasting a whole lot of resources spinning up sandboxes that no one's ever gonna use. Uh, and so that's where... But maybe it's like you've got people that are gonna come

  144. 32:11

    in and bounce, right? They'll come in, try it for a second, be like, "Hey, does this thing work?" And then bounce. Uh, in that case, what you wanna be able to do is you wanna be able to have the sandbox there, have it be able to run their initial command, and then have it go away once, uh, you know, once they've done the command. And maybe, you know, they'll come back again in 10 minutes or so. Uh, you wanna have that sandbox come back once they're there. Uh, so that's where for the sandbox life cycle, um, y- you need to make sure that you're basically using as little sandboxes, uh, a- as you can at a time. However,

  145. 32:41

    you know, let's say you've got people that are running long, long-running processes. Right? So let's say you're doing data science, web scraping, whatever. Uh, you don't know how long, uh, the user's tasks are gonna run, right? So maybe you've got the users coming in, they're running a task that takes 10, 15 minutes, uh, and then they're doing whatever, right? So it's how do you get that trade-off, uh, between how long you keep the sandboxes sitting around, uh, versus not eating money into the sun? So with this task, what we're gonna do, uh, is we're gonna take an approach, uh, to our sandbox management, where what we're gonna

  146. 33:11

    do is we're gonna give sandboxes a really long time, uh, to live, right? We're gonna say, "Hey, when a user comes in and runs a command, we want that sandbox to run for as, you know, as long as we want it to. But once the sandbox finishes its task, we're gonna give it a very short timeout, so it'll stick around in case they give it, you know, another command. But then it'll shut itself down, and then that way we can free up those resources to spin up another sandbox for somebody else."

  147. 33:36

    So we're gonna go ahead, open this up. Right. Uh, this one, uh, there's no code involved, right? And so all we're gonna have to do here is think about, hey, when we've got these, uh, people coming in, uh, running, uh, commands, how long should things, uh, uh, the state machine, uh, for how long should this run, right? Uh, so you can read up here, right, this thing can take a long time. Uh, let's say that we've got a task, and we think, uh, the initial task, uh, we should have the instructions on here somewhere.

  148. 34:07

    But we've got a policy. But basically what we wanna do is we want to, uh, allow this task as long, as lo- as long as it wants to, but afterwards, we wanna shut it off after a minute. So what we'll do is we'll say, "Hey, we want this command to be able to run for an hour." Okay, so we set our command to run for an hour. We can then run our task. Okay. And this is all simulated 'cause we're not gonna have you guys actually wait an hour for a task to run in a sandbox. Uh, and what we'll do afterwards is we'll say, "Hey, I want this sandbox to only s- uh, stick

  149. 34:37

    around for a minute, uh, and then it should shut down afterwards." All right? Very simple what I just did there, uh, 'cause we're cheating. But, uh, that's basically, uh, what the goal is here, is thinking about when we've got a, a sandbox that we're just gonna have it, uh, do a task. Uh, we want it to stay, uh, around as long as the task is gonna take, uh, but then afterwards, once we get that result back, we can be like, "Hey, uh, have it go away later." Cool. So everyone able to get to this slide?

  150. 35:07

    Or is that, um, did I spin that, uh, go through that too quick?

  151. 35:12

    Mm. Mm. Up, thumbs up. Okay, I see a lot of thumbs up. Great.

  152. 35:19

    All right, and then if you wanna cheat, uh, you can just put in the code here. So it's WRITE_SIZE_RUNTIME_LIFETIME. Uh, let me zoom in on that for you guys.

  153. 35:30

    All right. Uh, so we can come around or, uh, would it help if I went through that again real quick for anybody here just to show how that worked?

  154. 35:39

    Okay, I'll go through it one more time. All right. So we'll just spin up this, uh, that task again.

  155. 35:46

    We have a question over here.

  156. 35:47

    Yeah.

  157. 35:47

    Is that just like-- You said task, you mean like one Linux command?

  158. 35:51

    Exactly.

  159. 35:52

    Okay, cool.

  160. 35:52

    Or let's say the user has like, um, th- they're having to read a large file, or they wanna do video processing, right? That can ta- You don't know how long that's gonna take. Uh, and so what you wanna do is you wanna give the sandbox a long life cycle, uh, so it can complete that task. Uh, but then after it finishes the task, you don't wanna keep it sitting around using resources. Uh, so you give it a shorter timeout afterwards to be like, "Hey, unless the user starts interacting with this, shut-- pause it, shut it down, put it into so you can resume it later."

  161. 36:22

    But if it's like a web server or like a Cloud Session that might exit-

  162. 36:25

    Right

  163. 36:26

    ...

  164. 36:28

    you can't close the server.

  165. 36:28

    E- Exactly. Well, Cloud Session is similar, right? So with, with a Cloud Session- Yeah, speed matters, right? And so it's all these things, uh, you end up d- dealing with like a speed versus cost trade-off, um, where if a sandbox is running, things are fast. Uh, if it's paused, i- it's cheap 'cause you can have as many of them, you know, you can have lots of them paused and not be paying for them. Or, you know, if you're CPU limited, you just, you can't run that many sandboxes at a time. Um, but like for something like Claude, you don't know when the AI is gonna come back to reuse

  166. 36:58

    that sandbox, right? Maybe it's using 20 sandboxes, right, and it uses this sandbox, and it's gonna... So, uh, th- that's where it's... Yeah.

  167. 37:06

    So you end up having to, like, do daemons in your sandboxes to, like, read the web server and you can't actually-

  168. 37:11

    Yeah. Well, okay, so if, if, if you're talking daemons where it's like you want something that's persistent and long running, sandbox may not be the best thing to do there. Uh, you can, right? You can definitely have a sandbox that you're like, "Run it for a day," right? Uh, but there's probably cheaper, uh, or more ki- kind of traditional methods to do that. But again, we're running all this stuff inside of a sandbox right now, inside of multiple sandboxes, so you can definitely run web servers inside of them. Um, but they're just not gonna be optimized, like, as far as CDNs and that. Yeah. Um-

  169. 37:40

    Yeah. Can you use this in your main development line-

  170. 37:43

    Yeah

  171. 37:43

    ... where you start with the vanilla sandbox, you keep adding it, other folks get some replica of that sandbox-

  172. 37:49

    Yep

  173. 37:49

    ... and eventually you bring it all together and you, this is the production sandbox?

  174. 37:54

    Yeah, so how to f- uh, fork things, exactly. But th- that's kind of where a lot of people are using these things, is that you can give the AI a sandbox, say, "Hey, let's install the base dependencies," like you would with your GitHub actions. So you've got kind of a basic cache thing, and then you can pull it in and then build on top of it, right? And then that way if you're doing, like, work trees, right, you've got a single one, each work tree has... Yeah.

  175. 38:15

    Yeah.

  176. 38:15

    Yep. Okay. So anyway, so for here, uh, go back to our life cycle, right, real quick. What we wanna do is basically we wanna give it a really long runtime, right? And so here we're gonna give it an hour.

  177. 38:29

    And then, but once it, once it finishes the command, we wanna spin the, that, uh, th- th- the, uh, spin the timeout back down to a minute. And so what we're gonna do is we're gonna set our command timeout, we're gonna run a fake task, uh, and then we're gonna set our, uh, whoops, set our lifetime back to something very, uh, short.

  178. 38:53

    Okay. And then that should give us, hey, we've solved this, uh, challenge.

  179. 39:00

    Cool. Yay, complete.

  180. 39:05

    All right, uh, n- so and then this gets us into, uh, our next, uh, lab, which is our background process leaks. Right, so, uh, definitely when you're coming back and reusing the same sandbox over and over again, uh, you're gonna run a, uh, run into orphans, right? So basically, you start a process, uh, and then you go and do something else. You come back to your machine. If that machine is still running, right, you've, you end up, uh, ac- accumulating a whole bunch of different processes running over time. Uh, and that can kill performance in the sandbox. Uh, for some reason, with sandboxes, we th- see this

  181. 39:35

    a lot. Um, but it kind of makes sense 'cause, like, you're pausing your process and coming back to it rather than getting, like, a new image every single time. Uh, so it's... The advantage of, right, resuming this old state, things are fast, things are where you left them. Downside is things are where they left them and you ha- if you haven't cleaned up after yourself. So a lot of times what you'll have to do is you'll have to come in and try to figure out, how do I, uh, kill my orphans?

  182. 40:01

    Actually, I think a lot of sandbox management ends up being killing orphans. Uh, you end up with, uh, a lot of the time also is, uh, related to this, is, uh, when you start up sandboxes, you need to keep track of them. Uh, and if you lose track of your sandboxes, uh, you need to figure out h- how do I go find the sandboxes that or- or that I'm not managing anymore, uh, and make sure that I remove those resources. Um, so here is a demo of how to go, uh, find our orphans, right? And so, uh, we have our code here,

  183. 40:31

    uh, that we can, uh, in this nifty terminal, uh, we can go ahead and, uh, we can run some code to figure out how many orphans we have. Right, so right here we can see we hop into the sandbox, and we've got, uh, three orphans, uh, uh, that are alive, right? So we've got three children, right? And then if we, uh, search for them, what we can do is o- uh, often your, uh, processes that you orphaned are gonna be the same. Uh, here, uh, we're able to grep for them by this nifty, uh, somehow they're called orphan worker.

  184. 41:01

    Uh, in real world, unfortunately, they're not normally named this. Um, but, uh, for simplicity's sake, uh, here we're gonna be able to easily find them by just looking for orphans. And so all we need to do now is kill our, uh, leaked children.

  185. 41:22

    Okay. And actually, while I'm doing this, uh, so how many people here have run into orphan problems on their just regular servers as far as having a bunch of processes running? Yep. Okay, these guys over here, right? It's, like, a very common problem that you're gonna run into.

  186. 41:37

    Um,

  187. 41:39

    but, uh, so right, we're gonna kill 448. Uh, then we can run our script again, see how many orphans we have. Now we know there's two running. All right, so we still got our two orphans, and we're gonna kill 449 and 50. All right, and now this should notice. There we go. Once we've, uh, gotten rid of all of our orphan proc- processes, uh, the web page should notice, uh, that they're no longer there and give us the code to get to the next level. Cool. So everyone able to clean up their box and, uh, get this

  188. 42:09

    code? Cool. Anyone having a hard time finding the processes that are running around and, uh, hasn't been able to get them yet? Great.

  189. 42:21

    So we will take this flag and drop it into here.

  190. 42:29

    Okay. Uh, so, uh, next thing for sandbox management that we wanna, uh, cover real quick- Is, uh, our, uh, workspaces. All right, so when you first start dealing with sandboxes, uh, one of the first things that you do is like, "Hey, I've got all these servers, I will just hand them out to users." Uh, here we're kind of coming into we've done this naive approach to the sandbox management that we're just round robining the sandboxes. So rather than, uh, keeping track of what user gets what sandbox, uh, we're just, you know, you come in, you get a sandbox and then, uh,

  191. 42:59

    right? And so what we wanna do here is, this is the first, uh, code exercise, uh, is we're gonna wanna go look at our Python code, uh, that's on... Uh, so we have a little basically, uh, handler inside of here, uh, and we're gonna wanna modify it, so instead of giving us a round robin, uh, sandbox assignment, we're going to keep track by user of what sandbox they get. Um, we didn't wanna blow up our servers, uh, with spawning a sh- ton of sandboxes that we don't need, so we have a bunch of mocks, uh, that are gonna let us, uh, test what happens, uh, w- with our management. And so if we open

  192. 43:29

    up this, uh, terminal, we can go into our, uh, router.py, router-js. I think most of us know Python, so we're gonna go with the Python today.

  193. 43:40

    All right. Um, we have a Read Me in here as well, but we're running out of time. So, um, actually, let's see if I have instructions on this page, too. Yeah, so basically it's, right now we've got a, a code that runs round robin. We wanna change our code so that instead of b- being round robin, we're gonna assign this to users. Uh, the way that you do it normally is you'll have a database keeping track of, hey, here's all my users, here, here's all the active sandboxes for each of the users. Uh, keeping this persistent generally is useful, uh, so that way you can go

  194. 44:10

    back hisor- over historic data and figure out, like, were we reassigning too many sandboxes to users or what's going on. Um, but for here, uh, we're just gonna have this assignment dictionary. All right. And so all we're gonna need to do is we've got this assign sandbox function down here. And so right now it's just assigning it, and so what we're gonna do is we're just gonna quick modify this, right? So we've got that nifty assignments. Oh. Okay. Uh, 15 minutes left.

  195. 44:41

    So check if, uh, we've already kept track of this guy in our assignments.

  196. 44:46

    Also very hard to, uh, type on a screen that's way down there. Uh, and we're gonna return assignments of a user ID.

  197. 44:58

    We've already got this function here that's gonna get us the sandbox ID from a sandbox. So we'll say sandbox ID equals that thing.

  198. 45:07

    All right, and basically we're just following this. Uh, you should have instructions. Oh, we don't have instructions. All right. So we're gonna cheat. So, uh, uh, if anyone is... Has anyone already gotten past this level? Okay, so a couple of us have already gotten past this level, uh, so it is solvable in the time that we had here. Uh, but so all we're gonna do is we're gonna assign, uh, this sandbox ID here, and we'll return the sandbox ID.

  199. 45:34

    Right. So, uh, what this is gonna do is now i- uh, assignments is gonna be our cache of, hey, this is the ID that each user gets assigned. Uh, and then we'll return that. Otherwise, we'll have to create new sandboxes for everybody, uh, and return that sandbox ID. Okay. Uh, okay. Great question. All right. So basically what we wanna do is all we need-- Uh, the, the point of this is just saying we want to keep track of what user had what sandbox, so you're creating a map of user to sandbox ID.

  200. 46:05

    Right? That's, that's, that's the only point. Uh, so kind of starting with, hey, w- we're not keeping track of any of our sandboxes. So now it's like, hey, each user's gonna have a dedicated sandbox so they can rerun code, uh, in the same sandbox every time. Okay. And then we should have some tests in here,

  201. 46:26

    uh, that we can run and hopefully, uh, y'all are better coders than me. Uh, line 20. What did I mess up?

  202. 46:50

    All right, it'll run the test real quick.

  203. 46:53

    Failed test. And I'm gonna fail my tests.

  204. 46:57

    That is not-- False is not true.

  205. 47:04

    Uh, one second. You haven't added it to the assignments yet. Sorry? You haven't added it to the assignments yet. How do you get it then? Oh, good point. Exactly. See, that was a test of the, uh, audience, and good job, you passed.

  206. 47:20

    So right, we need to add it to the assignments.

  207. 47:33

    And so now that, uh, we're actually adding our sandbox ID to the assignments, uh, we should then, um, pass. Great. And now that we've passed, we should pick up over here that we have passed. Or...

  208. 47:49

    Come on.

  209. 48:03

    All right, so remember how I told you about how we learned about the, uh,

  210. 48:08

    uh, we learned about the, uh, orphans? Well, we have quite a few orphans running now. Okay, there we go.

  211. 48:16

    So took it a minute,

  212. 48:19

    but let's, let's just get to the next leg. Uh, right? Cool. Everyone able to get that?

  213. 48:26

    Cool.

  214. 48:30

    And we'll paste that here.

  215. 48:35

    All right. Uh, file system permissions. All right, so, uh, I, I think- Yeah, both Claude Code and Codex have this, is that a lot of times you don't want your agents writing to every single part of your file system. You're gonna have different parts of your file system, uh, that you want people writing to and reading to. Um, normally this would go into your AGENTS.md saying, "Hey, this is what I want you to write to, what I, uh, want you to read to." Uh, for these purposes, we're just gonna hard code this.

  216. 49:02

    But same thing, uh, right? So you're gonna have, uh, often you wanna have read, uh, write parts of your file system. So here,

  217. 49:11

    let's take a look. Uh, so, uh, we've got a cache config, and so, uh, then cache.

  218. 49:25

    All right. And then we can do the same thing. So here, uh, basically we've got a template cache and a runtime cache. Uh, so this is where when we build our templates, we're gonna have a cache that we're putting our template, uh, information to. That might be, uh, read-only. We've got a different cache that we want to, uh, write to, so all we're gonna do is for our cache store, we're gonna have, uh, refer to this, uh, cache here, which we're gonna spell correctly.

  219. 49:50

    And, uh, there we go.

  220. 49:54

    Cool. So that was a very fast one, where all we're doing is making sure that we're, uh, writing to our, uh, runtime cache. All right. Uh, at this point, let's see. I think y'all can do the rest of this on your own. We've got about 10 minutes left, uh, to talk about sandboxes. Um, so yeah, uh, hopefully that was a good introduction to hear, uh, issues that you're gonna run into with your sandboxes as far as managing, uh, what's able to do what, what's able to read to what.

  221. 50:24

    Uh, we've got about, ooh, eight more levels. Um, actually, I'm, I'll leave this to you guys. Do you think it'd be helpful to continue walking through all these or more fun to just talk about sandboxes? All right, let's talk about sandboxes. Great. Um, let's see. Uh, so who here, who here is currently running sandboxes in production? Oh, wow, those guys back there are running sandboxes. Anyone else here running sandboxes? Those guys over there. Uh, what are you guys doing with sandboxes? Uh, we run agent

  222. 50:53

    workloads. Agent workloads? Cool. Do you use E2B? Uh, no. Oh, okay. You know what? No one's perfect. Not yet. Not yet. Okay. Uh, we've, uh, Travis is over there. You should go talk to him. Uh, and then what do you do?

  223. 51:08

    Uh, we run agent workloads with Okay, cool. Woo-hoo, E2B, great. Um, yeah, yeah, yeah. So, uh, any of the, uh, issues that we just covered, ha- have you run into any of those? Uh, not yet. Oh, okay. So far we're cool. Gotcha. Yeah, yeah, and so we are, uh, w- part of what we do is we try to make it so you don't run into issues. Um, but your user-- you're always gonna run into something where it's, uh, something is happening on the sandbox, you need to hop in, debug it, and figure out why is this thing breaking. Um, great, and then we had some guys doing sandboxes over here.

  224. 51:38

    Yeah.

  225. 51:39

    So I have a slightly different use case.

  226. 51:41

    Yeah.

  227. 51:42

    I built my, uh, not production, but vanilla sandbox on the laptop, like on this one.

  228. 51:49

    Okay. So y- Sorry, uh, just so everyone else can hear. So you've got your vanilla production on your, on your laptop.

  229. 51:53

    On laptop.

  230. 51:54

    Okay.

  231. 51:54

    And then I-

  232. 51:55

    Oh

  233. 51:55

    ... I want to spread across on several GPUs, which are SL- like small, what do you call, small part but, uh, GPUs.

  234. 52:02

    Okay. Uh-

  235. 52:03

    And collect the data back-

  236. 52:04

    Okay

  237. 52:04

    ... and then process it over here. And they are across, what do you call, networks, different providers-

  238. 52:09

    Yep

  239. 52:10

    ... everything like that.

  240. 52:10

    So basically you're saying you, you've got your code local, local. You want to put it onto a bunch of distributed GPUs, process the, uh, the workload, and then send it all back to, to local. Okay.

  241. 52:20

    But I want the whole sandbox back.

  242. 52:22

    Yeah, so you want the data back.

  243. 52:23

    Yeah.

  244. 52:23

    Yep.

  245. 52:24

    So is that something you have?

  246. 52:25

    Yeah, so right now we don't support GPUs, but that's o- of course, a common use case that you can s- you know, when you're spinning up these contain, uh, whatever they are. So for us, we use VMs. A lot of people also use containers. Uh, common use case is you'll be like, "Hey, uh, take the output of whatever's running into there and send a post request home with the data that I want it to, uh, to have." Uh, that's, uh, like all these, depending on how, how big the workload are, uh, you're gonna have different bottlenecks and challenges. Um, where we've had, you

  247. 52:55

    know, it's one of the weirdest things about sandboxes, right, is that, uh, when you're just running a couple, you've got certain problems that are like, "Hey, why is this behavior, you know, why is one or two sandboxes breaking?" Once you're dealing with a lot of sandboxes, uh, then your problems become how do I manage all, you know, just e- just the logging overhead of how do I keep track of these? How do I make sure that I'm assigning all these? Um, uh, keeping track of the network from the machines that it's running on, right? So each machine can only be talking to so many sandboxes at a time. But, uh, yeah, definitely if you're running on GPUs,

  248. 53:25

    E2B doesn't do GPUs. Um, and I guess, so yeah, that's where s- the sandbox solution might not be... Although why would you need sandboxes for that as opposed to just having the GPU?

  249. 53:39

    Because I can get it at any place at whatever cost, different, uh, cost

  250. 53:43

    Gotcha.

  251. 53:44

    So wherever I get it, I'll put it in a private data center, then throw it.

  252. 53:48

    Yep.

  253. 53:48

    And I want the whole thing back again. So it's like all secure, close it, and whatever it wants, and it comes back.

  254. 53:54

    Makes sense. Yeah, and that's, that's a common use case. Uh, I just don't know enough about that, I guess, to give a good answer. Cool. Uh, and then so we had someone else over here using sandboxes. Okay.

  255. 54:04

    Yeah, just

  256. 54:07

    running Claude in Kubernetes.

  257. 54:08

    Something Kubernetes.

  258. 54:09

    Claude in Kubernetes.

  259. 54:11

    Claude in Kubernetes?

  260. 54:13

    Yeah. Um, one question I have is about storage. Like over time, the, there's gonna be a lot of storage after seven days. So how do you manage that?

  261. 54:24

    Yes. So okay. So one of his questions is, well, how do you deal with the storage, uh, over time, right? So how do you prune, uh, the resources that y- you develop? So one of the problems with sandboxes is, uh, you can create lots of them, but then it's where do you store them? Uh, that's an active challenge that we're currently working on. If you, if that interests you, we are hiring. Um- Uh, and again, uh, if, if you do Go, uh, we're, uh, hiring Go engineers. So our backends Go, uh, a lot of really fun challenges there. How do we, uh, optimize speed, uh, performance, and then

  262. 54:54

    especially on storage, uh, how do you deal with the compression, uh, and, uh, uh, that sort of challenge, uh-

  263. 55:01

    Volumes, Matt.

  264. 55:02

    Sorry?

  265. 55:03

    Volumes.

  266. 55:04

    Volumes. Right. Oh, yes, we also have volumes. Um, and so that's where if you have shared resources that multiple things can share, uh, volumes are great for that. Um, but yeah, it's, uh... I think all this is very early, and nobody knows where it's gonna go. Uh, so as the smart pe- you smart people in this room are gonna come up with better solutions, uh, th- that optimize things. Uh, don't know if that's a good answer. Uh, but yeah, so for, for your question on Claude, right, so it's... You can run Claude in sandbox. It's pretty easy. Um,

  267. 55:31

    uh, where, where do we go? So it'd be, right, so the same way that we, uh, spun up these guys, we can just be like template equals Cl- I think Claude should spin us up a Claude code.

  268. 55:42

    I'm guessing at what things do. Okay. Yeah, I think we have Claude on here. Right, so you c- yeah, you can run Claude in a sandbox, um, or you can run OpenCode in a s- right, uh, Codex in a sandbox. Uh, basically these, uh, give us places where you can have agents run remotely, or locally you can have the agent talk to it, um, and, uh, you get the same thing. But yeah, as far as kube, uh, so we are Firecracker, uh, based. Um, and so we're trying-- There's some technical issues, uh, with kube,

  269. 56:12

    uh, that we have that we haven't quite solved. Uh, so right now we're, uh, running Nomad, uh, but working on getting, uh, i- into kube. Yeah. Cool.

  270. 56:22

    How about coordination across sandboxes? So say you've got one orchestrator agent that's using one that has like a team of seven agents

  271. 56:30

    Yes. Yeah, yeah. So that's definitely common u- uh, definitely doable. Uh, we didn't do that for this exercise. Uh, so well, we kind of did for this where you were the coordinator, right? So basically the sandbox, uh, you clicked on a thing and said, "Hey, start a new sandbox, uh, w- from this template." Uh, that wasn't automated in this case 'cause I didn't wanna deal with, uh, making everyone grab their API keys. Uh, I also didn't trust myself to... This is being recorded, uh, and to not expose my own API keys by putting them up on the screen. But

  272. 57:00

    yeah. Um, but yeah, definitely that's kinda cool that you can get that inception that you have a sandbox, and then it spawns more things. Cool. Uh, we got time for two more. Yeah.

  273. 57:11

    So what's the ecosystem of these sandboxes like after the Claude, the Hermes, and all of these guys? Do you see that there'll be like a Docker ecosystem of, uh, sandboxes, or everybody will go their own way?

  274. 57:22

    Yeah, so the question is, what's the ecosystem of the sandboxes? Uh, is it gonna be like Docker? Uh, a- and I mean, all this is early, right? So like these AI agents are what, like two years old. LLMs as a service, five years old. Um, nobody knows what the ecosystem's gonna look like. But definitely as people build cool things, uh, we tend to standardize around them, share how to do it, uh, and then that leads to, uh, yeah. So sh- sh- right, and that's where our templates come in, right? Where again, you know, here I had a template, and 200 of you guys

  275. 57:52

    were able to then reuse that same sandbox, uh, without too much of a, a... Yeah. So we're going that direction, but time will tell. Yes.

  276. 58:02

    You know, like everything you know about products, do you run any of those local sandboxing stuff for your locations anymore?

  277. 58:09

    We actually... Okay. So internally at E2B, uh, we're running... I mean, yes, we run locally, especially for testing, right? Um, because, uh, running in the cloud is just a pain. But as far as like, uh... And so yeah, actually we're, we're, we're doing a lot of local, uh, uh, development.

  278. 58:26

    Okay. So like, uh, I mean, are you like blocking out network egress from that or like where is

  279. 58:33

    your

  280. 58:33

    Yeah, so I, I, I... I mean, that's where-- So our sandboxes have different rules that you can set for the, uh, network. Uh,

  281. 58:40

    uh, that's where if you give it access to secrets, you definitely wanna, want to, uh, prevent what it's gonna be able to send out. Uh, if it doesn't have secrets or it's just got dev keys maybe. And so that's, right, it's the ease versus, uh, security. Cool. Anybody else? Going once, going twice. Anyway, hope you all had fun. Uh, the-- You, you can go ahead and complete this, uh, these challenges on your own. Uh, thanks for coming. And again, I'm Matt with E2B, and, uh, it was a pleasure meeting y'all.