AI Engineer World's Fair 2026

Cooking with Codex — Charlie Guo & Gabriel Chua, OpenAI

Read the talk

Cooking with Codex

Selected presentation frame from Cooking with Codex — Charlie Guo & Gabriel Chua, OpenAI at 346 secondsOpen full source frame
The six-step outline for working with Codex begins with context and building, then moves through taking action, collaboration, repeatability, and embedding Codex.

Charlie Guo and Gabriel Chua build a live translation app from a Slack conversation, then connect context, browser control, goals, threads, hooks and automations into a repeatable development process.

From a talk by Charlie Guo and Gabriel Chua

At a glance

Ideas worth remembering

  • App shots carry visible application context and embedded text into a task; plugins add reusable practices, service connections and tools.

  • Choose browser and application control by the task: deep local inspection, operation of software without an API, or access to an existing signed-in browser session.

  • Long-running goals need verifiable completion criteria, visible progress and a way to surface blockers. A technically satisfied instruction can still miss the intended result.

  • Subagents separate work for the model; visible threads make workstreams inspectable by the person. Remote threads can delegate browser testing to a local thread and receive findings back.

  • Hooks run deterministic checks at workflow checkpoints. Heartbeats resume a waiting thread, while scheduled automations can start separate tasks.

  • A repeatable development loop connects incoming feedback to implementation, testing and review, with human input retained where access or judgment requires it.

A Slack conversation becomes the starting specification

A Friday brainstorming conversation contains the demo idea: build a real-time translator. It also contains a suggested translation capability and the idea of using a plugin to supply macOS development practices. In this workshop, Charlie Guo and Gabriel Chua of OpenAI turn that conversation into a build request without first rewriting it as a specification.

The opening examples establish how far the work can range: a synth plugin for Ableton, Premiere editing, games with separate art and mechanics work, and a self-driving golf cart. The workshop then follows six steps: give Codex context, build something, let it take actions, collaborate on complex tasks, make the work repeatable over time, and embed Codex elsewhere. Context is the first dependency because the useful specification may already live in a conversation, document or application.

An app shot captures the Slack window. It resembles a screenshot, but includes embedded text as well, so Codex receives both the visible conversation and its words. The accompanying request is short: “Hey, can you just build this?” The capture supplies what “this” means—the product idea and the surrounding brainstorming—while the request tells Codex to act on it.

Selected presentation frame from Cooking with Codex — Charlie Guo & Gabriel Chua, OpenAI at 462 secondsOpen full source frame
A chat message asks Codex to build from the captured conversation shown above it.

While the translator builds, another artifact demonstrates a different source of context. A repository has code, Git commits and a history of changes; Codex uses that material to produce a Sites page explaining what is new in the Agents SDK, with code snippets and changes such as sandbox agents, Codex as a tool, tool search and real-time agents. Sites provides a shareable reading format for knowledge that would otherwise require someone to inspect the repository.

0:240:40
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:18 · section reference included

Remote requests and plugins supply execution context

Codex Remote separates the place where a request starts from the machine that executes it. Charlie sends a request from the ChatGPT mobile app to a connected laptop or remote server: explain how the Agents SDK repository's primitives fit together, using the style of an IKEA instruction manual. The laptop remains on while Codex works. The phone becomes a way to give direction and check for blockers without sitting beside the executing machine.

Back in the translator build, plugins supply several kinds of reusable context and capability:

  • Skills: best practices, prompts and team-specific procedures that describe how to do the work.
  • Apps: configuration and connectivity to external services.
  • MCP servers: tools the agent can call to perform the work.

Bundling these pieces gives the agent both guidance and access. A development practice alone cannot retrieve information from an external service; a connection alone does not explain the team's preferred workflow.

The macOS plugin used for the translator includes practices for AppKit, Liquid Glass, signing and inspecting the built application. The OpenAI developers plugin can create an API key when needed, removing a manual trip to the developer platform from this workflow. These plugins address different dependencies: one supplies platform knowledge; the other handles service setup.

A successful conversation can itself become reusable context. After working through an unfamiliar workflow, Charlie asks Codex to “make this a skill” or package it as a plugin. That preserves the procedure for a later task, so the next run can invoke the learned workflow rather than repeat the whole discovery conversation.

9:179:47
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

9:17 · section reference included

The translator reaches a live German audio test

The translator is reported built in four minutes and two seconds. Gabriel selects German and configures audio input and output, then tries a microphone test. The first attempts require more audio adjustments. “You know it's a real demo, live demo, when you gotta do IT on stage.” Charlie suggests using local audio amid congestion in the room.

Selected presentation frame from Cooking with Codex — Charlie Guo & Gabriel Chua, OpenAI at 808 secondsOpen full source frame
The translator app is open as Gabriel prepares to test German audio input and output.

The next exchange produces the observable result. Gabriel says, “Hello, are you there?” The app responds, “Hallo. Bist du da?” It then translates his statement that he is at the AI Engineer Conference and his explanation that the microphone can pick up the speakers. The example has now moved from a Slack idea to an application producing German speech. This is a short live demonstration rather than a translation-quality benchmark, and the microphone picking up speaker output remains a practical limitation of the setup.

The causal steps are worth keeping distinct. The app shot carried the product idea into the task. Plugins supplied macOS practices and API setup. Codex produced a build. Choosing devices and speaking through the application tested whether that build worked in the room. A fast build still needed an actual input-to-output test; audio routing became the immediate problem once the software existed.

13:1013:40
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

13:10 · section reference included

Carry preferences forward, then review and share the result

Context can persist beyond an individual request. Memory, described here as off by default, can learn preferences such as using uv for Python package management. Chronicle is presented as a research preview for Pro subscribers that adds screen activity context: a question such as “Why is this failing?” can refer to the pipeline the person has been viewing. Custom instructions apply across conversations, much like instructions placed in a home-directory AGENTS.md.

Persistence has costs and controls. The personalization settings can skip memory generation in chats involving tools, because tool results may contain information the user does not want retained. Memory and Chronicle also consume additional token budget. The decision is whether the saved context is useful enough to carry into future work, rather than assuming every returned fact should become a lasting preference.

For context that has not yet been written down, dictation provides a low-friction route. Talking through an idea can both supply detail and clarify the idea before sending it. App shots serve the same purpose for material already on screen. Neither requires a polished prompt first; the important part is providing the information needed to identify the task.

Building also includes checking and distributing artifacts. Sites can package an event recap or a comparison of offsite locations so a team can make a decision. For the translator, /codereview starts a review against uncommitted changes. The repository visualization requested from the phone returns as an SVG, after which Gabriel asks to use the Image Gen skill. That exchange exposes a useful distinction: requesting a visual does not necessarily select a particular image-generation mechanism.

15:3116:01
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

15:31 · section reference included

Choose application control by the access the task needs

A dashboard export shows why an agent sometimes needs to operate an interface. The dashboard has date pickers and dataset controls; doing repeated exports by hand means pressing many buttons. Gabriel captures it in an app shot and requests two different outputs: commit counts for the last seven days as CSV, and pull-request data from the start of the month to the current day as JSON. Codex invokes computer use, inspects the application and begins clicking and downloading.

The two exports require separate choices of dataset, date range and format. Giving all three properties for each output makes the intended action clear. On macOS, the demonstrated computer-use interaction has its own cursor, so Gabriel can continue using the computer while Codex works in the relevant application.

The next request turns questions already drafted in Google Docs into a feedback form, with every question marked mandatory. The Chrome extension controls the live browser tab to create it. Having the question text is only part of the task: the agent must also enter the questions, configure their options and set the required state in the form interface. The form later reports completion while the workshop moves on to long-running projects.

Three control mechanisms serve different needs:

  • In-app browser: deeper integration with the browser engine supports inspecting layout, rendering and interactions in local web applications. It is described here rather than demonstrated.
  • Computer use: drives desktop or legacy applications through the interface a person would use, especially when there is no convenient programmatic interface.
  • Chrome extension: operates in the user's signed-in browser, with permission, so a task can use an existing session rather than start by trying to authenticate in a separate headless browser.

The choice depends on whether the difficult part is inspecting an application under development, operating software without an API, or accessing an already authenticated session.

Selected presentation frame from Cooking with Codex — Charlie Guo & Gabriel Chua, OpenAI at 1617 secondsOpen full source frame
A recap slide lists the in-app browser, computer use, and Chrome extension as application-control options.
21:2521:55
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

21:25 · section reference included

Long tasks need explicit goals and inspectable progress

The overnight projects cover three distinct kinds of development:

  • Greenfield application: a live Q&A site intended for a room of about 400 people, with admin, participant and onstage views, moderation, and a 16:9 stage presentation. It uses Convex for real-time interactions and is intended for deployment on Vercel. Its goal is still running after 15 hours, 30 minutes and 55 seconds when shown.
  • Repository adaptation: an existing open-source form builder becomes an internal Sites application using ChatGPT authentication to restrict access to the team. The reported run lasts five hours and 57 seconds.
  • Research implementation: papers become a Python package with a Rust backend, under constraints to use heuristics rather than dynamic programming. The reported run lasts seven hours, one minute and 40 seconds.

These are examples of work at different starting points, rather than interchangeable evidence that every large task is complete after a particular duration.

The Q&A specification names the users, views, presentation format and service choices. This is the useful part of prompt engineering in the demonstration: make the intended behavior clear. Gabriel initially talks through the idea on his phone, then asks Codex to polish the prompt for presentation. The information matters more than a magic phrase.

For progress that spans hours, Gabriel requests a series of goals in goals.markdown and a progressdashboard.html page. The filenames are his chosen artifacts, not a required format. They make the intended milestones and current work inspectable. Separate threads run code reviews and goal audits at milestones; an audit compares the work with the original goals and can tell the main thread to correct course or reconsider the plan.

Testing can require a machine the main task cannot access. The remote Q&A thread asks a local thread to inspect the application using the Chrome extension, then receives the testing results back. What crosses between threads is a task and its findings; browser access stays on the local machine. Slack updates provide another view of the work: what is done, what remains and which blockers need a person.

How does a remote build get browser feedback from a local machine? The diagram shows the handoff and return path. The remote thread can continue owning the project while the local thread supplies a capability unavailable on the remote host.

How it fits togetherRemote development delegates local browser testing

Runs the project and requests browser testing.

The testing request travels to the local thread; browser findings return to the remote project thread.

28:4529:27
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

28:45 · section reference included

Use side conversations to steer work and resolve blockers

A side thread lets a person ask about a running task without turning the main conversation into a status interview. /side opens a place to ask what happened in the last hour, what comes next and what needs help. If the proposed next step is wrong, the side thread can send a correction to the main thread. Side threads are ephemeral, so any decision that should last needs to be communicated back.

The form-builder project also demonstrates communication between two main threads. The remote task tests whether it can spawn a thread in the local project and use that location for future Chrome testing. The local thread then opens Chrome and creates a form end to end. This gives the remote implementation a functional check through the actual browser workflow.

Feedback can update an existing verification task. In the Retrodex interface work, Gabriel sends a screenshot and the blunt assessment “Not very good.” The main agent forwards the screenshot and directive to an already running subagent. The unwanted behavior is specific: messages are pushed to the bottom. The verification task then sends a short message to check the behavior. The important transition is from general dissatisfaction to a concrete interaction that can be inspected.

The progress dashboard shows milestones alongside thread responsibilities, including orchestration, communication, visual UX and the App Server protocol. It also exposes a genuine blocker: the remote Q&A work lacks authentication for Convex and Vercel. Gabriel asks a side thread to explain how to unblock it. A CLI login flow provides Vercel access, while a shell-based setup places the Convex deployment key in the remote environment instead of pasting it into chat.

The internal form builder returns as a working product flow: create a draft, preview it, publish a shareable link and view responses. Hosting it on Sites limits access to people who can access the ChatGPT workspace. Adapting the existing repository therefore changes both how the team distributes forms and who can use them.

34:5835:28
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

34:58 · section reference included

Separate continuation, delegation and deterministic checks

Compaction lets a long conversation continue across multiple context windows by carrying forward condensed context. Gabriel connects this capability to model training across multiple windows and offers user anecdotes about long threads. The practical benefit is continuity during hours of work; those anecdotes do not make compaction a guarantee that every important detail survives.

Goals answer a different question: when should the agent stop? At each turn, the task evaluates whether its success criteria have been met and whether it should continue. Those criteria should be specific and verifiable. Charlie's “monkey's paw” warning captures the failure mode: an agent can technically satisfy a request in a way the person did not intend. More time does not repair an ambiguous definition of success.

Delegation has two useful forms:

  • Subagents: keep task context separate and return work primarily to the main model. Named roles, developer instructions and model choices can be configured in subagents.toml; a reviewer might prioritize speed or thoroughness.
  • Visible threads: keep separate workstreams available for the person to inspect and direct. In Charlie's game project, art, music, animation and mechanics threads all check with a creative-director thread before considering their work complete.

The deciding factor is how much of the delegated work the person wants to follow. Separate context helps in both cases, but visible threads make ongoing coordination part of the user experience.

Hooks put deterministic operations at selected checkpoints. Configured through config.toml in the Codex home directory, they can send conversations to logging, scan inputs or outputs for secrets, inspect proposed tool calls, or run validation when a turn stops. These are distinct from asking the model in AGENTS.md to remember to run a linter: the hook places a programmed check at a defined moment.

The Agents SDK repository provides a simple example: a Python script tidies the repository at the end of each turn. Predictable housekeeping does not need another language-model decision. As tasks run longer with less direct supervision, hooks keep such checks attached to the workflow while goals define what completion means.

Selected presentation frame from Cooking with Codex — Charlie Guo & Gabriel Chua, OpenAI at 3003 secondsOpen full source frame
The Codex settings screen shows a configured repository hook.
42:2842:58
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

42:28 · section reference included

Timers resume waiting work and turn feedback into new tasks

Provisioning a remote host introduces waiting that neither the person nor the agent needs to watch continuously. The DigitalOcean plugin starts creating a droplet, then a heartbeat automation resumes the same thread in about five minutes to check readiness. Once the droplet is ready, the thread takes the next step and provides a deep link into Codex settings that preconfigures the SSH connection, creating an SSH key if necessary.

This is a timer-driven continuation: check a condition, wait again if needed, then proceed when the condition holds. The same pattern can monitor a deployment pipeline at five- or ten-minute intervals. Automations that create new threads serve another purpose: a weekly recap or a one-time analytics review can have its own conversation and be archived afterward.

The translator returns as the input to a development loop. A Slack channel contains demonstration feedback, including a request for a sound check before the audience sees captions. Gabriel asks an automation to read and reply every half hour, separating compliments, bug reports and feature requests. Then he steers the task while it is running: bugs and feature requests should also create a new thread in a worktree, address the issue, open a PR and request review.

What must happen between a Slack comment and a reviewable change? The diagram separates intake from implementation and review. This is the proposed automation workflow, rather than a demonstration of a completed sound-check feature. Further steps could build a test version, send it to someone and return their feedback to the PR reviewer.

Automations can also maintain the context used by later work. One shown automation checks the presentation through Chrome every five minutes and estimates whether the workshop will overrun. Other proposed uses review conversations for skills to improve or remove, or compare drafted emails with the replies the person actually sends. In the email example, a second automation writes lessons into a Markdown file that the drafting automation can use next time.

Charlie's own notification workflow reads updates from Slack, email and Linear and maintains daily notes and workstream notes in a local Obsidian vault. Later requests can draw on those notes to explain what changed since the last check-in. The timer therefore does more than repeat an action: it keeps a useful account of ongoing work available for the next conversation.

How it fits togetherTranslation-app feedback becomes development work

Comments include compliments, bugs and feature requests such as a pre-caption sound check.

The proposed half-hour automation routes fixable feedback into a worktree task and returns a PR for review.

50:4051:10
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

50:40 · section reference included

Retrodex puts a new interface on the Codex harness

The App Server protocol connects interfaces to the core Codex harness. The workshop describes it as underpinning the Codex app, CLI and VS Code extension, carrying capabilities such as tool calls, plugins, compaction and steering. An application built on App Server can let users work with their ChatGPT subscription and token budget. Capabilities that depend on the local computer, such as computer use, are not automatically supplied by embedding the core harness.

Selected presentation frame from Cooking with Codex — Charlie Guo & Gabriel Chua, OpenAI at 3581 secondsOpen full source frame
A slide describes embedding Codex into a product with the App Server protocol.

Retrodex is the fourth overnight project: a custom interface built on App Server. It offers model selection and reasoning-effort controls, with increasingly energetic visual effects as effort rises. Gabriel sends “Hello,” allows access and receives a reply. The attempted follow-up display of matching threads in the Codex app is inconclusive on stage; the demonstrated result is the request and response through the custom interface.

The interface is playful, but the integration point is practical. A team can put Codex inside its own product or internal agent platform while retaining the core harness capabilities. Gabriel suggests inspecting the open-source Codex repository and asking Codex to explain its App Server as a starting point for building such integrations.

58:3759:07
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

58:37 · section reference included

From an individual request to a software factory

The closing questions shift attention from a successful request to the process around it. What should a loop check at each turn? What stopped the agent: missing context, missing permission or an unavailable capability? Can the solution become a skill or plugin for the next run? How should work be divided across agents or visible threads? These questions connect the workshop's individual features to the recurring problems of development.

Selected presentation frame from Cooking with Codex — Charlie Guo & Gabriel Chua, OpenAI at 3812 secondsOpen full source frame
A closing slide asks how to scale work across agents and where human involvement is needed.

Human involvement also needs a reason. API-key creation once looked like an unavoidable manual step; the developers plugin can handle it in the demonstrated workflow. Authentication for the remote projects still needed Gabriel's intervention. The useful decision is to retain people where their participation is needed and remove routine handoffs that only interrupt the work.

The translator makes the larger ambition concrete. A conversation becomes a build; a live audio exchange checks its behavior; Slack feedback becomes input to a proposed implementation-and-review loop. Charlie calls this building “the machine that builds the machine”: a software factory that produces individual changes through a repeatable process. Its useful checks still have to be designed—success criteria, testing, review and the handling of blockers do not appear merely because the loop runs.

The recording ends by moving from the tour into a hands-on build session, with OpenAI staff available to answer questions and unblock work. That is the natural next experiment: pick a real task, provide its context, observe the result, and turn the parts that worked into a process worth repeating.

1:02:131:02:43
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

1:02:13 · section reference included

Resources

Read the complete timestamped transcript
  1. 0:18

    All right.

  2. 0:21

    Good morning, everybody.

  3. 0:23

    Morning.

  4. 0:24

    Uh, thanks for bearing with us. There was a long line outside, so we just wanted to make sure everybody could get in and get a seat. Um, thank you so much for being here. This is the Cooking with Codex workshop. We're so excited to have you. Uh, my name is Charlie. I'm on the developer experience team at OpenAI.

  5. 0:40

    And my name is Gabriel. I'm also on the developer experience team, uh, flying in from Singapore. Nice to be here.

  6. 0:49

    Uh, so before we kinda dive into the nuts and bolts, you know, I wanted to do a quick review of what have folks been cooking with Codex?

  7. 0:58

    It is a cooking show

  8. 1:01

    .

  9. 1:01

    Uh, these... You might have seen some of these tweets. Um, but you know, there's a real range of stuff that people have been making, some of it pretty, pretty incredible. Uh, I think this one, uh, a session with nearly 300 sub-agents. Has anybody beat that in a given session? No? I think, I think Dom might have, but we'll see. Um, Codex working on a synth plugin, uh,

  10. 1:31

    for Ableton. Um, really, really cool stuff. I think on our team, um, we have a, a team member, Brent, who's doing, at this point, video editing with Codex-driving Premiere, right? Which is not what you would normally expect to see from, from a coding app.

  11. 1:51

    Uh, playable, you know, games. I think, like, gaming to me is, is really a place where, where the models and the app can shine. The ability to go and check your work, um, for the game as it's done or to hand off between different threads and, you know, manage audio versus art design versus game mechanics versus, you know, difficulty balancing. Um, I think, I think the different threads and sub-agents are amazing here.

  12. 2:18

    And last but not least, um, you know, building a self-driving golf cart with Codex, uh, which I think is just absolutely incredible. Um, it kinda makes me wish I was back at school so I could be hacking on, uh, stuff in my nights and weekends again.

  13. 2:38

    But where have people been cooking? I think that's pretty fun. So you s- we sort of see it everywhere. Everyone's Codex mixing everywhere, anywhere, all the time. I mean, literally in between meetings. You've seen people, like, holding their laptops everywhere, keeping it open. It's become a meme at this point. Fortunately, we now have Codex on mobile. So this, uh, gentleman in Japan was able to, like, bounce between ideas and bouncing between cities while on the bullet train or while walking a dog.

  14. 3:09

    And you don't just ha- you don't necessarily have to use Codex on the Codex app or use it within, like, the ChatGPT app. You can just put Codex anywhere. With it being open source, there being the app server, which we'll talk about briefly at the end, you can put it here. So this, Daniel, for example, put it into three E Ink devices. So you really can Codex mix anywhere. So if it's your first time sort of using Codex, uh, do check it out. Do give it a download. What's really nice is that if you have existing

  15. 3:38

    configuration for other agents, you can just bring it in and just get started really quick. And today we wanna, through the cooking show, help you cook really fast and cook something really cool. So as a quick agenda, we're gonna talk about what's new, and then the real, pun intended, meat of the session is the six ways to cook with Codex. Some closing thoughts, and in the second hour of the session, we'll have, like, a mini hackathon where we're gonna have live cooking right in this room, and we'll be, like, walking around, helping to

  16. 4:08

    be your sous chef to your main chef. So what's new? What's fresh from the oven? A few... A bunch of things, I think, fundamentally to just really help you give context to Codex really fast. There was record and replay. You spend time maybe describing, "Oh, you gotta press that button. You gotta press this button." Maybe just show Codex by doing it. It builds on sort of our technology with computer use and so forth. Thread handoff, really fun. You can sort of push things from the cloud to the remote to your local and vice versa.

  17. 4:38

    We'll demonstrate some of that. And the mobile app is now generally available, so now there's, like, sort of one-to-one device pairing, a much more secure way of working with Codex. So it's on- it's been 2026. This slide barely summarizes everything we've shipped, and we're only 50%, literally 50% of the year. We- some things have been really fun, the plug-ins, computer use, multimodality, being able to use Codex on a mobile, app shots. That's a personal favorite. These things really make

  18. 5:07

    Codex and the Codex app really fun to use, and we'll be talking about some of these. Uh, they are our sort of key ingredients. So that's cooking with Codex.

  19. 5:18

    Uh, we wanted to, like, start with this quote where it's from our favorite movie, Ratatouille. "Anyone can cook, but only the fearless can be great." And today, with the sort of ingredients we're gonna share with you and some of the recipes, but with your ingenuity, we think you're gonna cook some really fun stuff. And these are the six steps to cooking with Codex. Number one, you gotta give Codex your context. There's a lot of context in, like, your documents, in your brain, in your conversations, in your email, in Notion, Linear,

  20. 5:47

    so forth. You gotta give Codex that context. That's a really important ingredient. And then once you have that, you can literally just build things, and that's already just step two. For the power users in the room, there's some interesting stuff where Codex can actually take action. You might notice right now, like, on-- I'm on Chrome, and you just see this thing, nice thing at the top saying, "Codex started de- debugging this browser." What's happening in background is actually Codex is reading the slides as I'm speaking. And then h- la... You can sort of collaborate with Codex on long-running tasks, and you wanna work over time where-

  21. 6:18

    We're gonna put this, make it repeatable, sort of ensure consistency, and really, like, build those loops. And finally, we're gonna bring Codex everywhere. If you wanna embed it in your own product or in your own internal platforms, that's also possible. So let's just dive right in for the first two steps. Cooking with Codex, where you give it your context, and you can just build things. So o- a few days ago, like on Friday, uh, Charlie and Christine and myself, we were thinking, "Yikes, it's, it's Friday. It's Friday night." We sent it at

  22. 6:48

    4- 4:44 PM. "We don't have a demo idea for this workshop. Uh, we have any suggestions?" Then Christine suggested, "Hey, let's build a real-time translator." Charlie, uh, suggested, "Hey, let's use GPT real-time translate." It's one of our latest models. It's really good. And then we can actually even have a plugin. I've never built a macOS app, but with the plugin, all the best practices are there. So maybe you would, like, copy-paste, you would, like, uh, Google or search for what do these things are, then you go to Codex. But I'm gonna do something fun. So I'm gonna take my two thumbs over here, my l- my right thumb and my

  23. 7:18

    left thumb. I'm gonna put it on the Command key. I'm just gonna press it. And then that's an app shot. All that context on Slack, all the conversation, that brainstorming is just right there, and let's just ask Codex,

  24. 7:32

    "Hey, can you just build this?"

  25. 7:38

    Yeah. Sh- Cooking Show's over. No, I'm kidding. And we're gonna let Co- Codex cook on that. And what happens with an app shot is that when you take a s- it's, it's like a screenshot, but all the text is embedded as well, and it's very nice because that's where you put all personal context with conversations. There's brainstorming. You don't have to, like, repeat yourself. You just give Codex anything, it's like, just figure it out. So we're gonna l- look, let, let that work on it for a bit, and we're gonna try something else as well. So we have the Agents SDK. It's open source. It's pretty nice. And, uh, you got,

  26. 8:08

    like, some things that have come up in the recent months. So I was thinking, "Hey, maybe we could build, like, a website to sort of just briefly describe what's new," and we're gonna use this thing called Sites. So Sites is something that we launched earlier in the m- in June for in preview for enterprise and business plans. It lets you sort of create a simple site, and it's really about just creating artifacts that you can share. So here we sort of worked with Codex for a bit, and let's just see what Codex sh- cooked here. So you're gonna open it up, and it's all what's new with the Agents SDK. You have the... And it's pretty nice. You have, like, six changes, sandbox agents,

  27. 8:38

    Codex as a tool, tool search, real-time agents. We'll be talking about that in a f- in the next few days as well. And then you can see, like, the code snippets are here. So yeah, it's pretty nice. With Sites, you can then create artifacts that you can share with your team. So that's another example of bringing context because you have a code base. It's a really large code base. There's Git commits. There's history to it. There's a lot of things. There's that context, and that's step one. And you're just gonna give it to Codex to build it, and that's step two. You can just build things. So let's give them that last, last example here. Maybe, like, now wh- uh, Charlie's out on the streets, th- uh,

  28. 9:08

    taking a walk with his kids, his dog, and he's wondering, "Hey, maybe he wanted to build something." So I'll let Charlie, uh, take the stage here.

  29. 9:17

    Got it. Cool. Uh, so what I wanna show you is, uh, Codex Remote. Gabe mentioned we might be out in the world, right? We might be, you know, on the beach. We might be, um, camping, something, but we still wanna get work done. We still wanna keep cooking. So Codex Remote is a feature that is now available in ChatGPT, a mobile app, and allows us to connect to a, you know, laptop or remote server and to issue commands that can be executed, uh, remotely. So in this case, let's say we wanna visualize something about this Agents SDK, right? Um, for me personally, you know, last weekend,

  30. 9:47

    I was building a bunch of IKEA furniture, so that's kind of on the brain. Uh, why don't we try to visualize how this repo works in the style of an IKEA manual?

  31. 10:00

    Can you create a visualization of this repository in the style of an IKEA instruction manual? Uh, I wanna understand how all of the primitives fit together.

  32. 10:12

    So this is gonna send from my phone.

  33. 10:15

    And here we go.

  34. 10:20

    So you can use Remote to keep working. Uh, you know, you leave your laptop on. Codex will keep the, the laptop open while it's cooking in the background. Um, and it means that, you know, you can sort of free yourself from having to be at the computer all the time or having to walk around with your laptop open all the time, which is 100% a thing that I've been doing lately.

  35. 10:37

    It's almost like a regular occurrence. Every weekend, everyone's, like, leaving their laptop in their room at home, and he's, like, walking out, touching grass, you know, still checking in, "Hey, Codex. Is everything okay? Is there anything I can do to help unblock you?" And in a sense, everyone's become a manager. So let's see. Go back to the live translation. We'll see what's going on over here. All right, it's building. It's d- some web search. It's finding out what's going on. So I think we've sort of covered very briefly, uh, sort of another concept called plugins. So just going to the slides. So you go there.

  36. 11:08

    Oof.

  37. 11:14

    Oh, here it is. Yeah. So plugins are a bundle that sort of bundle together skills, which are best practices, external prompts, uh, maybe practices that are custom to your team, apps which have the configuration and s- and the, provide the connectivity to external services, and MCP servers, which pro- provide the tools that are necessary. And wow, Codex already sort of popped it up there. We'll just skip that in a bit. So that's plugins. And if you just take a look at some of the plugins we were

  38. 11:44

    using here. So if you just go over to the Codex app and you just scroll right up here, there's this button called Plugins, and you'll see a bunch of plugins. So for example, what we were using was the built macOS plugin, which has a bunch of best practices, uh, in terms of the skills for how to use AppKit, how to use Liquid Glass, how to sign and inspect the built application. And we were also actually using the OpenAI developers plugin. So it's 2026. Um, we're not gonna go to the developer platform and click, uh, create

  39. 12:14

    API key, save it somewhere, pray that you actually save it somewhere. Now, with Codex and the plugin, it can just directly create the API key for you as necessary, and it really keeps things really smooth as you... It doesn't really interrupt you in that flow. So those are the sort of two plugins we are using, but there are other plugins that may be relevant like, say, uh, Gmail, Outlook, uh, maybe you use Microsoft Teams, you use Slack, and it... That provides all the relevant connectivity you need for your agent to work to bring context into the kitchen.

  40. 12:45

    I think, I think I'd add to that. You know, the Codex app, um, comes with the skills to build your own plugin. So in giving the app context, a lot of times I will work through a difficult problem or a workflow that, you know, I'm not sure how... I wasn't sure before how to tell Codex how to do it, and when I get to the end of that conversation, I will often say, "Hey, make this a skill," or even better, "Make this a plugin so that to- next time I need to do this, you can just invoke the plugin and get the job done, and we don't have to step through all of this work again."

  41. 13:10

    So this is what Codex built in, say, four minutes and two seconds. Let's see if it works. So it's a live translation. So let's, let's maybe do, like... I, I understand there's some German speakers in the room, so let's go- let's do German, and let's just make sure the audio is, uh, working as well. And let's see what Codex can do. So again, I've never used... built a macOS app. I'm a data scientist by training, not very- not the best at, like, sort of working with, like, uh, WebRTC, WebSockets. But with the Codex app and the plugins,

  42. 13:40

    I'm able to get that sort of best practices through the pl- through that. So let's just do that, make sure the input, we change to the phone. The output is here. All right, fingers crossed. This is a live app. It's, it's built. So this is, like, we don't know how this is gonna taste, and fingers r- really, fingers crossed.

  43. 14:10

    Oops. Hello there. Hello. Mic test, one, two, three.

  44. 14:19

    Okay, let's change the audio a bit.

  45. 14:28

    You know it's a real demo, live demo, when you gotta do IT on stage. It's real.

  46. 14:36

    Mic test, one, two, three. Are you there?

  47. 14:40

    A lot of congestion in the room. I would suggest local.

  48. 14:43

    Yeah.

  49. 14:46

    Thanks for the tip.

  50. 14:49

    Hello, are you there?

  51. 14:51

    Hallo. Bist du da?

  52. 14:53

    So right now I'm at the AI Engineer Conference.

  53. 14:56

    Also, im Moment bin ich auf einer KI-Ingenieur Konferenz.

  54. 15:01

    And I'm, like, winging this live. I have really... It's actually not the most ideal EV setup.

  55. 15:04

    Und ich improvisiere das live. Es ist eigentlich nicht das ideale EV Setup.

  56. 15:10

    Yeah, 'cause, like, actually right now, the speakers, my mic can actually pick up the speakers.

  57. 15:13

    Ja, weil ich gerade jetzt die Lautsprecher... Mein Mikro könnte tatsächlich die Lautsprecher aufnehmen.

  58. 15:19

    How was that?

  59. 15:20

    Great.

  60. 15:20

    All right.

  61. 15:31

    So one thing is that context is not just about best practices or external information. There's a lot of context in our heads. Whenever you talk to Codex, there's a lot of im- a lot of things that Codex can learn about you, and that's when memory comes in. It's really useful. It's... I think about it as bringing context over time. So it is off by default, but if you go to Codex app, you can turn it on. And let's say, like, for example, I use UV, and then Codex will learn that over time, that I prefer using UV for Python package management. But memories and context isn't just

  62. 16:01

    textual. It's also about what we're doing, what's on screen, and so forth. And we have a research preview called Chronicle for pro subscribers, also worth checking out. And with that, you could literally, for example, be on, like, GitLab, GitHub, looking at your pipeline, it failed, and you just ask Codex, "Why is this failing?" And it sees, ah, you were looking at your GitLab, your GitHub pipelines, and then it'll sort of figure it out. So to sort of set that up, it's quite straightforward. You can just go to the Codex app. You go into settings, and then it's called personalization, and you can enable it. So we can enable

  63. 16:31

    memories. And also, like, sort of given the, um, as a safety feature, you can also sort of skip, uh, generating memories in chats that involve tools because sometimes the tools might return information that you might not want persistent over time. So that's also worth considering. W- right now I'm using a business plan, so that isn't Chronicle. But if you're on a pro plan, you will see Chronicle as an option to enable here as well. And then there's this thing called custom instructions. So here you can sort of ena... It's, it's like setting the agents.markdown in the home directory.

  64. 17:01

    It applies to all conversations. So for example, I said here, "Always reply like a pirate." Then in the subsequent commentary, Codex would always reply like a pirate, a trivial example. So that's context, and to s- summarize, I'll let, uh, Charlie, uh, sort of summarize our first two dishes of the day.

  65. 17:22

    So to recap, uh, part one, right, we wanna give Codex content- context. Um, you need to know which ingredients you're actually using, um, for your recipe. Um, that includes bringing in external content via plugins, right? Uh, me personally, I'm a pretty heavy user of Slack, Gmail, Linear. Um, and then it makes it really easy to just say, like, "Hey, go look at my, you know, unread Slack messages for today and create a daily agenda for me to, to think about when I get into work." Um, you wanna be able to, like, keep the flow when you're adding your own individual context for the task that you're

  66. 17:52

    working on. Uh, I love dictation. I think just being able to, you know, as we've been demoing up here on stage, say, "Here's the context. Um, go do it." I think dictation to me is a really underutilized feature because we speak so much faster than we can type, right? And these days, you know, we're, we're beyond the, the era where you have to make sure every single word is correct when talking to a large language model. You can just give it reams and reams of your own rambling, uh, in transcription form, and it'll generally figure out the right thing to do, right? And you're able to

  67. 18:21

    actually talk through the idea and, and get to a better space by the time you send it to the model. Uh, likewise, app shots I think are, are incredible and also very heavily underused. Um, the ability to just, like, from any app... I think I, I don't copy and paste, you know, anything from Slack anymore. I just app shot my, my Slack window, or I just app shot my browser and send it directly into Codex, um, and tell it, "Go figure it out," right? And, you know, 90-plus percent of the time it can figure out the right thing to do. Last but not least, we've got personalization.

  68. 18:51

    Um, I think, you know, Agents MD, uh, which hopefully folks know and love, uh, and then Memories as well, uh, in the Codex app via Chronicle. Um, worth noting, I think Memories, uh, and Chronicle will use up a little bit more of your token budget, but, uh, it can be a really magical experience to have the app just kind of understand what it should be doing based on conversations you've had it in the past. It really levels it up and makes it start to feel like a true collaborator. Um, there's a ton more stuff, I think both in like contacts and personalization in the Codex app. Strongly recommend you go check it out.

  69. 19:21

    Um, Pets is like a universally beloved-

  70. 19:24

    Yeah

  71. 19:24

    ... uh, personalization feature, which we're not really gonna showcase here today, but, um, is... would, would... Yeah, definitely go make your pet.

  72. 19:32

    Cool.

  73. 19:34

    And once you have all the ingredients, it's... You- you're here to build. And to really help you build, and you wanna distribute it, Sites is a great thing. At OpenAI, we've been using Sites extensively. Uh, I, for example, have been using Sites to summarize like an event, like an event after... like d- like after an event like today's, I would ask, "Hey, Codex, can you summarize what happened today? Create a site, and then we can share it with the team, the rest of the organization." I create a site to sort of... Uh, we, we were planning an offsite, and then, uh, uh, Corey, uh, Corey on the team created a site to help us plan,

  74. 20:04

    choose different locations, understand the trade-offs. So you can create artifacts to help you make decisions. Specific for code, you can even j- you can also hel- it helps with code review. So let's just qu- give a quick example. With the live translate what we had just built, you just do /codereview, and then you can review against uncommitted changes, and then voila, you're firing off a code review, uh, task. It's a bit like having that, uh, you know, co- sous chef, like, double-check everything. You know, you see in the movies, right, before the chef, like, send, sends out the dish, they always do a quick taste test. I think that's a bit like that.

  75. 20:34

    Image Gen, where you can create images, and let's see what the example, uh, Charlie kicked off created. Oh, okay. It created SVG, but not too bad. Not too bad. So then what you could also do is, uh, use the Image Gen skill and ask Codex to sort of cook it. So we'll come back to that in a bit, but let's try that. Let's use this.

  76. 20:55

    And Spreadsheets, Document and Slides, we won't be talking much about it today, but you can sort of build these artifacts as well, and in the subsequent talks in the rest of the days, like the AI engineers, uh, Jason, for example, will be sharing more about how he's sort of building with Codex for sort of these, uh, deliverables. So those are the first two steps, c- context and building. But it doesn't stop there. So if you're sort of new to Codex, it's your first time, so you're still using coding agents, I think that's really helpful, but I'm, I'm pretty sure this is AI Engineer World's Fair. We're

  77. 21:25

    in a room full of power users. So for the next four steps, we hope we're gonna share something that's practically useful for you, and the first one is computer use. So Codex can sort of take actions, like as you see right now in the browser, Codex is debugging something. I won't tell you what's happening. I will, we'll reveal the surprise at the end. But it's sort of working in the background, thinking about something. So let's give an example. So I, I used to work in a large organization before OpenAI, and not everything had a API. You had a lot of enterprise software that had like... they were legacy. There were different

  78. 21:55

    dashboards. You wanted to download data. So let's say for example we have this dashboard that somehow pulls data from the OpenAI Python GitHub repo. There's this nice date range picker, but i- imagine I had to cl- get data for y- every month from the start of the year, the year to today. I would press Monday to Sunday, Monday to Sunday, Monday to Sunday, Monday to Sunday. It's a lot of buttons you have to press. Then there are different data sets you could choose. So what if Codex could help me with that? And the answer is yes. So again, I'm gonna take an App Shot, my

  79. 22:25

    two thumbs, and then just take it. Okay, let's just make sure we're in a new conversation. And we'll just ask, "Hey, um, can you help me download the, uh, number of commits for the last seven days as a CSV and the number of pull requests, the pull request data as a JSON from the start of the month to today?" Uh, yeah, use like computer use and, uh, go for it. So again, it's like two different dates. I'm asking it for the, like

  80. 22:55

    Monday to Sunday, like all the last seven dates, days, but also from like the 1st of June to today, which is the 29th of June. And you would see, okay, Codex has, uh, correctly invoked the computer use skill, and very soon it's gonna like start looking at the application. Um, you see, uh, here that it looked at application, and very soon in the chain of thought summaries, you would see more of it, and very, very soon, again, my hands are not on the keyboard, and you can see this is my cursor. I'm just gonna like park it over here right next to enterprise. Uh, you will see another

  81. 23:25

    cursor pop up.

  82. 23:26

    There it goes.

  83. 23:27

    And yeah, there it is. It's, it's moving. It's downloading it. Again, my hands are here. Charlie's hands are here. My phone is, like, on l- is locked, and Codex is downloading all this data by itself. It's a bit like a mag- It's, like, pretty magical the first time you see it. And that's... We're gonna let Codex, uh, work in the background. That's, that's the beauty of computer use of Codex. You don't have to wrestle with the agent for control of your screen. You can let it work in the background.

  84. 23:48

    I, I think it's also worth pointing out that on macOS, you know, it doesn't, uh, steal your cursor, right? Like Gabe's still opening the app. He can still use a different tab in Chrome.

  85. 23:56

    I can go scroll on Twitter and something. Yeah.

  86. 23:58

    Yeah. And so you can still drive your computer while things are just happening, you know, in the apps that Codex needs access to.

  87. 24:04

    So we're gonna let that work. But sometimes not everything is like a native application, and there's someti- sometimes, you know, like maybe using Google Drive, you're creating a feedback form like for today's workshop. You know, I mean, feedback forms are, you know, it's... I, I personally find it very cumbersome to create a feedback form. You need to like, uh, create a question. You need to like fill in the different options, and I always forget to mark it as mandatory even though I want it to be mandatory. And it's... You press so many buttons with creating these feedback forms. So I already have the questions drafted here, and

  88. 24:34

    let's see if Codex can help me with that. So I just create a new thread. So again, App Shot.

  89. 24:42

    It's quite addictive, I, I must say, the user interface. And we'll just say, "Hey, um, can you create these, uh, this feedback form for me using the questions in the Google Docs? Um, remem- remember to make... mark each question as mandatory." And just to give Codex a bit of a h- I mean, just to make things for exposition purposes, what we're using here is the Chrome extension, which is slightly similar to computer use. So that's, this is computer use. But here it's directly controlling a specific

  90. 25:12

    application, which is Chrome. So we're gonna fire that off, and we're gonna see very soon another cursor come up. So we're gonna split these two side by side over here, and you would see Codex thinking about the task. It realizes, hey, I have the quest- form questions already, then it still needs to drive the live tab to actually create it. And again, like what we saw earlier, it's gonna sort of press start pressing buttons. So let's just wait for a bit to see the first, uh, button get pressed, and then we'll sw- swap back to the slides in a bit. So, ah, there it is. You see it again. My cursor is

  91. 25:42

    here, and then you see the blue one's here. So you can continue to go do different tasks. Let's say, like, maybe I'm interested, uh, to s- learn something about the Codex repo. We were talking about memories. So I can just, like, ask, "Hey, um, can I configure the model that's used for memories in Codex, and what's the config, uh, term I need to change?" So you see, while you're continuing to do other tasks, Codex in the background is sort of doing the button pressing for you, and we'll fire that off. So that's computer use, and there are, like, different

  92. 26:11

    flavors of it, like, you know how maybe, you know, your favorite dishes come in different flavors. And I'll let Charlie just really give a nice summary of what, where the different flavors are.

  93. 26:24

    Yeah. At this point, there's, there's, like Gabe mentioned, a few different ways that Codex can start controlling the applications that you're using. Um, and it's kind of important to understand, you know, when you should reach for each one. The three that we currently have today are computer use, which you saw, uh, the Chrome extension, which you also saw, uh, and the in-app browser, which I think we did not demo. Um, the in-app browser, why don't I start there, is a feature where in the Codex app, you know, it has, it's the ability to open, uh, a browser tab on the side just like you would kind of a, a normal, um, you know, desktop coding editor these days. Uh, but in addition to

  94. 26:54

    just opening and rendering web pages, uh, Codex has really deep hooks into that browser engine, uh, deeper than it has with the Chrome extension. Um, and so as a result, uh, I think the in-app browser is really great for when you're building local web applications, and you wanna do a lot of deep testing, um, on the interaction design, on how it's rendering because Codex can see much further into the, uh, into the, the layout at that point. But going back to the two that we talked about today, um, we have computer use, right? Which is great for, um, I mean, anything

  95. 27:24

    on your, on your laptop or your computer. Um, but it really shines when you have applications that otherwise don't have a programmatic interface to them, right? Um, it's great if it's a web app and you can reverse engineer the API calls. But, you know, there's a lot of desktop software out there, there's a lot of legacy software out there which doesn't easily expose that kind of interface to an agent. And so if you just need Codex to drive the app in the same way that you, a human being, would, computer use is really incredible and kind of unlike anything else that, that I've tried. It, it, to this day, it's still just so

  96. 27:54

    magical to watch those little cursors fly around. Um, then there's the Chrome extension, uh, which is your sign-in browser, and I think that is the most useful for when you need to hand off, um, state or authorization to the agent. So I think we've all, you know, probably run into the issue where we wanna build an agent to go and, you know, use the World Wide Web the exact same way that we do, and the first thing that it immediately runs into is, like, oops, I can't log into your Gmail account. Um, and worse than that, like, Google has detected that I'm on a headless Chromium and is just sending me this infinite CAPTCHA loop,

  97. 28:24

    right? Uh, and so the Chrome extension allows Codex, with your permission, to go and read from certain websites as you, uh, and borrow your credentials so that you don't have to worry about the authentication step as you're building.

  98. 28:38

    Cool. Um, so that's Codex take action, and then next we're gonna move on to collaborating on complex tasks.

  99. 28:45

    So this one's a really fun one. So I think people have really talked about how Codex is really good at, like, complex refactors, long-running tasks. And in the sort of next 15 to 20 minutes, I wanna share some tips on how you can sort of work best with Codex on these complex tasks. So it's gonna be a bit hard to, like, kick off a 20-hour long-running task, and we're all just gonna sit here, enjoy lunch, bre- dinner, and breakfast. So I started off some tasks the night before, and there are, like, four different tasks, each

  100. 29:15

    covering a slightly different flavor. So this one is still running for, like, 15 hours. It's running on my remote, uh, uh-

  101. 29:23

    You can-

  102. 29:24

    ... instance.

  103. 29:24

    You can see at the bottom for-

  104. 29:25

    Yeah, over here

  105. 29:26

    ... familiarity, yeah.

  106. 29:26

    Like, I'm running a goal.

  107. 29:27

    15.

  108. 29:27

    It's been running for 15 hours, 30 minutes, and 55 seconds. So I gotta scroll, scroll, scroll, and here I am, right at the top. So it's a bit of, like, a greenfield project. I wanna build a live Q&A site. Uh, it's gonna be in a room for, like, 400 people. There's, like, a admin view. There's a participant view. There is a onstage view. And then, like, I describe, oh, maybe the admin can do, like, AI-generated text. There's some moderation features. On the stage view, there's a 16 by 9. Uh, run every application

  109. 29:57

    through our moderation API. Make the experience feel like something OpenAI would launch, so on and so forth. Uh, here I'm being particularly prosaic in the prompt for exposition purposes. You could just ramble to Codex as well, and it'll get, it'll get the job done. But here, just for, like, e- exposition purposes, I know people like seeing me take photos. This is, like, an example of how you're being clear. I think one misconception these days, like, people always ask, "Does prompt engineering matter?" Insofar as the instructions are clear, it still matters. Like, I can't read

  110. 30:27

    your mind. Codex can't read your mind. So you just need to tell Codex very clearly what you want. But it's... Prompt engineering, it, where you see as, oh, I gotta use this magic phrase, I don't think we're in that state anymore. Reasoning models are really intelligent, and these sort of, the real thing is about context. So it actually goes back to our first step, giving context clearly. So how I wrote this prompt was I was just talking to Codex. I was ra- I was yapping it on the, ya- yapping to it on the phone. I just, "Hey, Codex, can you, like, polish this prompt so I can show it on stage?" So yeah, it's building this thing, and I'm using, like, a few different things, like Convex. There's a, a plugin for

  111. 30:57

    Convex. It helps to do, uh, real-time interactions. Gonna deploy on Vercel, and of course, the OpenAI, uh, plug- Plugin. And all right, the feedback form is done. So that very nicely came up. And yeah, that's one example. So it's like a greenfield project. But, you know, not everything is like you're gonna start from scratch. So here's an- another example that ran for, uh, five hours and 57 seconds. So here it is an open source repo for a form builder, and I wondered, hey, could I, like, repurpose this and build it internally in Sites? So maybe you

  112. 31:27

    don't wanna buy SaaS. It's a small team of, like, seven of you. You wanna just have a internal form builder. You wanna use the ChatGPT authentication that comes with these sites to sort of ensure that only people within your team use it, and then you gotta let Codex sort of build it. And I, I will show you an example what that looks like later. Then the last example of a long-running task that was running, this one ran for, uh, let's see, I think seven hours, one minute and 40 seconds. So it's, like, some, uh, gradu- uh, research papers from my graduate school days. Uh, it was, like, written

  113. 31:57

    in, like, Julia, then I think, like, R. And I was like, "Hey, can I, like, write a Python package with a Rust, a Rust back-end?" And then I told Codex again, like, give some constraints. I don't wanna use dynamic programming. You gotta use, like, some heuristics. And then Codex read the papers and did it. So those are the three long-running tasks, and I wanna, like, give some examples and some of the best practices that come with it. So in one of these, for example, what I told Codex was this over here. You can define a goal. So there is actually a skill in the

  114. 32:27

    OpenAI skills repo called define goal. And I told Codex, "Hey, don't create one goal. Create a series of goals and put it in a goals.markdown file." Again, it's just very arbitrary, but it's just a collection of files that you sort of see, and you can, you can inspect it. And then create a dashboard, progressdashboard.html, so you can go back to it and see what's going on, because in this seven hours, you don't really know what's going on. You wanna have like some, some way to, like, see what's happening. And rather than having this one thread do all the work, it's actually spawning new threads. So

  115. 32:57

    if you go to the... See, I've pinned the project. But if you go to the, the thread, but if you go to the project, you will see it's, it spun many other threads. So for example, at different milestones, it ran co- co- code review, and then it also ran a goal audit. So it's seeing, like, hey, are you on track? There's this initial thing you talked about, like goals.markdown. Are you deviating from it? If it is, then it will nudge back the main thread, "Hey, uh, remember this," or, "Do we need to rethink the plan?" And so forth. So it, it ran at different times. That's why there's, like, different milestones, M2, M3A,

  116. 33:27

    and so forth. Uh, there's another example here with, like, the live Q&A site. Uh, this is still running. So this is running on a remote session. So it's, like, my [REDACTED] remote. And I wanna, like, do testing. So what I told Codex is that, "Hey, you can spawn a local thread on my machine and use the Chrome extension to then inspect it." So you would see, for example, this is a thread that came, and you will see it's, it, it says here, "Sent by Codex from another thread." So the remote thread can communicate with the local thread and says, "Hey, you gotta do click ops. You gotta do

  117. 33:57

    browser. You gotta do QA testing." Then the local thread did all the QA testing, then it send the results back to the remote thread. So that's another thing that's, uh, worth, uh, thinking about for these long-running tasks. And let's see what else we can call out here in the long-running tasks. Ah, communication. So one other way of thinking about this is that maybe you could just ask Codex to send you Slack messages. So here I have a few, uh, demos. So I was building a live Q&A site. And without having to set up additional

  118. 34:26

    infrastructure, you can just use the Slack plugin and just say, "Hey, um, reply to this thread and give me, like, a summary of, uh, what's done, what's not yet done, and are there any blockers?" 'Cause sometimes there are blockers. Maybe there's, like, a authentication issue. There's a API key that you need to give to Codex and so forth. So I asked Codex to, like, speak to me, and over the weekend, like on Sunday while I was out with my friends, I was just regularly checking, is there anything I can help unblock Codex for these different tasks? Um, let's just give some screenshots of what that process looked like. So just take a look at it.

  119. 34:58

    So another nice feature with Codex is that while you have the main thread, so right here in, you see in the screenshot on the right, left is the main thread. On the right, you can start a side thread. So I- you literally see here and it says, "Let's focus on signing in with ChatGPT first. I wanna sleep soon. I assume you need me to do that." 'Cause I know, like, okay, I think Codex needs probably some help with that, because I'm running a, uh, building a new instance on the app server, a new user interface, so it needs to get some credentials. So it needs to know where my credentials are. And it says, "Okay, actually, the credentials are

  120. 35:28

    already here." So it say, "Okay, update the main thread." So you can use the side thread to update the main thread. So let's do an example of that live as well. So here, this one is still running. You could just do /side, and it creates a side thread. And you can just ask Codex, like, what's going on? So maybe, uh, we just ask. It takes a while, this, I think due to network, but let's say, uh, "Hey, um, what's going on? Can you list, like, uh, what you've done so far, maybe in the last hour? What do you plan to do next? Is there anything I can do to unblock you?" And then from this, when you

  121. 35:58

    get, uh, the agent tells you what it's gonna do next, you can also get a sense of, like, what... Maybe it's, like, on the wrong, wrong track. Then you can sort of tell Codex, "Hey, maybe you don't want to do that. We should do this instead." Then the side thread can communicate with the main thread. So that's a pretty nice, uh, feature.

  122. 36:12

    It's worth noting with side threads, those are ephemeral, so if there is any context that you wanna keep long-term in your project, make sure you're communicating that back to the main thread because they will get stale after a while.

  123. 36:23

    Yes. Uh, here's another example where what we were doing earlier was a side thread talking to the main thread. This example is where two different main threads are talking to each other. Specifically, on the right s- uh, the left side, um, we have Co- uh, Codex saying, "Test spawning a local thread in the local form builder project. If it succeeds, make that the default location for all local, future local threads, especially for Chrome or end-to-end testing." So again, on the... What's happening here is

  124. 36:53

    that it's running on a remote service, but so it doesn't have access to my Chrom- Chrome browser. Then it send a message to the local host, and then it can then from there, uh, see, uh, Chrome appears to be installed, and then from there it can do the testing with Chrome. So there you can get the remote session to work with the local session. And this was the test that it did. So while it's working in the cloud, locally on my laptop, like maybe like 3:00 AM while sleeping, Codex opened Chrome. It went through the form-building flow, and it created a form end to end.

  125. 37:23

    Subagents are also a particularly helpful thing in these long-running tasks because you wanna make sure context is well separated, uh, agents sort of focus on different tasks. And w- here's one nice example. So it was building a application, and I just, like, took a peek at it, and I just... n- not necessarily best practice, I just sent it to the main thread saying, "Not very good." Yeah. And then Codex said, "Yeah, agreed. That's technically cleaner, but it's still a bad chat interface." And you notice here what it did was that it messaged the

  126. 37:53

    subagent. So the subagent was already pre-spawned, but then the main agent s- messaged the subagent. And you would see this is the message over here. So for example, it says this. So it sent it, and it says, "Update your verification task with this new user screenshot and directive." So there you have this sort of interagent communication. Uh, so this was what, what it did thereafter. So it did its own tests 'cause I sent a screenshot with the messages flushed to the bottom. I didn't want that. So then Codex did its own test, and it says, uh, "Send Retrodex visible

  127. 38:23

    okay in one short sentence," and it sent it. So Codex itself did its own QA testing to sort of verify my feedback. Um, and this is an example for the live QA where the main thread was spawning different other side threads. For example, the audit of the goal, the code review, but also doing that communication back to Slack. So this thread over here where it says, "Provide Slack thread URL," that's over here where I was providing all these updates. So a separate sort of session was managing that. We

  128. 38:53

    earlier talked about goal, and here is a dashboard that, you know, Codex created for the goal. So you would see it has all the different milestones, what's been completed, what's active, what's yet to be be- started. And then the dashboard also gave me a sense of, like, what the different threads are doing. Like, oh, this thread is the main orchestrator. There's a communication thread. There's one thread doing, like, visual UX. There's one doing, like, the managing the app server protocol and so forth, and even, like, creates it into, like, more conceptual work streams, like one work stream on

  129. 39:23

    auth, one work stream on testing, one work stream on UI polish. But throughout the night, sometimes Codex does get blocked. So then, uh, with the example of Convex and Vercel, it said... Oh, um, I, I marked it as blocked because the critical path, it, it didn't have the authentication. It didn't have, like, the identity to access these external services. But it still, like, did everything it could up till that point, so it said it was blocked. So it would come up in the UI, and then I just selected it, and I say, "Add to side chat." And it says, "Can you please explain what's going on?" 'Cause I just

  130. 39:53

    woke up. Like, I, I didn't have context of what's going on, but Codex was telling me it was blocked. So I'm trying to get Codex to, to catch me up to speed. So here it's, like, saying, "Explain step by step what I can do to unblock you here." So it gave me different options, how I can give the code, uh, Vercel, uh, details, the Convex details, and then I sort of did it. So here it... I sort of... I selected it and say, "Do that." So then the side thread told the main thread to sort of create the CLI login flow, and I would just click it over here, and I would be able to

  131. 40:23

    then, uh, authenticate and give Ver- uh, Codex, which again is running on a remote instance, access and the right permissions to manage my Vercel project. Then in the case of Convex, there's a deployment key. So I was like, "Hey, I'm not gonna paste that into the chat." So I asked Codex, "Give me a secure way to do it." Then it gave me this, like, bash code that I could just paste in, and I could then set the key into the remote environment. So that was, like, a way where I could col- collaborate with Codex in a remote instance, do different testing, managing permissions, and so forth. Um,

  132. 40:53

    so let's see what it actually built from this entire process. So let's look at a site builder, site one. So this is it, the internal form builder. So to recap again, there's an existing repo where... of a form builder, but I wanna, like, customize it for, like, maybe my internal team. We're, like, maybe five of us. I don't wanna spend money on SaaS, so we can build our own form builder. So it creates a new draft. You could call it, like, AI Engineers, uh, Workshop. Uh, hello. And then just say, like... Uh, yeah, I should, I should

  133. 41:23

    ask the Chrome extension to do this, but you know, for exposition purposes. Okay, and you can, like, preview the form, and then you can sort of share it thereafter. You can, like, copy, publish the link. Then there'll be a shareable link. You open the link, and then someone else in your team can fill it in. And again, because this is hosted on sites, only people that have access to your ChatGPT workspace can then access this form. So that's really nice. Yeah. Then you can see all the responses and so forth. So this was again what we did. So to summarize on the

  134. 41:53

    complex tasks, we had three different projects. One was a machine learning one where I wanted to re-implement a package with a Rust back end. One was to take a existing, uh, to build from scratch, like, a greenfield project, a live Q&A site. And the third was to take an existing open source project and customize it for my sort of needs. And Codex was able to sort of get it done in those cases. So underlying this are a few sort of, uh, concepts and primitives here. Let's

  135. 42:23

    go back to the slides.

  136. 42:28

    So first thing is compaction. You saw how Codex worked for seven hours. It worked for five hours. It worked for, like... Right now I think the other one's, like, 12 hours. And many times it's compacting its context window. This is not a new technology. In November '19, when we released GPT 5.1 Codex Max... Yes, the name was GPT 5.1 Codex Max. We natively trained the model on multiple, uh, context windows through a process called compaction, and that's what lets Codex be really good at it. Um,

  137. 42:58

    some examples of what people have said about compaction. To quote, "Did OpenAI basically solve compaction? I pretty much never had issue with 5.5 in Codex across ultra long threads spanning wen- many compactions." Uh, this individual from Japan said, like, uh, "In fact, Codex has strong compaction, so even if you make a lot, a lot in the same thread, it's less likely to forget. And in that sense, because it only remembers the good stuff, there's some token efficiency to that, I suppose." And as always, people think it's pretty cool, too.

  138. 43:29

    I'll let, uh, Charlie share a bit more about, like, specific, uh, primitives and concepts that enable this, such as Goal. Yeah.

  139. 43:36

    Thanks, Gabe. So as you saw, we, we have a goal that's already been running for fifteen hours, and, uh, I think, you know, Goal is, is one of these primitives that, um, is incredibly powerful if you can use it in the right way. Uh, one of the biggest things that I tell people who are just trying it out for the first time is be really specific with the criteria that you're giving the model to, to know when it's done. Um, and specifically, you want that criteria to be as verifiable as possible. Uh, sometimes I do feel a little bit like Goal is, is kind of like a genie in the sense that

  140. 44:06

    when it works well, it is insanely magical. Uh, and when it doesn't, sometimes I feel like I've entered into a monkey's paw type situation where, you know, it has done the thing technically that I asked it to do, but, like, very much not in the way that I was expecting. Um, so it-- when you give it a task, it'll run off. It'll keep evaluating itself against the success criteria, against the verifiable criteria that you've given it. Um, and at each turn, it'll check, "Hey, did this actually complete? Um, and if not, should I keep going?" Right? Um, there are quite a lot of use cases, uh, for very large

  141. 44:36

    projects in this way. Um, you know, Tama, as you can see, used a /goal that was running for forty hours, uh, that did a re-implementation of Doom in native Swift code, um, if that's the type of thing that, that you wanna give a shot. Um, but also things like code migrations, large refactors, uh, retrying loops, uh, experiments, and, you know, like games or, or full one-shotting apps with fairly detailed specs. Um, I think Goal is, is pretty incredible.

  142. 45:04

    Uh, we've also got subagents which we, uh, saw a little bit of a preview of here, right? Subagents in Codex, um, you know, by default will get used out of the box. You can tell, tell the model, "Hey, uh, you know, think about delegating to subagents here." Um, but you can also customize them. So if you want-- if you know you have a subagent that should be your, um, you know, code reviewer, right? You can have a specific prompt set of developer instructions here. Um, if you know you want that code reviewer to just be focused on, um, like, just be a

  143. 45:34

    lightweight model so that it runs quickly, or if you want it to be a much more intelligent model so that it's very thorough, you can customize that in its own, uh, subagents.toml file. Um, and you can give these names, and you can say, "Oh, go," you know, "check with the code reviewer. Go check with the docs researcher, um, in order to, to implement," right? Subagents, uh, can work in tandem with, uh, thread-to-thread handoff in Codex. So, um, thread-to-thread is, is quite new. I think we just released it in the last week or two.

  144. 46:00

    Two or three weeks.

  145. 46:01

    Yeah. Um, and so, uh, that's just an even higher level of abstraction. Um, and the way that I think about balancing them is, um, when the s-- the delegation of work needs to be visible to you, the user, versus when it just needs to be visible to the model. So subagents we started with because I think that is when the work, you know, primarily needs to be visible to the model. Um, if you wanna delegate, you know, some, uh, like a subcontractor, you know, I don't really care what they're talking about. I just wanna make sure that the work comes back good.

  146. 46:31

    Subagents are a great fit, and the model is great at figuring out when it should be, um, handing off into subagents. Uh, it also keeps the, the context separate. I think unlike, you know, a lot of other tools that, um, we've, we've brought to Codex to help keep context windows efficient, um, subagents are one where, uh, you know, you're keeping the context separate, and the mo-- the, the models are communicating with each other. Thread-to-thread handoff is useful for when you still wanna be in the loop, and you still wanna be able to, like, look at everything that's happening. Um, and perhaps even larger than that, you are mentally thinking about

  147. 47:00

    separating these two, uh, these two threads, right? So, um, one example is, uh, over the weekend, I was working on a game, um, and I had a creative director thread, right? I could have used a subagent, and that would have worked fine, but I wanted to see all of the different things that were happening, and I, I didn't want it to just be delegated into the ether via subagents. So with the creative director thread, I opened up, um, different threads, one to work on the art direction, one to work on the music, one to polish the animations, one to figure out the game mechanics. And I was using those,

  148. 47:30

    um, and I told them in the agents.md file, "When you're done, go and check with the creative director thread," which I'd given a bunch of extra context on, you know, the look and feel I wanted for the game. "Go check with the creative director and get sign-off before you consider this task done," right? Uh, and they all did so, and, you know, it was actually really interesting to watch the threads just constantly sending messages to each other back and forth as they worked on the game.

  149. 47:57

    Uh, we didn't quite show it here, but we also have hooks in the Codex app. Um, if you wanna set up your own hooks, you can do so in your config.toml file, uh, in your Codex home directory. Um, these are really great for when you wanna start introducing, uh, deterministic behavior at specific checkpoints in your software development life cycle, um, or in your just, you know, general, uh, development life cycle. So I think in these cases, you might use hooks to do things like, um, you know, send your conversation to, uh, a logging engine, right? Um, you might want to

  150. 48:27

    scan inputs or outputs. I think there's quite a lot of security use cases for hooks, right? Before Codex takes a specific action, maybe you wanna look at the prompts and figure out, "Hey, have we accidentally pasted an API key that's, that's gonna go somewhere?" Um, or before it takes, uh, an action to call a tool, "Hey, is this tool, you know, on, like, not on our approved whitelist or something that we need to be extra sensitive about when calling the tool to evaluate if it's, if it's safe to use?" Um, and we can also do custom validation checks as well when the turn stops. These days, you know, you can tell the

  151. 48:56

    agents.md, "Hey, go and use the, the linter," right, "before you consider the work done to check your work." Um, and it's really good about doing that. But if there are things that are a little bit more sophisticated or a little bit more complex that you wanna do to validate the output, uh, you can use hooks to, to take care of that as well. Um, you know, I find that when building more and more complex projects, some of the most important things are figuring out, like I've said, the verification criteria or the success criteria. You wanna give the model as many boundaries as it can, as you can, and let it fill in the lines, um, in

  152. 49:26

    accordance with, like, the boundaries that you've given it.

  153. 49:31

    Cool. Um, and that brings us to working over time.

  154. 49:36

    Thanks, Charlie. So

  155. 49:40

    just to quickly share, on hooks, uh, this looked to my mind as well. So we have the OpenAI agents SDK, and that's actually a quick example of, uh, hooks right over there. So if you go to settings, all the hooks are s- sort of there, and you would see the hook for the agent's SDK repo. And here it's sort of doing at the every end of its turn, a Python script to just tidy the repo. So again, there are like some of these housekeeping that maybe you don't want a language model to do, you can run it deterministically with a hook. And hooks can become particularly useful with long-running tasks because

  156. 50:10

    there is that trade-off. You're giving the agent more autonomy, it's gonna work off... work longer and longer without your supervision or abstracted supervision. And to that, to that end, hooks become useful in getting the guardrail. So on working over time, what we're trying to show here is that now it's our collaborator Codex, we have some of the power tool, power user tools like computer use. How can we, like, stitch this all together and let, you know, Codex work on its own? I mean, for exposition purpose, here we are, like, prompting and, like, being more bit

  157. 50:40

    methodical, but in reality, you can just, like, let Codex work on it, like, with thread handoffs. In practice, I could have all my demos in one thread and just ask that thread to segregate, to separate it, then send it and delegate it to the various threads. So here's one quick example of a, a automation. So right now you notice, like, on my, uh, screen over here that these, like, blue orbs. So this is just showing that it's a remote host, and maybe you want to create one. So we have a nice plugin here with the DigitalOcean plugin, which we just,

  158. 51:10

    uh, released last week. And then you sort of do it. You can just try it in chat. I'm just gonna, like, fire it off as well. And what it's gonna do is gonna, like, provision this infrastructure. It's gonna take some time, and I'm not gonna, like, stand here and look for, for it to provision this infrastructure. Uh, and Codex is not gonna, like, just keeps, um, like, waiting. What it's gonna do is create a heartbeat automation. So it's gonna, like, monitor and so forth. So it says yes, and let's just show it in the spirit of this being a cooking show, an example that was already done.

  159. 51:40

    So here, provision a DigitalOcean droplet. Okay, yes. And then it created the automation here. I set a Codex heartbeat to resume this thread in about five minutes. So it's gonna keep checking every five minutes. Is it ready? Is it ready? Is it ready? Is it ready? And once it's ready, it will take the next step. So in this case, it did it. So, okay, the droplet's ready, and it even gave a nice link. So if you click this link, it's a deep link. It's a link, like a deep link to the settings page within the Codex app, which automatically already sets up the

  160. 52:10

    SSH instance for you. So you just need to press it, and then it'll pre-configure everything. It knows where... It'll create the SSH key for you if necessary, and so forth. So that's nice. And you can sort of ex- extend this to other parts of, like, say, the software development life cycle. Uh, maybe you're building something, you need to deploy, you have a pipeline that triggers event-based, and it takes time. So you get a Codex to, like, run every five minutes, every ten minutes, check is it ready, is it ready, is it ready? And maybe depending on the event, then it sort of spawns different tasks, and so forth. So that's a heartbeat automation. It runs within the

  161. 52:40

    thread. But we can also do automations that sort of spawn new threads. Uh, so one example of this, so right now I have actually this Slack channel. So let's see if it's up. So remember the macOS translation app. So in the last, say,

  162. 53:01

    forty-five minutes, we had a bunch of users all called Gabriel Chua, and they've been using it, and they, they've been giving a lot of feedback on the translation app. For example, "Oh, the screen feels focused. I can tell where to start," so it's positive. "Is there a way to do a quick sound check before the audience sees caption?" So forth. So there's a lot of feedback, people are giving feedback, and I wanna, like, you know, give this feedback the attention it deserves. So I'm gonna create an automation to do that. So just make sure we are in a... I'm just gonna close this since we don't need it already.

  163. 53:32

    Over here, we create a new thread. And then I'm just gonna, like, do a app shot.

  164. 53:40

    Um, every half an hour, can you just, like, look at the comments and, uh, suggestions in this channel and reply to them? Uh, I think there'll be, like, a few categories. There's one category where it's, like, compliments. There's one category where it's a bug report. There's one category where it's a feature request, uh, so forth. Let me just, like, fire that off. So it's gonna create an automation. And I s- I just sent it, but I realized, oh, wait, um, if it's a feature request or a bug fix, I wanna do something different. So I'm gonna, like, steer

  165. 54:09

    this. Oh, um, if it's a, uh, bug fix or a feature request, can you also then, like, spin a new thread in a work tree and actually address that and then open a PR and then, like, uh, request Dom to review it, and then so forth, and ask Dom to, like, approve it within, like, twenty-four hours? If not, just keep reminding him. So yeah, it's, it's, it's a comical example, but here it's where I'm tr- I, I... what I was first doing was steering. So I'm gonna... I don't have to wait for Codex to finish. I can sort of steer it midway. And what

  166. 54:39

    I'm trying to do is, in addition to all this feedback and classifying it, if it's a, it's a bug report, if it's something that, uh, we can fix, Codex would then sort of spawn a thread. So we're putting all these primitives together, the automation, the plugin, which can, can read Slack, the fact that one Codex thread can communicate with other Codex threads and even create work tree threads and so forth. So imagine then you wake up, and this is a bit like that software factory, where all the PRs are ready, then you can even add additional com, uh, layers to it where it's au- this review. You, th- you build a test version.

  167. 55:09

    You send it to someone. They test it, and if they give the thumbs up, then it's good to go. That feedback gets fed back to the PR, then the PR reviewer sees all this holistically. So yeah, i- there's a lot of it you can chain together to build that automation. So then every half an hour it would do this, and it will find Dom as well in the, uh, channel. So you, we were talking earlier about how we've actually been running a automation in the background, and let's find that. So let's see Here it is. So

  168. 55:40

    earlier, while everyone was coming into the room, I was asking Codex, "Can you run an automation every five minutes, taking a screenshot of what's happening on screen?" The Google dr- uh, Chrome, uh, Google Drive presentation. We only have sixty minutes. And I'm asking Codex to estimate to what extent are we gonna overrun. So right now, it's at checkpoint thirteen. It's seeing if the demo is there. It's flashing twelve, uh, ten twelve, and so forth. So Codex is doing its own assessment of, like, to what extent are we gonna finish on time. So it's a bit of a meta example of how you can use automations and chain it together with other tools,

  169. 56:10

    like in this case, the Chrome extension.

  170. 56:12

    It looks like it's telling us we might be a little bit behind the ball here.

  171. 56:14

    Yeah. It's, it's a... You know, it is what it is. So automations have been really useful in my personal life to an e-extent, really, like, getting a lot of things done. Uh, some automations that may be useful too. Uh, one automation you can run, maybe every Friday you ask Codex to look at all your conversation treads and see are there skills that you could improve, are there, like, things you could add to the agents or markdown, are there skills you could just remove, maybe skills that you haven't been using. A second automation or approach that may be useful, maybe you're using automations to draft r-

  172. 56:44

    email replies. You can have one automation that drafts the email reply, but you can run a second automation that sort of cross-validates against the eventual reply you sent 'cause that is the ground truth. Then it updates, like, a markdown file with the best practices, tips, and so forth. And in that regard, the email drafting automation can get better over time. So, um, I'll let, uh, Charlie just recap what automations are and share a bit more about the app server as well, our final step.

  173. 57:15

    Cool. Uh, so I think, yeah, automations, uh, there's two kypes, two types of automations, right? Um, I think the more powerful one these days tends to be the in-app or in-thread heartbeat automations. And so those just live in a single thread. They run on a timer. Um, they, you know, keep going over and over until they're, they're done. Um, these, I find, like, have just almost unlimited use cases. Uh, me person- Gabe shared some of his. I think me personally, um, I feel like I'm drowning in Slack and email and linear notifications every day. Uh, so I have an automation that just sort of, uh, goes through and checks the

  174. 57:45

    latest updates from, you know, all the sources that might need my attention. I have a local Obsidian, uh, vault, and so it's just constantly updating my daily notes. It's constantly updating my, uh, notes on different work streams, and it's keeping a lot of context there. And that way, I can just pull from that later if I say, "Hey, you know, what's going on with the latest, uh, AI Engineer Worlds Fair talk? Like, you know, how far are we on the demos? How far are we on the notes?" Um, and then I can just... Codex can just quickly tell me, "Here's what's happened since the last time you checked in." Uh, but automations can also live in a new thread. If you set them up in the Codex

  175. 58:14

    app, uh, under Schedule, the Scheduled page, um, you can have them spin out in their own thread as well. Um, and then that's useful if you just wanna say, "Look, every week, I want you to do a recap of what all of my direct reports have been doing," or, "I want you to look at all of the analytics from this, uh, system, and then we're just gonna do a one-time look together, and then I'm gonna archive the thread and not worry about it again."

  176. 58:37

    Uh, I think last but not least, we've got embedding Codex anywhere. Um, and so I think we're... For this one, uh, I wanna talk a little bit about the app server protocol, right, which is, um, the core protocol that underpins the Codex app, the Codex CLI, uh, the Codex VS code extension. Um, the app server is basically, uh, what allows all of these services to talk with the main Codex harness, right? Um, and it's integrating all of the tool calls and the plug-ins and the context compaction that you know and love.

  177. 59:07

    Um, we have, you know, built a lot of the components of the Codex ecosystem to be open source because, uh, we really want to encourage developers to build wherever you are and for whatever use case that, that you have. And so to that extent, um, you know, if you're not aware, you can embed Codex into your product using app server. Uh, I think this brings a lot of the, uh, power of Codex directly to you. Um, and it brings the ability for your users to, uh, use their ChatGPT subscription and, perhaps more importantly, use their ChatGPT token

  178. 59:37

    budget, uh, to power the products that you're building. Um, and like I said, it has all the same underlying features that we've been demoing, you know, compaction, steering. The ones that need to live on your computer, like computer use or, or, you know, not gonna be there. But for everything that exists in the core, uh, you know, agent harness, we've got that in the app server.

  179. 59:58

    So as a quick demo of that. So there was a fourth long-running task that I kicked off last night. It was actually to build my own user interface for Codex, the app server, called Retrodex. So it's been working for a while. I think there are different tasks. I worked with it. And this is sort of the work in progress, Retrodex. It's using the app server. Uh, let's just... Okay, can zoom in, then optimize it for the stage. But yeah, you can, like, choose the model, five point five, five point four mini. The reasoning effort, so I told, like, oh, make it a bit fu-funky when you change the reasoning effort. So at

  180. 1:00:28

    low, it's like a small glitter. At medium, the color changes. You think at, like, high and extra high. It's like, ooh, a bit more. So there you could just, like, simply say, like, uh, "Hello," and let's see. This is in Retrodex. And I'm just gonna send hello. And then allow access. It's gonna run, and it's gonna reply back, "Hello." And what's nice, and to just prove that it's all running on the same app server, is that if we go to the Codex app in the same project, uh, Retrodex, you should see the two right here. Let's

  181. 1:00:58

    see.

  182. 1:01:04

    Mm. Yeah, it, it should. We could change their treads, I believe. But yeah, that's an example of how you can build... It's, it's a trivial example of building a, a funky user interface. But let's say you wanna build agents on the cloud. You wanna manage different agent providers. Uh, you're considering in large org-organization your different sort of providers you're using, and you wanna build a unified layer. The app server is a useful way for you to integrate it into your own platforms. And because it's open source, if you go to... You Git clone OpenAI/Codex, you can just ask Codex about Codex. You could just say,

  183. 1:01:34

    like, uh, "Hey, can you tell me more about the Codex app server?" And then you can fire it off, and you can start building your own applications or integrations from there immediately. So that's the app server. And we have a nice blog post talking about it, and there've been more engineering blog posts about, like, our Windows sandbox, first of its kind, how we built the agent loop, amazing stuff. And it's a good reference point also. Like, if you're an AI engineer, seeing the, uh, Codex repo, it's a great source, and we'll have, uh, further talks in the subsequent

  184. 1:02:04

    days about this. So I'll let Charlie share some closing thoughts after seeing this entire cooking show.

  185. 1:02:13

    Uh, thanks, Gabe. And thanks, everybody, again for being here. I think, um, hopefully you've learned a bit today about how you can use Codex to kinda level up your workflows. Some of the co- closing thoughts that I wanna end on here are, you know, questions that I repeatedly ask myself as I'm trying to get to the frontier of what Codex can do, right? Um, they start pretty simple just by asking, have I tried asking Codex to do it, right? Um, I think, like, one of the things around the office that we say, uh, more frequently than ever now is, "Have you asked Codex?" Um, but I think as you evolve that workflow,

  186. 1:02:43

    right, you wanna go from just asking in individual turns to building loops, right? We've all heard about loop maxing. That is the new thing. Um, and I think we wanna think about how does that, how does that work, right? Like, what is it that the loop needs to be checking at each turn, right? Um, as you go down this path, there's gonna be additional questions like, okay, you know, as it's progressing, how, how... Like, what blocked Codex from actually getting to the answer, right? Um, was there context that it needed? Were there permissions that it needed? Are there ways to build that context into repeatable

  187. 1:03:13

    tools like skills or plugins that Codex can use the next time it runs into this problem? Um, how do I start scaling, kind of from one agent to many, right? Like, as you think about the work, is there ways that you can be prompting Codex to implicitly delegate the structure of the things that you are building, um, or to start orchestrating across multiple threads to have different Codex instances communicating with each other. And then ultimately, uh, you know, if you take a step back and ask the question, okay, why does this process need a human at all, right? I think, like, um, as we've shown, you know, one

  188. 1:03:43

    thing that you might assume, uh, when building with Codex or when building with the OpenAI API, is that a human needs to be involved to, like, go to the platform page and go get an API key and then paste it into the environment, right? I think, like, uh, as we've shown with the developers plugin, that is not the case anymore. The plugin itself can just go fetch the API key and set it up locally for you. Um, and ultimately, you know, we do still want humans in the loop, like, where they need to be, but, uh, for everything that's not that, like, how do we figure out how to take the human out of the process to unblock ourselves?

  189. 1:04:12

    And how do we figure out how to build the machine that builds the machine, right? I think, like, it is one thing to sit with a coding agent and build a single piece of software. It is another thing to start zooming out and say, "Actually, I want to build a software factory," right? I want to build a system that, um, that then builds individual pieces of software on an assembly line as we go. Um, and so a lot of these things are kinda how I start to get from, um, just an individual ask to something significantly bigger.

  190. 1:04:43

    Cool. And I think with that, um, we are-

  191. 1:04:46

    Do a re- a recap of-

  192. 1:04:46

    Yeah, yeah

  193. 1:04:47

    ... things. Okay.

  194. 1:04:49

    Go ahead.

  195. 1:04:49

    So, um, thanks so much for joining us today on a 9:00 AM on a Monday. It's the... I hope it's a great way to kick off your, uh, AI Engineer World's Fair. And that's not it. So tomorrow, we have our opening keynote. It's gonna be fun. Uh, do check it out. We also have Jason on the developer experience team giving a talk about how he gets the most out of Codex. It's a nice, uh, foil to today's presentation about how he's using it on a day-to-day basis. And then he has a workshop after that where he goes through step by

  196. 1:05:19

    step, like setting up the plugins, doing app shots and so forth. So today's session, it's more, uh, like a whirlwind tour. This workshop will go through the step by steps of how to get s- to set yourself up for success with Codex. We then have a session on h- uh, from the harness about the harness, and then sessions about voice agents. Charlie will be talking about that as well. Uh, Dom will be doing a session on building on the Codex harness, and then another session explaining behind the Codex

  197. 1:05:49

    harness. So the first one's you build on top of it, then the second one is about they explain the harness. Then lastly, a session about LLM inference and production, and so how we sort of do that at OpenAI. So thank you for joining us, and that's the end of the first segment of today's workshop, where we sort of give that whirlwind tour of everything Codex, and hopefully you have something fun to code with. So it's time to build, and to help you with that for today's session, uh, first things first, if you've not download the c- downloaded the Codex app,

  198. 1:06:19

    do give it a download. And for credits, so we are giving 100 dot USD in Codex credits and 100 USD in API credits. So we'll lift- leave this QR code up there. Everyone's, uh, phones up and s- take a s- photo. Yeah. I could, like, app shot this and say, "All right. Send this via email to everyone who attended." But we don't have your emails. Yeah.

  199. 1:06:39

    Cool. And I would also mention, uh, please stick around, uh, to the end of the build session, um, if you can. We brought, uh, swag for everybody. So at the end of the two-hour mark, um, on your way out, you can pick up a sweatshirt, I believe.

  200. 1:06:52

    We'll love to see what you're building as well. Yeah.

  201. 1:06:55

    Yeah.

  202. 1:06:55

    Cooking. Cooking.

  203. 1:06:56

    We will be here. We will be here in the audience.

  204. 1:06:57

    Oh, yeah, yeah.

  205. 1:06:57

    We've got, uh, OpenAI staff here as well, um, to help unblock you or to answer any questions about, uh, using Codex.

  206. 1:07:03

    Thank you.