AI Engineer World's Fair 2026
Building on the Codex Harness — Dominik Kundel, OpenAI
Read the talk
Building on the Codex Harness
Dominik Kundel explains how the Codex App Server lets a custom application reuse Codex’s agent machinery, add its own instructions and tools, and control threads—then demonstrates the same integration inside Doom.
From a talk by Dominik Kundel
At a glance
Ideas worth remembering
App Server exposes the Codex harness through client messages and events, including threads, streaming output, goals, plugins and filesystem capabilities.
Developer instructions add application behavior; dynamic tools add actions; deferred tools let the agent discover functions without loading every definition up front.
Bundle a known harness version and generate bindings for it so the client’s expected capabilities do not depend on the user’s installed CLI.
The Doom demo shows that a custom interface can expose new domain actions while retaining the harness’s local coding and workspace capabilities.
Choose codex exec for non-interactive scripted tasks and App Server when the application needs ongoing interaction, multiple threads or deeper harness control.
Build the interface without rebuilding the agent
Building an agent harness can feel daunting before the interface work even begins. Dominik Kundel, who works on Codex developer experience at OpenAI, starts with a way to reduce that burden: build on the harness that already powers Codex. The Codex App Server exposes that harness through a protocol, allowing another application to use its capabilities and integrate deeply with its operation.
Both the harness and App Server are open source under the Apache 2 license. They can be inspected, forked and adapted. Model choice is also configurable: Kundel describes support for other providers with a Responses API-compatible interface, including convenience settings for LM Studio and Ollama. That compatibility requirement matters; choosing another provider still requires an API the harness can communicate with.
The protocol follows a JSON-RPC style and already connects the harness to the Codex app and IDE extensions. Third-party integrations in Xcode and JetBrains use the same machinery. A Claude Code plugin also uses App Server to hand tasks to Codex or ask it to review code. These interfaces differ, but their underlying agent integration follows the same path.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A message becomes threads, configuration reads and streaming events
The application you build is the client; Codex App Server is the server. The client sends messages to request actions and receives events as work proceeds. Kundel makes this exchange concrete with an inspector that intercepts traffic between the real Codex app and App Server. Sending a message in the app produces a stream of protocol activity.
After restarting the example to clear older buffered events, the inspector shows thread starts carrying configuration and project information, configuration reads, and delta messages carrying streamed data. A visible response therefore sits on top of several kinds of exchange: establishing the work’s context, starting the work and delivering its output incrementally. File search and other key app functions also pass through this connection, making them available to a custom client.
At the time of the talk, App Server exposes more than 120 client messages. They cover several distinct parts of an integration:
- Client setup: initialization configures details specific to the application.
- Agent work: thread modifications and turns control conversations and tasks; goals support a client’s own goal-mode interface.
- Supporting capabilities: plugin listing, file search and filesystem changes expose functions beyond response text.
Agent Client Protocol provides a similar client-to-agent concept. Kundel’s distinction is specificity: ACP is more generic, while App Server exposes control over individual parts of the Codex harness. That is useful when an application needs those controls rather than only a common interface to an agent. App Server also supports sign-in with ChatGPT, allowing users to bring their subscription and consume their own Codex usage limit through the custom application.
A small, copyable example makes the request envelope easier to inspect than the projected debugger. This initialization message uses the fields in the official App Server integration guide:
json
{
"id": 1,
"method": "initialize",
"params": {
"clientInfo": {
"name": "example_client",
"title": "Example Client",
"version": "1.0.0"
}
}
}
For the documented standard-input/output transport, send each message on one line. Wait for the initialization response before sending the initialized notification; starting a conversation is a later step. The example shows protocol structure, not a complete authenticated client.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Add application behavior through instructions and tools
Developer instructions let an application add behavior on top of Codex’s existing system prompt. The base prompt remains available in the repository, but Kundel recommends starting with additions rather than replacing it. This preserves the work OpenAI puts into model-specific prompt tuning while giving the application room to explain its own functions and expectations. The Codex app itself uses this approach for capabilities that differ from the IDE extension or CLI.
Tools supply the actions those instructions can refer to. An integration can use MCP or define dynamic tools directly. Dynamic tools expose additional functions the agent can call, so the application can contribute capabilities that the underlying coding harness would not otherwise know about.
Deferred loading changes how those functions become available. Instead of injecting all their definitions into the system prompt up front, the integration marks tools as deferred and lets Codex find them through tool search. The functions remain discoverable without filling the initial context with every possible action. This introduces a discovery step in exchange for reducing the amount of tool information carried before it is needed.
Configuration can also vary by thread. A client can select a model provider or model and adjust sandbox permissions, including which network connections are allowed and which files the agent can modify. Instructions describe the application’s intended behavior, tools expose its actions, and these settings control the environment in which a particular thread works.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Ship a binary whose protocol your application understands
App Server comes in the same binary as the Codex CLI. An existing CLI installation can therefore start it, but the Codex app bundles its own binary and updates it with the application. On Windows, the app ships both Windows and Linux builds: WSL mode uses the Linux binary, while native mode starts the Windows one.
Bundling solves a practical version problem. The protocol changes between releases, so the application needs to know which capabilities and message shapes it is building against. TypeScript bindings or JSON Schema can be generated for a particular App Server version. Shipping that version alongside the client keeps the implementation and its protocol definitions aligned, rather than making behavior depend on whichever CLI the user happens to have installed.
App Server supports both standard input/output and WebSockets as transports. The talk leaves that choice to the application’s use case. For integration, the Python SDK handles much of the lifecycle management and offers a familiar interface, including help with sign-in with ChatGPT. It reduces the amount of process-management work the client must implement itself.
The other route is to ask Codex to build the client, pointing it at the protocol documentation when needed. Kundel describes using it to create a Windows-native Visual Basic client largely in one shot, and suggests combining it with plugins for tasks such as building macOS apps. That is a personal implementation example, not a guarantee about how much work another application will require.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Doom supplies the interface—and tools that change the game
A custom client does not have to resemble a chat application. It can trigger the agent inside an existing interface, run it in the background or put it on unusual hardware. Kundel takes the idea into a video game: run Codex inside Doom. The demonstration begins with a small reminder that embedding an agent does not remove ordinary interface problems—a dual-monitor setup briefly interferes with control capture.
The application uses Electron and Cloudflare’s Doom WASM renderer. Most of the game engine is implemented in C, and it loads Freedoom with a game file that Codex helped modify. The Codex interface is rendered within the game engine. Input passes through Electron’s IPC connection to App Server, and a greeting produces an agent response inside the game.
The useful change comes when dynamic tools connect the agent to game actions. A request for armor and health causes tool calls that grant them; a request for all weapons and ammunition calls the corresponding tools. The causal chain runs from the player’s request, through the agent’s selection of exposed functions, to a change in the game. Rendering a reply makes Codex visible in Doom. Exposing actions lets Codex affect Doom.
Player input reaches the same harness through Doom’s interface → Electron IPC → local Codex App Server. From there, the available tools lead to two different destinations:
| Request | Tool destination | Observable result |
|---|---|---|
| Grant armor, health or ammunition | Dynamic game tools | Game state changes |
| Show the current directory | Shell in the agent workspace | A filesystem path returns to the game interface |
The shell is an alternative action available to the harness, not a step after granting game items.
Kundel demonstrates that second path by asking for the current directory. The local App Server calls a shell, obtains the working directory and renders it in the game. The agent still has its own workspace, so coding from inside Doom is possible. His recommendation is more restrained: the Codex app is probably the better place to do that work. The game demonstrates the range of interfaces and actions available, rather than a claim that every creative interface improves productivity.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Use codex exec for scripts, App Server for managed interactions
codex exec provides a non-interactive way to start Codex tasks. It suits scripting, including CI/CD work such as automatically fixing a pull request. When a script needs to hand off a task, that interface can be enough.
App Server becomes more useful when the application presents and controls the agent over time, needs goal functionality or starts multiple threads concurrently. An evaluation harness is one example. Instead of building its own mechanism to parallelize many codex exec calls, the client can send messages to create threads and start tasks through the protocol. The distinction is the control the surrounding application needs over the work.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Keep the version predictable and the customization additive
The closing advice turns the integration into three concrete engineering decisions:
- Ship your own harness version. A user’s installed CLI may lack a feature the client expects, or differ in ways that break experimental functionality. Bundling puts that version choice under the application’s control.
- Expose application functions as deferred tools. The Codex app uses tools for actions such as starting automations. Custom clients can expose their own functions while keeping all those definitions out of the initial context.
- Start with developer instructions. Add the application’s requirements before deciding to replace the base prompt.
That final recommendation carries a model-performance concern. Harness prompts are tuned with the model’s expected behavior in mind; Kundel describes keeping them “in distribution.” Replacing the system prompt discards that work and makes the client responsible for whether its replacement preserves good agent behavior. Developer instructions offer a smaller first change: explain the application’s needs while retaining the base that was built for the harness.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
Further reading
Official protocol documentation and integration examples for the harness used in the talk.
Read the complete timestamped transcript
- 0:12
All right everyo- hi everyone. Um, I want to start with a quick raise of hands. How many of you have built your own agent harness? All right. How many of you, um, feel like it's way too intimidating to build an agent harness? Okay, cool. So, uh, over the next 18 minutes or so, I wanna talk to you all about how you can actually build on top of the Codex Harness, and hopefully make it a bit easier and less daunting for you to, um, actually build your own interfaces on an a- on an Agent Harness,
- 0:42
and hopefully also a little bit fun. Uh, my name is Dom. I work, uh, at OpenAI on developer experience, specifically for Codex. Um, and I have one more raise of hands. I promise I'm gonna give you all a break. Um, but how many of you have heard of the Codex App Server? All right, this was a bit of a trick question because if you attended the keynote yesterday, all of you should have raised your, your hand, um, otherwise you probably didn't listen to Roman. Um, but, uh, we talked briefly about it yesterday, but we didn't really explain what the app server is.
- 1:13
Um, so the app server's actually the protocol that wraps the Codex Harness. Um, if you've ever looked into something like Agent Client Protocol, this is a very similar concept where basically we have a protocol that allows you to directly build on top of the, uh, harness and deeply integrate into it. And we'll talk a bit in the talk about what the differences are between, uh, App Server and ACP. Uh, the big thing here to keep in mind is both the Codex Harness as well as the App Server itself are open source, so
- 1:43
we're gonna cover a couple of different things on how you can engage with it. Um, but more importantly, we're gonna both dive tomorrow in the afternoon into how a couple of these things work behind the scenes. But you can also always just ask Codex to read the repo and dive really deep into it. It's Apache 2 licensed, so feel free to fork it, do it, make it your own and have- have fun with it. You can even connect it to any other model provider as long as they have a responses API capab- uh, compatible API. Uh, we even have, uh, convenience
- 2:12
settings to connect it easily to LM Studio or Ollama. Um, in terms of how it works, the App Server protocol is a JSON RPC style protocol that we actually use to power the Codex app, the IDE extension, uh, our VS code extension, so all of our first party interfaces, but also third-party ones. So if you've ever used Codex through Xcode, JetBrains, IDEs, um, or even some open source projects like Theos T3 Code or Remotex, you can actually... Uh,
- 2:42
all of these are powered by the same harness, the same app server. We even used it earlier this year to put Codex into Claude Code, um, using that same app server. So if you're still on u- uh Claude Code, uh, and you wanna use Codex from there, either to have it review your code or, um, you know, pass off tasks to Codex to handle it, you can actually use that plugin, and it uses that same app server, that same harness behind the scenes. Uh, if you've looked into things like ACP,
- 3:13
this should seem pretty familiar. At a high level, basically the app server, the way it works is we're gonna send messages between the client, which is the application you're building, and the server, which is the Codex App Server, uh, to perform different actions and receive events. But rather than showing you a bunch of hypothetical messages, um, I went a bit overboard and actually built a little inspector that intercepts all of the events between the real Codex app and the, uh, Codex App Server, so we can look at
- 3:43
some real events. Um, let me actually switch over here to the Codex app. So I loaded up that same inspector in the, um, in-app browser here of the Codex app. And so if we send a message, we're gonna see here a bunch of different events coming in. Um, this might seem daunting at first, but as you're sort of going through some of these messages, a lot of this should feel relatively familiar. So, um, we can look through here, um,
- 4:13
some of these might be a... Let me do this once more. Um, I think I had some old ones buffered, so just gonna start a new one. Um,
- 4:28
there we go. Also, don't mind my Codex pad up here. Uh, but you're gonna see like different messages from like threads being started, uh, with sort of your different configurations of what is actually happening, what project are we working on, uh, over to configuration reads, um, where is new threads being started, and then here's like some of your delta messages of like d- uh, data actually being streamed in.
- 4:58
And so there's a lot of additional information that you can find here. And one of the big things that makes this protocol so flexible is that we're really exposing anything that is really in the app, um, a key functionality. Meaning if we're streaming in things, we're doing file search, all of that information is actually handed back and forth between the app server and the, and the app. And so if you're building your own harness, uh, harness integration, your own client, you can actually leverage, uh, the same capabilities.
- 5:29
In fact, we have right now over 120, um, so different client, uh, client messages that we continuously add more to. Um, and this goes from everything from initialization, where you're able to configure some of the more like specific things to your client, um, over thread modifications and, uh, turns, which are sort of the things that you would assume are like the most classic things, but also stu- uh, stuff like setting a goal if you wanna implement your own version of goal
- 5:59
mode- Um, and listing plugins, doing all of the different other aspects that you might see in the app, all the way to things like file search and modific- uh, modifying the file system. So you can really have that full access. We also try to make it easy so that you can, uh, configure the underlying agent, and that's a bit different to ACP, where it's a more s- more, uh, generic protocol. We can actually give you more control over the individual parts of the harness as well through the protocol, and we'll look into that in a second. The other thing that is ex-
- 6:29
exciting for a lot of people is you can actually, you- if you're using this protocol, you can have sign-in with ChatGPT as part of that, meaning that if someone is using your app, they can still bring their own ChatGPT subscription, and it goes off their Codex, uh, use- usage limit. So if you've used something like T3 Code or similar, you're actually using that. In terms of customizing the agent, the mo- the biggest two things then you might wanna look at if you're building on top of this app server is developer
- 6:59
instructions as a way for you to augment the instruction, the system instructions of the actual app, as well as, uh, bringing tools. So with developer instructions, we're actually appending on top of the system prompt of Codex. You can look at the actual system prompt of Codex as well. It's in the repo. Um, but we generally recommend that you actually append to that using the developer instructions. That way, you're still leveraging all of the work that we put into tweaking the prompts for the individual models, while still
- 7:28
customizing it for your specific use case of the harness. This is, for example, what we do when, uh, you load up the Codex app, since there's some functionality that is, uh, more specific for the app over, like, the IDE extension or CLI, for example. And then the other part of this is dynamic tools, where you can expose, uh, additional functions that you wanna, uh, that wa- you want the agent to be able to use. You could use MCP or other things similar to what you would do in, like, the agent client protocol, but you can also just
- 7:58
straight up define tools, and you can actually mark them as deferred loading, which means that rather than injecting them into the system prompt, it, uh, allows Codex to use a tool search tool, which we'll cover more tomorrow in the other session. Um, but it allows it to y- find these tools through tool search rather than polluting the, uh, overall context window. And then on top of that, you can even, on a per-thread basis, modify any other configuration that you might have for your agent. So think about,
- 8:28
uh, setting a different model provider, choosing a different model, or even doing fine grain control on the sandbox, like what network connections, uh, it should have access to or what files it can modify. In terms of how you can actually use the app server, the first thing to, uh, know is that the app server actually comes as part of the Codex CLI. It's the same binary, so if you have the Codex CLI already on your system, you can just, uh, use that to spin up the app server right now.
- 8:59
But you can also look at how we're, for example, shipping it with the Codex app. So the Codex app has the, uh, binary as part of the actual bundle. So if you're looking into your installed Codex app, you should find it in the contents of that bundle, and it changes with every version of the app that you're updating. We're shipping a new version of the CLI behind the scenes. On Windows, we actually ship with two versions of the binary, both the Linux compiled one and the Windows one, so that if you're loading the app in WSL mode, it
- 9:29
uses, uh, the Linux one, and otherwise it starts up the app server in Windows native.
- 9:35
Um, from here, you actually have a lot of different, um, additional convenience functionality. One, you can have Codex generate you, for that particular version of the app server, the TypeScript bindings or the, um, JSON schema, which is incredibly helpful since this, uh, protocol regularly changes between versions, which is also why you actually wanna bundle it with your application, which is what we're doing with the Codex app, so that you have full control over what are the actual capabilities of the version that you're
- 10:05
expecting to build against. Um, in terms of protocols, you can spin up the app server both using standard input output, so stdio, or you can use WebSockets. It's gonna depend on your use case, and honestly, the easiest way to figure out what works better for you is probably to just ask Codex. Um, from here, there's two ways that you can integrate it, um, most easily. The first one is the Python SDK that we recently re- uh, released. It handles a lot of the life
- 10:35
cycle management for you, so you don't have to deal with any of that. And then on top of that, it makes it a bit easier for you to, like, integrate things like sign-in with ChatGPT, and gives you, like, a familiar interface. That being said, we are at an AI engineer conference, so, um, probably most of you are gonna choose option two, which is you just ask Codex. So unsurprisingly, Codex is incredibly good at implementing this protocol, and you largely just have to point it at, um, the documentation. In most cases, you don't even have to do that, but it
- 11:05
sometimes helps it to nudge it. Um, and from here, you can combine it with any of our plugins, like build macOS apps, um, if you wanna quickly spin up something or completely go wild. I used it, for example, to build a client using, um, Visual Basic, uh, to have a, like, truly Windows native app. Um, and it did it largely in one shot. So, uh, you can build really impressive applications with this in a fairly short amount of time without you having to learn all the nitty-gritty parts of buil- building an agent.
- 11:36
The most interesting thing for me with this a- uh, with this app server though is that you're not limited to sort of the classic agent chat interface. You can really be creative, whether it's, um, you-- how you actually, like, trigger it inside of your UI and your existing app, whether you're running it in the background, whether you put it on weird hardware like Maddie does all the time, um , or if you're like me, you think about, "How can I shove this into video games?" Um, and recently, um, I think
- 12:06
like a month ago or so, I thought about, how can I actually run Codex inside of Doom rather than running Doom inside of Codex? Um, and I figured, let's actually do that live, and I'm gonna dismiss a couple of messages here. Uh, so I have this Codex app, Doom version, and I boot it up in the wrong window, so let me pull that over.
- 12:29
There we go.
- 12:36
So this is... Oh, it messed up the controls. My bad. One sec. Something with this, like, dual monitor setup sometimes messes with the control capture. So let's hope that this works better now. Yep. Okay, so this is, um, an Electron app that has, uh, uses Cloudflare's Doom WASM renderer. Um, and it's, uh, so the most of the game engine itself is implemented in,
- 13:06
uh, C, and then it loads the Freedoom game, except that I actually manipula- had Codex change the game file. So you can see here, this is rendered natively in the video game using all of their, that game engine. Um, and if we actually interact with it here, we're booting up Codex, and again, this is rendered natively inside the, um, actual video game, and then passes the, passes it to the app server using Electron's, um, IPC protocol. So if we say hello here,
- 13:36
um, we can see Codex actually responding. So this is using 5.5. The fun part, though, is you can actually use dynamic tools to expose additional functionality. So naturally, um, we can get, um,
- 13:53
we can get all of the, um, armor and health, and it's calling the tools, gives me all that, or even, say, like, all the weapons with all ammo. Um, and it's calling all of the different tools. It's, so it's interacting with them. But again, r-remember, this is still an app server actually running on my machine, so if I say, like,
- 14:18
um, "What's the directory?" It's actually calling zshell, getting the current working directory, and you can see here it's r- uh, rendering. It ha- it has its own workspace. So if you want, you can vibe code directly into code, i-in-inside Doom. Uh, not sure if I would recommend it. Uh, I think the app, uh, the Codex app is slightly better for this. But you can really have fun here and, and implement your own interface.
- 14:43
Awesome. Uh, I talked with a couple of you already during the show about, uh, the app server, and one question that came up a lot of times is: What's the difference with Codex Exec? So if you've used the Codex CLI before, you're, might be familiar with this. There's another command called Codex Exec, which spins up a non-interactive interface for you to, uh, pass, uh, to kick off tasks to Codex. This is great for scripting. If you're building a CI/CD implementation and you want to script certain
- 15:13
acts like auto-fixing a PR, um, this is a great opportunity. But if you're trying to build, like, a more comprehensive client, uh, where you're interfacing, uh, where you're showing the agent, or more importantly, if you're trying to do something where you're spinning up a bunch of different, uh, Codex threads at the same time, for example, um, you're running, um, you're, you're running your own, like, eval harness, or you wanna use, like, /goal, that's where the app server really shines because you have that full control
- 15:43
over the harness, and you don't have to figure out your own way of parallelizing all of the Codex Exec calls rather than just sending off a bunch of, like, new thread events and spinning off tasks that way.
- 15:56
Before I let you all go to the keynote, um, I have, uh, three more tips just to recap, um, what we talked about, things that you should keep in mind if you're using the app server. The first one is you should actually ship your own version of the harness. Uh, this will eliminate so many headaches for you. If you're, if you're relying on the user's installed CLI, you're gonna constantly run into things from, uh, like, breaking changes, even though we are trying to minimize them. But if you're relying on experimental features, that might happen. Um, but you
- 16:26
also don't have to worry about whether a user has actually updated their version and has the latest feature that you wanna use or have them nudge the, them updating it. So ship your own version of the harness. The other part is you should absolutely expose new, um, functions as part of your interface. We do this with a Codex app, for example, for it to be able to, um, c-control itself as part of the Codex, uh, thread, like spinning up automations, et cetera. But mark them as deferred. That way, you're not polluting
- 16:56
the context window unnecessarily with all of the different additional functions that you're adding, and instead, still have them available for the agent without, um, necessarily confusing it. And if you are changing the instructions, try to start with over, uh, like, adding developer instructions rather than changing the base instructions. It might feel like you wanna immediately rush to, "I'm gonna write my own system prompt," but there's a lot of different things that we constantly have to think about when we're building a harness to make sure that the actual
- 17:25
prompt is in distribution and make sure that the agent performs well. And you're losing out on all of that if you're actually jumping in and, like, writing your own system prompt. So be careful there, uh, if you're doing that. And with that, thank you so much. Uh, thanks for taking your time. Feel free to scan that QR code if you wanna have the slides, and I'll be around for, uh, like, the next 15 minutes at the OpenAI booth if you wanna ask me any questions there. Thank you so much.