AI Engineer World's Fair 2026
Designing CLIs for Agents, Not Humans — Pedro Lopez, Airbyte
Read the talk
Designing CLIs for Agents, Not Humans
Pedro Lopez explains how Airbyte exposes the same data platform through MCP and a CLI—and why discovery, authentication, structured output and client limits determine which interface works for an agent.
From a talk by Pedro Lopez
At a glance
Ideas worth remembering
Keep interfaces thin over shared APIs, and use OpenAPI specifications to help them follow a frequently changing platform.
Progressive discovery keeps initial context small: MCP retrieves connector schemas on demand, while CLI skills load detailed command references when needed.
MCP authentication must fit real client UIs and long-lived session expectations; protocol features such as URL elicitation still depend on client support.
An agent-oriented CLI uses structured input and output, predictable command names, discoverable skills and noninteractive configuration.
MCP makes conversational access convenient and updates centrally; a CLI supports longer tasks and Unix output processing, while adding package-versioning and runaway-loop concerns.
One data platform, several ways to reach it
Company data lives across tools: support tickets in Zendesk, payments in Stripe, contacts elsewhere. Airbyte’s move from data movement into agent tooling builds on connecting those systems. In this talk, Pedro Lopez, a software engineer at Airbyte, introduces the Context Store: a search-optimized index between those sources and an agent, designed to let the agent retrieve and combine information across them.
An SDK serves developers building their own agents; a web UI provides direct access. MCP and the CLI bring the same platform into existing agent environments such as ChatGPT and Claude. Each interface needs to support three kinds of work:
- Manage resources. Work with organizations, workspaces and connectors.
- Connect services safely. Authenticate third-party systems without exposing their credentials to the agent.
- Ask and act. Search the Context Store, read records and write changes to connected systems.
MCP and the CLI remain thin interfaces over shared platform APIs. They accept a request, call the platform and format the result. That separation matters when the platform changes more than ten times a day: duplicating its behavior in each interface would create several implementations to keep synchronized. OpenAPI specifications supply the API source of truth, and much of the interface code builds on those specifications.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
From available connectors to five paying customers with tickets
MCP offers a convenient entry point for someone who wants company data inside a chat application. Airbyte demonstrates its ChatGPT listing, while installation outside a marketplace can use a hosted MCP URL. The user connects a service to an existing conversational interface rather than setting up a command-line workflow.
The demonstration starts with a small discovery question: which connectors are available? The response identifies four connectors in the workspace. The next request is more useful: “Show five paying customers with Zendesk support tickets.” Answering it requires information from two systems. The agent searches the Context Store, combines Zendesk support information with Stripe payment information, and returns five people with open support tickets.
The observable change is from a list of available data sources to a cross-system customer answer. Connector discovery establishes what can be queried; retrieval supplies the relevant records; cross-referencing combines the payment and support conditions. The demonstration establishes that result, but does not specify the identity-matching rule used to associate records across Zendesk and Stripe.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Describe capabilities before executing them
A connector can expose many entities and actions. Multiplying those by all available connectors could produce a large catalog of MCP tools, with a description for every operation. Airbyte instead keeps the connector execution interface to two main tools:
- Describe. Retrieve the connector schema so the agent can learn how to query it.
- Execute. Perform an operation, including reading, writing or deleting records.
The choice moves detailed knowledge into progressive discovery. The agent learns the relevant connector’s capabilities when it needs them, rather than receiving every entity and action in its initial tool context. Lopez reports that keeping descriptions narrow helped MCP performance; the talk gives a design observation rather than a quantified comparison.
Where does the operation-specific information enter the agent’s context? The diagram shows the extra discovery step between a compact tool interface and execution. The schema carries the detail that would otherwise have expanded the initial tool catalog.
Starts with a small connector execution tool interface.
Two general tools expose connector-specific capabilities on demand, keeping the initial descriptions small.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Authentication has two jobs—and client support decides the flow
First, the user must authenticate to Airbyte’s MCP server. In the demonstrated OAuth flow, clicking Connect opens an Airbyte sign-in, the user accepts the requested privileges, and the connected services become available. Other authentication mechanisms are technically possible, but the client’s installation UI may make them impractical. The Anthropic custom-connector registration shown in the talk provides no place to enter a bearer token, so OAuth is the usable path there.
Session duration is part of that experience. Users expect MCP access to behave like a mobile app: available when needed, without repeated sign-ins. Airbyte therefore keeps long-lived offline sessions separately from web-app sessions, with their own lifecycles. Reusing a web login’s lifecycle would miss the expectation that the agent connection stays available independently.
Second, the user must authorize the third-party service itself. For a request to connect HubSpot, Airbyte generates a link. Opening it displays a widget where the user selects entities to connect, then redirects to HubSpot’s OAuth flow. This gives the user a place to complete sensitive authorization outside the agent’s conversation.
MCP elicitation offers a protocol mechanism for requesting information from the user. Form-mode elicitation is unsuitable for sensitive credentials; URL-mode elicitation is closer to Airbyte’s link-based flow. Yet specification support does not guarantee client support. In the recorded Claude Desktop example, URL-mode elicitation is unsupported, so Airbyte uses its own link flow. That limitation applies to the demonstrated client behavior, rather than establishing that URL elicitation cannot work in other clients or later releases.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The same customer question, through JSON and predictable commands
The CLI demonstration repeats the customer question. This time, the agent loads a skill describing the CLI, uses commands to discover the available connectors, queries the Context Store and aggregates the data. It returns the same five users. The data platform and task remain the same; the route into the platform changes.
Designing that route for an agent changes what counts as convenient. Airbyte accepts JSON input and returns JSON output. For complicated query inputs, a structured object lets the agent express the request directly instead of guessing how to distribute it across many flags. On the output side, JSON preserves structure that another program can select and transform.
That output format makes Unix composition useful. A response can flow through a pipe into jq, which selects fields or trims the result. A long response therefore does not have to enter the agent’s context in full: another program can reduce it first. This is a practical advantage of the CLI environment, where the agent can work with both the data interface and the surrounding operating-system tools.
Command consistency reduces another kind of guesswork. Airbyte uses a noun followed by a verb: connectors list, connectors create, and workspace list. Once the agent learns the pattern for one resource, it has a useful expectation for another. The goal is to make the right command “dead obvious” without requiring repeated documentation searches.
Repeated agent mistakes are also design feedback. If agents keep calling the same nonexistent command, consider making that command the supported interface. A recurring error may reveal that the naming scheme is less obvious than its designer thought. The proposed response is to improve the command vocabulary so agents can more easily find the intended operation.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Ship discovery instructions, and never wait for a prompt
Installing a CLI does not, by itself, tell an agent that the tool exists or when to use it. Airbyte ships a skill with the CLI to supply that missing discovery layer. Its description makes connector-related requests a reason to load the skill and invoke the tool; its instructions help the agent use the CLI effectively.
The skill needs the same restraint as the MCP tool catalog. A large main skill would consume context before the agent knows which details matter. Airbyte separates detailed command information into reference files and lets the main skill point to them. Both interfaces therefore use progressive discovery: expose enough to find the capability, then load the detail needed to use it.
Interactive prompts introduce a different failure mode. A human-oriented CLI may stop to ask for an answer, but an agent invocation cannot assume a person is present to type one. The process can simply get stuck. Inputs must be supplied through flags or other noninteractive configuration; credentials can be prepared in environment variables or configuration files before invocation. JSON handles complex query data, while these mechanisms handle setup that should not interrupt execution.
This requirement extends beyond authentication. Any command that pauses to ask a human a question can block the agent’s work. Preparing configuration out of band lets the command run with the information it needs already available.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Choose the execution environment as well as the interface
The closing comparison adds operational constraints to the design choices. MCP features vary across clients, and providers impose limits beyond the protocol itself. Lopez cites a two-kilobyte limit on server and tool-call descriptions in Claude, along with limits on tool execution duration. These are reported constraints from the environment discussed in the recording, not universal MCP limits. Execution deadlines make long-running processes harder to fit into that interface.
A hosted MCP server has an update advantage: the provider can update the server centrally. A CLI package installed elsewhere needs versioning and distribution, and keeping those installations current takes work. The CLI’s freedom from the same client-imposed limits also has a cost. Repeated calls can go wrong and leave the agent spinning in a loop; fewer external restrictions do not automatically produce better execution.
The useful division follows the task and user:
- MCP for convenient conversational access. Nontechnical users can connect data, prototype quickly, build one-off reports and work within the MCP ecosystem.
- CLI for longer work and large outputs. Technical users can run longer tasks and use Unix tools to redirect results to files, inspect the beginning with
head, inspect the end withtail, or transform structured output through pipes.
The customer example shows why both can belong in one product: each reaches the same Context Store and returns the same five customers. The decision turns on how the agent discovers capabilities, how the user authenticates, where execution happens, and what must happen to the output afterward.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Read the complete timestamped transcript
- 0:12
Hello, everybody. Uh, my name is Pedro. I am a software engineer at Airbyte, and today we're gonna be talking about how we built our Agent MCP and our CLI. Um, and before we go into the nitty-gritty of all of that, um, I wanted to give a bit of context on what we do at Airbyte and how all of this kinda fits into, uh, our product. So Airbyte's been around for a while. Um, we started as the open source, uh, standard for data movement,
- 0:42
um, connecting all of your SaaS tools that you might be using, like Zendesk, Stripe, uh, GitHub, et cetera, to, um, your data analytics pipelines. And now we're using that same expertise and have evolved to being the data and action layer for your AI agents. Um, and we do that by offering a, our Airbyte Agents product, um, that has what we call the context store, which connects to all of those tools that, that I mentioned before and sits in
- 1:12
between your agents, um, to provide a search optimized index of all of your data, um, that is able to be more efficiently queried by your agents. So you're able to gather information about all of these, um, records and, and, and contacts and stuff that you might have in these all, all these tools and join them together in an easy way, uh, through our context store. And so our goal is that we wanna make data available and actionable
- 1:41
everywhere. And so we have a couple different interfaces by which you can access this context store. Um, one is our SDK, which is more usable if you are, uh, building your own agents and your tools. Um, our UI, which is, um, just a, a quick way that you can use it through our web app. Uh, and then what we're gonna be focusing on in this talk, which is our MCP and our CLI, which is, um, a way to easily make that, um, available within existing, uh, agents like ChatGPT,
- 2:12
um, and, and Claude. So what exactly are we exposing through these interfaces? Um, three main things. One is the ability to manage all of the, the core resources of our product to our organizations, workspaces, and connectors. Um, and so all these interfaces should let you, uh, interact with that. Second, um, super important is be able to authenticate those third-party services that I mentioned about earlier, um, and do so in a way that's secure and doesn't expose your
- 2:42
credentials, uh, to, to the AI agents. And third, um, ask and act. So we wanna expose the ability to search across the context store, uh, be able to read and write, uh, take action on all those systems. Um, and so we want all of our interfaces to be able to span across these, uh, three capabilities. Um, so the MCP and the CLI, the way to-- The way that we've structured things is that they are s- thin unified interfaces into our
- 3:12
platform, not re-implementations of the platform. So most of the surface layer of these, uh, interfaces are presentational, where you take a command, uh, and you call the platform and format the results in some way. Uh, and everything kinda sits on top of our shared platform and APIs. Um, and because we wanna-- We deploy a lot, you know, our, our platform changes like over ten times a, a day, we wanna keep these things up to date as we go. And so a big part of this as
- 3:41
well is OpenAPI, OpenAPI, uh, specs as a source of truth for how to connect to, to our APIs, and a lot of those interfaces, um, are built on top of, uh, those specs.
- 3:56
Okay, so let's go into these interfaces. Let's starts with-- Let's start with the, the MCP. And so MCPs are really great for users that might not be technical. Um, and in our case, it's great for those users that just want to plug their data straight into ChatGPT or Claude in an easy way. Um, and the cool thing is they're, they're very nicely packaged. They're very, uh, easy to install. For example, here's our, um, listing on the ChatGPT app, uh, listing, the official, uh,
- 4:26
store that they have. Um, and we also have a application for the, the Anthropic one, uh, which we're, we're still pending on, if there's anybody from Anthropic out there that wants to approve us. Um, but even if you're not on these marketplaces, um, they are pretty trivial to install, uh, for a, an end user. You just give it the hosted MCP URL, and you're able to, um, connect pretty quickly. Okay, so let's look at, uh, an example.
- 4:56
This is our MCP in action. Here I'm asking, you know, which connectors do I have available, um, and it goes through, "Hey, I have four connectors set up in this workspace." And then I can ask it some questions about my data, like, "Show five paying customers with Zendesk support tickets." Um, and so this right now is looking through all the connectors that we have, searching the context store for that information that I, that I asked it about, um, and then combining that information across Zendesk and
- 5:26
Stripe, um, so that it can come up with an answer for me. Okay, and there we go. So it cross-referenced Zendesk and Stripe and then came back with, you know, five of those people that, um, have open support tickets. Cool. So let's go into kinda some lessons learned while building this, uh, MCP and things that you might wanna keep in mind if you're doing, uh, something similar. So one thing that we decided to do, um, after trial and error is to keep the tool count small. So,
- 5:56
um, we-- For the execution part of the connector where you're pulling that data, we only expose a really- Deal with two main tools. One is to describe the connector to get the schema of how do you query that information, and two is execute, that's able to read, write, and delete records. Um, the opposite end of the spectrum here would be since we connect to so many different, um, connectors and each of these has different entities, different actions you can take, we
- 6:26
could have exposed, you know, an MCP tool for each one of these. Um, but instead what we found is really keeping it narrow, using progressive discovery of all of the capabilities, um, that the connector has available, um, and really limiting the amount of context that you shove into the, those tool descriptions, um, really helps with the, the performance of the, of the, the MCP server. Two, OAuth is the way. So this, this refers to
- 6:56
how you authenticate the MCP server itself. Um, so this is, uh, an example for, like, Anthropic. You click a button, uh, to connect, and that spins up our OAuth flow so you can sign into Airbyte, um, accept the, the privileges, and then you're all set up with all the connectors that you've, uh, linked up. Um, and technically you can use other authentication mechanisms through MCP, but you're really gonna have a hard time, um, if you do, like, Bearer Auth or, or anything like that 'cause the
- 7:26
spec is really heavy on OAuth and all of the clients, the, the main ones at least, um, are really leaning towards the, the OAuth as a, a better UX for, uh, installing that connector. So that's an example for, uh, the, the Anthropic, um, custom connector, um, registration where they have no, no way of you to put, like, a, a Bearer token or anything. It's just all defaulting to, to OAuth. So that's is, is very important. And two,
- 7:56
users actually expect, um, to have long-lived sessions, uh, through these MCP servers. Um, unlike a web app where you might be okay with having to sign in again, um, a couple times, uh, users expect to use MCPs more like a mobile app where they don't really have to think about authenticating. It's just available, uh, whenever they wanna use it. And so that has its own little, um, you know, uh, things to keep in mind as you're implementing OAuth is keeping offline long
- 8:26
sessions, um, that are stored separately from the web app and have its own set of, uh, of life cycles. Cool. Elicitation, though, is not the way. So I'll get into what an elicitation is, but a big part of, um, the MCP server that I mentioned is you need to be able to connect your tools. You need to be able to give credentials in a way that's safe. Um, and to do that, uh, basically when you ask, you know, "How do I connect to, to HubSpot?" um, we generate a
- 8:56
link that you go to which you can... pops up this widget, and you're able to select, you know, which, uh, entities you wanna connect and redirect to, to HubSpot to go through the, that OAuth flow. Um, now technically, MCP has something in the spec that allows you to surface those kinds of URLs, um, and, and it process those inputs, um, called elicitation. It's a way that you can request information from the user. For the
- 9:25
longest time, um, it was only form mode elicitation, which is not, um, applicable for sensitive, um, sensitive, uh, data like credentials for an API. But more recently, um, this URL mode elicitation became a thing, um, that's very similar to what we're doing. But the problem is, uh, support across the MCP clients. So that's a big theme, uh, when working with MCP, is that even though there is a spec, the features that is, are
- 9:55
supported across the clients kind of vary quite a bit. Um, so this is an example from Claude Desktop where, you know, it just is telling me, "Hey, I don't support URL mode elicitation. I can't do that." And so we have to go back to doing our own thing, um, and not using the spec for that, for that kind of thing. Okay. Let's talk about the CLI a bit. So if MCPs are easy for these non-technical users, um, CLIs are kind of the more technical counterpart, might call it the kind of more powerful
- 10:25
i-in some ways that we'll get into. And this is the, the same example we saw earlier but this time with the, the MCP... or, or sorry, with the CLI, where we're asking, you know, which connectors do I have available. It's this time loading a skill that tells it about the CLI, um, and then, you know, using the CLI itself to query things. We got back the connectors, then we asked, you know, "Show five paying customers that have Zendesk support tickets," same as last time. Um, it's able to this time use the CLI to query that,
- 10:55
query the context store, aggregate that data, and then eventually it should come back with a result now. Okay. Um, so there's those same five, um, five users. So, um, on the design side for CLIs, I think one important thing to keep in mind is to design those CLIs with agents in mind. Um, some of the things that might be helpful for users when you're making a CLI are gonna be, like, blockers for
- 11:25
agents, um, and we'll get into that a little bit more. Um, the first design decision here that we made was take JSON in and return JSON out. So, uh, for humans it's very nice to specify, um, a bunch of flags, uh, to work with CLIs. But when it comes to these more complicated, um, inputs that, uh, you have to build to, to make these queries, and agents are
- 11:55
really good at working with JSON. Um, they're able to construct this, uh, in a much more, um, direct and, and less of a guessy way, um, than using those flags. And so on the output- JSON also makes it easy to deal with longer responses. Um, be able to leverage one of the key capabilities of CLIs, which is, hey, you have these, this Unix system with pipes that you can redirect the output to
- 12:25
different, um, programs. So you can use jq, for example, to select certain, certain things from the output and trim that. And so JSON really unlocks a lot of that, um, capability. Two, consistency is very much key here. So it should be dead obvious for agents to know how to do something across all of the resources that you're exposing. If you have structured your CLI in a way that it's, it's structured where there's a
- 12:55
noun and then a verb, phrase, this example, you know, connectors list, connectors create. Make sure that when you're implementing a workspace, um, CLI, uh, commands, they also follow that same format. So workspace list is another one that, that we have. And so it's very, um, critical that agents know the right way to do things, and it's easy for them to know the right way to do things without having to read a bunch of documentation. And another interesting fact here is,
- 13:25
you know, if, if you see agents making the same mistake over and over, um, calling the wrong command a bunch of times, consider making that your contract. Make it dead simple for agents to stumble across the, the right way to, to use things. Um, 'cause that usually means that something is not quite as obvious as it should be in your AP- in your CLI contract if agents keep making the same kind of mistakes. Three, skills are-- skills ship with, uh, the CLI tool. So skills
- 13:55
are a must, uh, when you're doing CLIs because how else, if you have an, a CLI installed, is an agent gonna know that that CLI even exists in the first place? So it helps a lot with discovery of, uh, your tools that are available. Um, and they also help with, uh, using that tool effectively. So, uh, we have a skill that gets surfaced. It has a, a description that is able to come up whenever you ask anything about connectors, it knows to bring up that CLI and invoke it.
- 14:25
Um, and then the way that you structure the skill is also something to, um, think through. You know, you don't wanna put too much into the skill, uh, that you're gonna bog down the context. Um, skills have this, uh, ability to be separated out into references, um, where you can have a main skill that points to other files in a, in a skill directory, um, that can expand information about the, the different commands. And that really, uh, helps cut down the, the amount of, uh, of context that you're loading in and, um,
- 14:55
leverages the progressive discovery, um, ability of, uh, of these agents. Cool. And then another interesting fact, uh, for Auth side. Auth was interesting on the MCP side. It is interesting in other ways for CLIs because we can't assume that a human is gonna be present to able to interact with those prompts. Um, when CLIs are being invoked by agents, it can't sit there and wait for you to type, um, a prompt into the CLI. You have to
- 15:24
expose things either through flags, um, or in the case of credentials, environment variables, and config files, um, so that things can be set up out of band and you don't have to wait for that, uh, whenever you're invoking the CLI. Um, and I would say aside from Auth, um, that's a general thing. A lot of CLIs that are built for humans, you type a, a command, and it'll wait there, ask you for prompts. Uh, and that is not, um, you know, able to be used properly by these
- 15:54
agents. Um, they'll just kinda get, get stuck there. Um, cool. So CLIs and MCPs, which one is better? Well, they both have kind of their own trade-offs. You know, the, with the MCP, the spec is not uniformly implemented across the board. Um, there's a bunch of limits that these providers decide to implement aside from the spec as well. For example, like Claude has a two-kilobyte limit on server and tool call, uh, descriptions. There's
- 16:24
a limit of how long the tools can run, so it's harder to do more long-running processes through MCP. Um, and uh, you know, there's, there's other kinda, kinda limits that they've introduced there. And then on the CLI, the cool things about MCP, like, you know, you could have a hosted MCP server that just is always updated. That's not quite the case with the CLI. You have to think more about versioning. It's harder to keep packages up to date, uh, with all these new features. Um, and there's no-- while the limits can be,
- 16:54
well, limiting on the MCP side, not having any of those on the CLI across the tools, uh, a-across the, the calls of the CLI can make things go wrong and continually go through loops that, um, you know, make the agent spin. And so really, I think both have their place. I think MCPs are great for those non-technical users to be able to quickly prototype things, build one-off reports, um, and even leverage that MCP ecosystem
- 17:24
to have things, uh, talk to each other. And then the CLI is better for those more technical users, more longer running tasks, maybe outputs that are super long and you need to leverage those, um, features that Unix and, and the pipes give you, where you can redirect output to a file, you can do a tail, do a head on that output. Um, that's where really the, the CLI, um, starts to, to shine. All right. That is basically all that I had.
- 17:54
Um, we have a booth over there at Airbyte, um, that you should check out and come work with us. Thanks.