AI Engineer World's Fair 2026
From Your Laptop to the Pipeline: Scaling Custom Agents with GitHub Copilot
Read the talk
From Your Laptop to the Pipeline: Scaling Custom Agents with GitHub Copilot
Jose Palafox follows custom agents from local instructions to shared organizational tools and unattended CI workflows, where human checkpoints and execution telemetry make their work easier to control and improve.
From a talk by Jose Palafox
At a glance
Ideas worth remembering
Custom agents pair task instructions with a model choice. Exploration and second-pass review can use different configurations instead of assigning every job to the largest model.
Generated agent definitions still contain decisions worth editing: the changelog example needs source selection that excludes roadmap items from a summary of changes.
Sharing definitions and executing across repositories are separate problems. Marketplaces distribute agents; Orchestrate starts sub-agents in the repositories where their work belongs.
Agentic Workflows can combine scheduled discovery, human issue selection, automated planning and implementation, and a final human acceptance decision.
Instrumented pipeline runs make cost and tool failures inspectable, enabling investigations of expensive runs and comparisons of instructions or models.
Different interfaces, one Copilot harness
A useful agent on one laptop is only the beginning. Other developers need access to it, work may span repositories, and maintenance should continue without someone keeping a local session alive. Jose Palafox, who works with GitHub customers on adoption, opens with two operating modes: local agents used interactively and autonomous agents running in a pipeline. The talk follows the path between them.
Copilot offers several ways into that work. The IDE and CLI support direct interaction. A CLI session can be delegated to GitHub for background execution, while /remote provides a way to steer a session from a mobile device. The SDK offers another entry point: an application can expose agent functionality inside Slack or Teams, rather than requiring everyone to work through a development interface.
What stays common as the interface changes? Requests pass through the Copilot API service—the harness—before reaching a model provider. In Palafox’s overview, the harness chunks inputs, handles authentication, records metrics and audit logs, and applies checks intended to prevent repetition of open-source content. Model responses return through filtering before the interface displays them. The diagram makes the shared infrastructure visible: changing the entry point does not bypass the service between the user and the model.
Interactive development or agent functionality exposed through an application.
The interface changes where a request starts; the harness still mediates model access and the returning response.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Choose a model for the job the agent actually does
Typing /agents in the CLI reveals the embedded agents. These provide a concrete starting point for understanding customization: an agent specifies instructions and a model, allowing different kinds of work to use different configurations.
- Explore: Codebase investigation uses a smaller model, Haiku, rather than automatically spending on the largest available model. Its specialized task is to examine the repository.
- Rubber Duck: A planning review can use a different model provider from the one that wrote the plan. Palafox illustrates this with a plan written by Opus receiving a review from GPT, seeking a second perspective from different training and weighting.
That second pass is a review strategy, not a guarantee that another model will catch every mistake. The useful design choice is to separate the jobs: investigation, planning, and review need not share a model or instructions. When the built-in choices do not cover a company’s compliance requirements, custom tools, or required checks, a custom agent can encode those specifics.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Scaffold a changelog agent, then correct its task
The agent builder first asks where the agent should live. Project scope ties it to the Git project being worked on; user scope makes it available across projects. Copilot then scaffolds the definition from a short description. The same kind of building process is available through the IDE.
The example is a repeatable request: visit the GitHub changelog and summarize the changes. From that one-line intent, the builder generates a description, formatted front matter, success criteria, and a proposed methodology. This removes much of the file-writing work, but the generated methodology still deserves attention.
The generated definition proposes starting with the GitHub roadmap. That introduces a mismatch: a summary of recent changes could pull in planned work. Editing the source order to go directly to the changelog narrows the agent’s evidence to the intended task. The observable change here is in the instructions; the recording does not show a before-and-after summary. The causal point is still concrete: source selection determines what material the agent can treat as a change.
Once that definition works in the harness, the repeated request has a reusable home. Instead of reconstructing the task each time, a developer can invoke the agent. The next problem is distribution: how does the corrected definition reach everyone else?
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Start each cross-repository agent where it needs to act
Sharing an agent does not by itself solve cross-repository execution. A compliance agent starting in the repository where its definition lives may lack permission to act elsewhere. The work needs to start in the target repository with the permissions appropriate to that repository.
Orchestrate addresses that placement problem in the Copilot app. A series of GitHub issues supplies goals, and sub-agents start in the repositories where those tasks belong. The example is an API change in a microservice that also requires a downstream change: separate agents can work on the affected repositories simultaneously. Shared agents from a marketplace or organization distribution can be used for those jobs.
This adds horizontal reach to an interactive workflow. Repository maintenance raises a different question: whose identity and budget should the work use? Palafox wants project maintenance to run on behalf of the repository, rather than making every task a personal action charged to the developer initiating it. That leads to agents inside CI.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Use repository events to start work and issue comments to advance it
GitHub Agentic Workflows builds on a simple execution mechanism: the Copilot CLI can run headlessly, taking a one-shot prompt through -p. Running it inside a GitHub Actions runner—or another CI system—places the agent in a pipeline. Repository events can then start the work: a release might trigger documentation updates, while a weekly schedule might refresh context files.
Code deduplication makes the workflow more tangible. Palafox describes agent-heavy development accumulating duplicate functions and poor namespacing. A scheduled maintenance agent can examine the codebase and turn possible improvements into issues. Discovery produces reviewable work items before the system spends on planning and implementation.
In the illustrated flow, the scheduled agent finds three improvement opportunities. A developer reads them and decides which are worth pursuing. A slash command in an issue comment advances a selected issue to the next phase, where a product-manager agent generates a product requirements document and scopes the work. The plan then passes to Copilot for implementation. The developer returns at acceptance to inspect the code and decide whether it belongs in the project.
Where do people control the flow? The diagram separates autonomous discovery from the two human decisions. The first checkpoint controls whether an opportunity deserves more work and expense; the second controls whether the resulting code is accepted. A discovered issue therefore does not automatically become an accepted change.
A maintenance agent looks for deduplication opportunities.
Scheduled discovery can proceed unattended, while issue-comment commands and code review govern the next phases.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Reuse the workflow across an agent-led project
Pipeline execution removes dependence on a developer’s local session. A laptop failure or lost connection can interrupt local work; an agent running on the platform continues independently of that device. Slash commands make it possible to direct the work without sitting at a development machine.
Palafox presents the GitHub Agentic Workflows project itself as the maximalist example: he describes it as built entirely by agents, with a team goal of doing development from mobile devices through slash commands. He estimates roughly 200 agents handling different pipeline tasks. That is a reported project example and an approximate count, not a requirement for adopting the architecture.
One selected agent checks daily reports from other agents. It runs on a daily schedule and produces Markdown reports, actions, or input for another agent. The same event-driven pattern therefore supports both code work and oversight of that work: one agent’s output can become the next agent’s input.
The GitHub Next project Agentics offers a smaller starting point, with roughly 30 pipeline agents described in the talk. Examples include deduplication agents, QA reviewers, and even an agent that makes jokes. Their practical value is reuse: teams can start with an existing pipeline task rather than designing every agent from scratch.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Make agent cost and failure visible enough to improve
The ending gives a second reason to move work into CI: visibility. Encouraging everyone to use AI does not reveal whether their agents complete the task, produce excessive output, or repeatedly fail at tool calls. It also hides opportunities to replace repeated model reasoning with a small Bash script for a search operation. Without seeing execution, an organization has little basis for choosing an optimization.
In the pipeline setup Palafox describes, agents are instrumented with OpenTelemetry, or OTel, and their execution can be examined in a logging platform. The deduplication agent now becomes an observable workload across repositories: how much does each run cost, what is its average token use, and which tool calls succeed? Moving execution supplies a place to collect that information; the instrumentation supplies the information itself.
That visibility supports several distinct interventions:
- Investigate expensive runs: When an agent exceeds its expected P90 cost threshold—the 90th-percentile reference for run cost—trigger an investigation that returns debugging information. The unusual run becomes something to explain.
- Repair repeated tool failures: Use recurring failures to identify a configuration worth changing, then A/B test different front matter.
- Compare models on the task: Test whether another model produces the required output, creating an opportunity to choose a cheaper model without abandoning the task’s quality requirement.
The useful endpoint is a workflow that can be improved from its own behavior. A local definition encodes the task; distribution lets colleagues reuse it; pipeline execution creates a repeatable place to run and observe it. For the deduplication example, that means the organization can examine both the proposed code improvements and the cost of finding them, then change the agent’s configuration based on what its runs reveal.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
Related talks
- Beating RL With Reflection: GEPA and Optimize Anything
Develops the optimization step suggested at this talk’s ending: use execution traces and diagnostic feedback to improve prompts, agent programs, and repository skills.
Read the complete timestamped transcript
- 0:13
Hey, everybody. Thanks for coming. Uh, my name's Jose Palafox. I've been at GitHub for six and a half years. So I joined the team, uh, right after the Microsoft acquisition of GitHub, and then, uh, shortly after, GitHub acquired a tool called Semmle. We, we made an acquisition of a security company. And so my role at GitHub really has been growing our security business for the last, uh, five or six years. Um, in the last year, I've really kind of focused on AI, but my role is to work in the field, um, working with customers on how do they adopt our technology, how to think
- 0:42
about, um, scaling our technology inside of their organization. So, um, today, what we're gonna talk about is really just like an introduction to agents on the platform, and there's sort of, uh, two different modes that I like to think of agents in. There's like autonomous agents that I have off running in the pipeline, and then I have local agents that I'm using, uh, to do various bits of work. Um, so I'm gonna kinda walk you through sort of how to build those, how to think about them, how to think about scaling them, um, think about different workflows that are possible using agents. Um, this talk is gonna be
- 1:12
fairly high level, so if you have additional questions afterwards, you wanna chat, feel free to grab me. I'll kinda hang out, uh, right after the talk in the back of the room. You can feel free to come by and, uh, chat with me afterwards. We can talk a little bit deeper about your use cases. This is a quick preview of just kind of things that are available on the platform. Just getting us started, um, I always like to baseline with sort of the overall data flow of how Copilot works on the platform because, um, a lot of people aren't familiar with how, uh, Copilot works under the hood and, and maybe aren't familiar with all of our different interfaces.
- 1:42
So, um, just to give you a quick preview of what the product has, um, you've got a handful of different interfaces that you can use to interact with the coding agent. Um, you've got the IDE. You've got the GitHub CLI, which is sort of a lightweight interface that sits on your desktop. Um, you can take any session from the CLI, and you can stream it. So if you want to, um, you know, close your laptop or walk away from your laptop for a period of time, you can either delegate the agent to github.com and have it go run in the background, or you can stream it using a command,
- 2:12
/remote, and it'll let you log in through your mobile device and actually steer the agent, uh, from your mobile device. Um, you've also got, uh, the SDK. So if you wanna create agents that live inside of Slack or inside of Teams, um, expose some individual functionality to a set of team members that maybe don't have a GitHub license, you can do that pretty easily through the, through the SDK. So you can build applications using the CLI, uh, or the SDK. Um, regardless of what door you go through on the top, you're gonna hit this blue
- 2:42
box here, and this is what we call CAPIR, the Copilot API service. And, um, this is also referred to as like the harness, right? This is the thing that does all the, uh, difficult engineering tasks, taking your input and chunking them and sending them to the models, making sure that you're not, um, repeating any content from open source providers, uh, dealing with authentication for you, dealing with metrics logging, dealing with audit logging, all of the kinda like to-dos of the infrastructure of the platform. And then, of course, once you send a message in, we take that, and we pass it
- 3:12
to the model provider. So, um, regardless of which door you go into, you hit the harness, and then we take, uh, your output, and we send it to the model providers. Um, once we get the output back, we filter it and then display that back to you. So that's sort of under the hood how the product works, just so you kinda have like a high-level, um, understanding of the different doors you can go through and, and where you can touch it. Um, now if you're starting in the CLI, for example, you have a handful of embedded agents, and I think it's useful to kinda look at the embedded agents, um, just to
- 3:41
understand what agents do because there's like a lot of different, uh, you know, potential ways that you can implement, whether it's an agent or a skill or, you know, it all just ends up being Markdown at the end of the day. So it's useful to kinda see how, um, we're implementing agents in the platform in a way that, uh, you can utilize. So I just type /agents, and all this does is drop up the, um, the default agents that are a part of the CLI harness here. You can see I have a handful of them, right? So if I wanna explore a code base, for example, I don't necessarily
- 4:11
wanna use the biggest model possible, uh, in order to do that. So we have a specialized research agent, which is our Explore agent, um, that you can call or invoke, uh, through a keyword or actually invoking like the slash command, um, in, in, in the CLI, and it'll spin up, you know, a small model using Haiku, um, and go and investigate a code base, right? If you, um, you know, have a plan that you've done, right, and you've written out a planning document, and you want a different model provider to come and audit the plan because it has different
- 4:41
training data and different model weighting, uh, then fine, you can do that. Rubber Duck is like a embedded agent that will validate a plan that we know was written by Opus with, you know, a pass from GPT or something like that. Um, you can see in all of these models they specify, or in all of these agents they specify the model, and that's one of the options that you have when you're building custom agents. You can always control the model that you're using, and then you control the instructions that you're giving it. So if these model or these, um, agents aren't doing, you know, something that you need done, right, if you have a
- 5:11
specific task or a review or you have a compliance or, uh, regulatory requirement that you have to meet and maybe you have a tool that you wanna run or, um, a series of checks that you wanna have that are custom to you, may- maybe you wanna write an agent. And, you know, if you, uh, want, you can, you can just use the agent command here, and it will, um, set you through like a, a work, uh, a workflow to add an agent into your, your CLI environment. So here, I'm getting to pick, you know, do I want this on the project level, like in the Git project that I'm working
- 5:41
on, or do I want this as like a, um, an agent that's available regardless of which project I'm working on at the user level, right? Um, so maybe I'll say I'm at it at, at the, at the org level or at my, my, my individual user level. And then I go ahead and create this with Copilot. Copilot will do all the scaffolding for me. So, um, if I, uh, want, I can generate an agent that maybe, uh, you know, looks at the GitHub change log, and it, um, you know, will summarize all of the latest features or something that, that Copilot did. So
- 6:11
I did that earlier today and just generated a sample agent, you know, using this, uh, agent builder that's built into the platform, right? And this is available in all of the interfaces. Um, you know, if you have the IDE, you have an agent's interface there as well that will do the same kind of build, build process with you. But what's nice is it scaffolds everything out. So from like a basic one-line prompt that just says, "Hey, I want to go to the change log and summarize all the changes," it'll write the description for me, it'll format all the front matter, it will, um, structure success criteria. It'll take an attempt at structuring methodology. Like,
- 6:41
it kind of knows how to write these files. So you get to be, uh, you know, somewhat lazy with it, right? You, you just write a general description of what you want to get done. We'll fill in all the specifics for you. You can, of course, come in here and edit all of this, right? So if you don't like, uh, the order that it's, um, picking, you know, sources in this case, um, we could change that, right? Like, this one wants to start with, you know, use the GitHub roadmap, but maybe you wanted just to jump straight to the change log so you're not getting roadmap items. You know, you can just edit it here, right? There's no, no big deal. But this does all the scaffolding
- 7:11
work for you. What's nice now is you've created like an agent that works in the harness. So if you have repeat tasks that you have to do, you can call this agent to do those tasks for you, right? Um, that, that's kind of step one. Uh, but once you do that, you're kind of like, "What's the next thing that I got to go do?" And I think, um, you know, the next thing is, how do I share this agent, right? How do other people, uh, take advantage of it? And the easiest way to do that inside of Copilot is to do that, um, by making a marketplace. So a marketplace is just a
- 7:40
repository with a little bit of extra, uh, JSON in one of the directories that says, "Hey, this is a marketplace." And what it does is it gives you some light package management capabilities inside the repo. So, um, if I subscribe to a marketplace, then I can update all the agents simultaneously, I can version them, things like that. I can notify people of security vulnerabilities. You know, there's features built in that allow me to communicate with the users that are downstream from me. So creating a marketplace is oftentimes one of the first things that I work with companies on when I go in and try to
- 8:10
talk with them about how to scale Copilot adoption inside their organization. Um, a lot of times there's not a central repository for teams to even contribute agents or skills to. So if you don't have a central location, uh, you, you should generate one, right? And there is a capability basically to, um, sync those with the, uh, the local client device and then push in and out different agents. Um, you can also use a specific directory inside of, um, your organization or a specific repo. So i- if you're inside of GitHub, I'm at the
- 8:40
org level here, and then I went into a repository, and it has to be named .githubprivate, right? This is a naming convention that we've used. But if you have a .githubprivate repository, um, you'll be able to expose agent files. So in this case, I have three agents here. And you can see from the cloud agent, um, I'm actually able to call that specific agent on the platform. So I showed you earlier in the diagram, uh, you know, you can, you can interact with all of these from different surface areas. This is the cloud interaction, right? If I want to delegate to a cloud agent that's hosted on the
- 9:10
platform, I can do that here. But what's nice is this compliance bot that I'm showing you here in the, the dropdown is the same compliance bot that's in the repository, right? So this is a way of, uh, distributing agents out to the team, right? So if I write a handy agent that checks the change log for me, then I can publish it here, and I can share that with everybody else, right? So, um, this will push to everyone's IDE environment or push to their CLI environment for them, and then they can just call the same agent that I'm working on. So now I have some reuse, right? Like I can
- 9:40
publish them in a marketplace. I can publish them directly as, um, assets inside of my business that anybody can call, right? And this is what's allowing me to do standardization, right? If you look at the cost model for all of these products, the cheapest way to use them is to get cache hits, right? You want exact matches between turns on common data. And having reusable components will drive up your cache rate, right? So if cost is a concern, um, building reusable agents and making sure that people are solving problems in the same way, uh,
- 10:10
rather than reinventing the wheel every time, it's a good way to drive down your, your overall bill, drive up your cache rate. So again, this, um, .githubprivate directory, this is how you would share an agent internally to everyone inside the org, right? So you can make a marketplace that will work across organizations, that could even work, you know, externally, right? So we have a repository called Awesome Copilot that has a bunch of, um, you know, example agents and things. You could register that as a marketplace. This is purely internal, right? And this is like an admin can push this out.
- 10:41
The last thing that I want to do sometimes, um, when I'm working, uh, with an agent is I, I want to be able to use, um, agents across repositories. And the biggest challenge with these, uh, compliance agents, for example, is that they're always starting in the repo where the agent lives. So if I have agents living in this repository, um, they maybe can't take action in other repositories, right? I need to actually start the agent in the other repo so that it has permissions in that repository. Um, inside of
- 11:11
the Copilot app, there's a command called Orchestrate. Um, and Orchestrate is our horizontal scaling for cross-repo. Uh, so this will allow you to set a series of GitHub issues as, uh, goals for the agent and then spin up sub-agents to address all of those tasks. But what's nice is that it will spin up agents in the repository that it's supposed to be working in. So if you want to take a task inside a microservice where you're like modifying an API, and then you have to make a change somewhere downstream because of that,
- 11:41
Orchestrate will allow you to set up several agents that are gonna make, uh, cross-repo changes simultaneously. So I can use the, um, the pre-canned agents, right, if they're available to me in those other, uh, marketplaces or being pushed to me by my org admin. I can kind of summon them here to go do work in individual repositories and keep them sort of on task using this, this Orchestrate tool. And that's really, like- Three layers of how do I use agents locally, right? Like, and, and using agents locally has a lot of utility, right? Um, but,
- 12:11
uh, sometimes I don't wanna run an agent locally. I actually want the agent to be run as a part of the repository, and certain agents maybe are gonna be performing, like, maintenance tasks in the repo. And so I wanna run them, um, in the repo, uh, on that repo's budget, not on my budget, right? Or I wanna run it using the repo identity, not my identity, right? Not every agent needs to be me taking an action. Um, you know, it's, it's maybe a maintenance task that has to be done on behalf of, of the project. Um, when we're working in that
- 12:41
model, we have a framework that you can use called agentic workflows. And agentic workflows is a pretty simple concept, like the Copilot CLI can be used in a headless mode. So I can call the Copilot CLI, I can use the -p command and basically pass it a single, uh, prompt, a one-shot prompt for it to go do. And then if I take that and I run that inside of a GitHub Actions runner or in some CI system, then I have an agent that's running in a pipeline for me, right? So I can trigger it on events that are happening in the repo, like
- 13:11
on a release, I wanna update, um, I wanna update documentation, or every week I wanna update my context files, or, you know. In maybe a model like this, I have various maintenance tasks that I wanna go do. Like, one of the biggest problems today that's happening in repositories where a lot of agentic development is happening is there's lots of spaghetti code getting added, lots of duplicate code. Like, the agents are lazy about, um, implementing, you know, uh, functions, and they're really lazy about namespacing, so they'll just, like, create constant duplication inside the code base.
- 13:41
So maybe you want an agent that, uh, looks inside the code base, um, identifies opportunities for, uh, deduplication of code, and then creates a series of issues for you. And if you use this framework, this agentic workflows framework, what's nice is we'll, we'll build in human-in-the-loop triggers for you so that you can review and, um, decide when you're actually gonna spend money on sort of the next phase of the flow. So in this flow here, you see I have the first agent maybe running on a cron schedule. It's looking in the repo. It
- 14:11
identifies, you know, three items where there's opportunity for me to make improvement in the code base. Okay, great. The developer reads that and then makes a decision about whether or not those things are worth fixing, and they express that through a slash command that they can, um, make as an issue comment, right? So they, they go to the issue and they say, "Okay, advance this issue to the next phase of the process," right? For us, that's maybe a planning phase, right? So maybe I'm gonna call another pipeline agent, right, which is the PRD-generating product manager agent, right, that's gonna scope the
- 14:40
feature for me. It generates a plan. It hands it off to Copilot in the platform for implementation, and then the developer comes back in the loop maybe at an acceptance phase, right? Let's look at the code and evaluate whether or not we actually wanna take it in the project. Um, now I've got pipeline agents, and what's really cool about the pipeline agents is how they scale. Um, you know, the, the, the, the challenge with working locally is you're on, like, a single session or you're on your device. If your device goes down, uh, or you lose network activity, you know, your, your workflow is
- 15:10
interrupted. These are running in the platform, so, uh, they're running autonomously from you. If you wanna see, like, a maximalist approach to how this works, the GitHub AW project itself, like the thing that, um, is the framework for building all these agents, it is built entirely by agents, and this team has sort of an internal goal that they're only gonna do their development on their mobile devices because they interact entirely with this application, uh, through those slash commands, right? Um, so just to give you a feel of, like, what a fully, you know, agent-led development
- 15:40
project looks like, this is, this is our, uh, our, our framework, right? And it has, I don't know, 200 agents in here that are all running different parts of the pipeline, right? Um, and all of these follow, uh, a similar, uh, workflow that we just showed on that, um, on that slide, right? That architecture is reusable. So, um, let me see if I can just pick one of these. I, I'm a little bit picking at random here, but, um, let's see what this one does, right? So this is, uh, checking daily reports
- 16:10
on other agents, right? So this is a meta-agent that's looking on things. It's running on a daily schedule, and it has, you know, its agentmd file that's working away here, right? So every time, uh, this thing runs, it's generating some markdown report or some action or input for another agent, and then we're just continuing to trigger it. Um, you can feel free to take a look through here. There's hundreds of these that are creating this project. GitHub also has, um, a bunch of them that we've bundled up. So inside of our GitHub Next organization, there's this project
- 16:40
called Agentics. Agentics has 30 or so different pipeline agents that you may be interested in just pulling off the shelf and running in your project. Um, some of these will do the code deduplication that we talked about. Um, some of them are different types of reviewers. So if you need a QA reviewer or, um, you know, you want one that's funny and just makes jokes, like, there's all kinds of, uh, different sample agents here that you can pull and drop into the pipeline. But what's nice about this is it gets the workload off of your machine, and I think when we shift back to thinking about cost, um,
- 17:10
getting the workload off of your machine is key for optimizing cost, right? If I'm, uh, having everyone in my organization, uh, work with AI, and I'm telling them to go use it, but I can't see how they're actually working with the tool, I don't have any opportunity to optimize, right? I can't actually see, um, whether or not the agents that they're using are too verbose or are even achieving the task. I can't tell when tool calls are happening, if they're successful or unsuccessful. Um, I, I, I can't tell if they're using agent reasoning for all types of
- 17:40
things or if they're, um, you know, cor- correctly, uh, building little bash scripts to actually go and do, uh, different search functions and things like that. Like, there's lots of optimization opportunities for me, but if I can't see it, I can't do any of that optimization. So by getting the workload off of the laptop and moving it into CI, I get to see what's happening, right? I have every single agent running, uh, and it's instrumented with OTel, so I can go to a SIM or some kind of logging platform and look. Does my code deduplication agen- agent, how does it
- 18:10
run across repositories? How much does it cost me, right? I can pivot on average token use. I can also do things like say, "Hey, every time that an agent exceeds its P90 for what I expect it to cost me to run, uh, go investigate why it exceeded its, i- its threshold," right? I want it to tell me, uh, what debugging information I can, I can take to, um, to evaluate it further and improve it, right? Or if I see that a tool call's failing over and over, maybe I wanna A/B test different front matter, or maybe I wanna A/B test different models and
- 18:40
see if they produce, like, the, the right output for me so I can pick a cheaper model at some point, right? All of this stuff is possible once I get the agents into the pipeline. Anyhow, so this was a quick and dirty kinda tour of agents in the platform, right? We talked through, um, building an agent locally, sharing agents through the platform, and then eventually moving those agents into the pipeline. Um, if you wanna talk more about any part of that, I'll be in the back. Feel free to come up and grab me, or, um, you can add me on LinkedIn and, and ping me. I'm happy
- 19:09
to continue the conversation. Um, I work at GitHub in a field capacity, so if you know your GitHub AE and you wanna chat with me, uh, feel free to reach out with them. Uh, we, we can get in touch and, and have a quick conversation as well. Thanks so much for your time today.