AI Engineer World's Fair 2026
The Software Factory: From Bug Report to Production Code — Davis Palmie, Factory
Read the talk
The Software Factory: From Bug Report to Production Code
Davis Palmie explains how agents can connect incident triage, planning, code, tests and deployment—and why the path to autonomy starts with shared context, narrow responsibilities and human judgment.
From a talk by Davis Palmie
At a glance
Ideas worth remembering
A software factory connects incident signals, planning, documentation, implementation, review, testing, deployment and monitoring; generating code is one step in that journey.
Start with a narrow, verifiable responsibility such as turning a Sentry alert and traces into a diagnosis for human review. Establish ownership and monitoring before connecting steps into a loop.
Model choice should balance task effectiveness, speed and cost, while deployment fits organizational policies and shared context connects the development lifecycle.
Documented builds, tests, reproducible environments, modularity, observability and security checks improve working conditions for both agents and engineers.
Humans retain architecture, strategy and validation responsibilities. Measure signal-to-production time, interventions, repair time, code shelf life and cost per PR rather than rewarding token consumption.
From completing code to governing agents
An autocomplete tool leaves the engineer responsible for nearly everything around the next token. A file generator moves more implementation into the model, but still leaves a human selecting and integrating its output. An agent changes that relationship: it can gather context, call tools and debug what it produces. Davis Palmie, a member of technical staff at Factory, opens with this progression to explain the next step: organizing agents into a system that takes bug reports and user feedback through to production code.
The abstraction ladder runs from tokens to files, then context and entire systems. Each step changes what the engineer controls. In the proposed software factory, engineers maintain guardrails, catch code drift and set priorities while agents execute work across the development process. The responsibility moves toward governing the machinery that produces software.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The first questions concern cost, model choice and responsibility
When agents produce a growing share of the code, model spending becomes an operating decision. The useful target is the least expensive model that meets the task’s effectiveness requirement. A routine task need not consume the same resources as a difficult one. Palmie preserves the point with an analogy attributed to Factory’s CEO: an algebra tutor can be cheaper than Albert Einstein.
Four questions frame the transition:
- Cost control. Access to multiple providers creates room to choose among different prices and capabilities rather than paying one rate for every kind of work.
- Access to capable models. Model selection can happen per task, potentially through dynamic routing, instead of requiring an organization to commit permanently to one provider.
- Engineering responsibility. Engineers build and maintain the system responsible for producing code, and set its direction. That brings product decisions closer to engineering work.
- Organizational responsibility. Context must cross team boundaries. Product, engineering and other teams share responsibility for the information and practices the factory depends on.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Connect the work around coding, starting with incident triage
“Coding was the easy part” is Palmie’s provocation. To test it in an organization, look at the time spent reviewing pull requests, reproducing bug reports, debugging, updating documentation and maintaining tests. His claim that the wait from written code to production can dwarf implementation time is a hypothesis to check against those engineering metrics, rather than a measured result presented here.
Those delays are connected. Documentation can fall behind code or strategy; tests can reflect assumptions that product plans have already changed; dead code can remain because nobody has enough context to remove it. An agent operating across these teams needs to bridge their services and understand the organization’s conventions. Otherwise, faster implementation feeds work into the same disconnected review, testing and deployment process.
A software factory therefore spans the journey from an input signal to production deployment. Its responsibilities include incident triage, planning, documentation, implementation, review, testing, rollout and release monitoring. That scope is also the reason to begin narrowly: automating everything at once gives teams little opportunity to understand individual steps or build trust in them.
Incident triage supplies a concrete starting point. A Sentry alert arrives. The agent reads it, gathers traces and other relevant context, develops a diagnosis and posts it in Slack. The observable change is that the engineer receives a contextual diagnosis rather than only the original alert. The engineer still reviews the diagnosis and decides what to do. This is a proposed rollout pattern, not a demonstrated incident with a reported repair result.
Where does this first automation stop? The diagram places the human decision after the agent’s diagnosis. That stopping point makes the agent’s contribution inspectable without immediately granting it permission to change production. Repeated correct diagnoses let the team learn how this step behaves. Each later step needs its own ownership and monitoring before the pipeline becomes a loop; eventually, connecting production results back to the original signal can support improvement across the organization.
An incoming incident signal starts triage.
The initial automation gathers context and posts a diagnosis. Engineers retain the decision about what happens next.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Choose models by task, deployment by policy and tests by intent
Model independence creates a three-way choice among speed, task effectiveness and cost. The Pareto frontier is the set of useful tradeoffs: improving one dimension may require giving up something in another. Access to several models lets the factory choose the tradeoff for each task instead of accepting one provider’s position for every task.
Factory reports comparable code-review performance from models priced differently, including a comparison at half the price, and describes open-source alternatives as potentially 10 to 30 times cheaper. These are reported task-specific comparisons without an evaluation protocol or deployment-cost breakdown in the recording; they motivate evaluating cheaper options rather than establish equivalent performance or savings for every workload.
Two further principles determine how that flexibility fits an organization:
- Sovereign deployment. The factory should fit existing security architecture and policies. Palmie describes Factory’s supported deployment choices as fully managed, bring your own machine and entirely air-gapped.
- Integration across the software development lifecycle. Documentation changes should inform code, code changes should inform documentation, and planning goals should inform tests. The intended behavior must survive the trip through strategy, product, engineering and testing.
The lifecycle principle makes the factory more than a collection of independently useful automations. Each part depends on the others: well-executed code can still implement outdated documentation, and passing tests can still validate the wrong product goal. Weakness in one part becomes weakness in the system that carries a request to production.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Eight pillars make a repository usable by agents and engineers
Building a factory exposes tribal knowledge, manual processes and outdated information. Factory calls the preparation work agent readiness. Its eight pillars describe whether an agent can work in a repository, check its changes and understand failures without depending on knowledge trapped in someone’s head.
- Validation. Linters and formatters provide guardrails for changes.
- Build system. Documented build commands and CI make the build process understandable.
- Feedback loops. Unit and integration tests give changes a way to receive feedback.
- Documentation. READMEs and
AGENTS.mdfiles supply written project context. - Development environments. Reproducible environments let an agent enter and use the project reliably.
- Modularity. Clear code boundaries and rules against sprawl constrain how changes spread.
- Observability. The team can find out why something went wrong.
- Security scanning. Proactive checks look for leaked secrets and vulnerabilities.
These improvements help the people using the repository too. A documented build removes guesswork for a new engineer as well as an agent; useful tests and observability make failure easier to investigate for both. Palmie’s road analogy captures the shared investment: if humans and agents walk the same roads, paving them well is “doubly useful.”
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Automate individual steps before closing the loop
Readiness supplies the working conditions; governance controls behavior once the agent starts. Engineers need visibility into performance and actions. Palmie describes Factory actions as auditable and governed through organizational standards, role-based access control and least privilege. Those controls answer different questions: what happened, who or what may act, and how much access the task requires.
The rollout sequence remains incremental. Automate individual responsibilities, build understanding of their behavior, then consider connecting them into a loop. Returning to incident triage, a correct diagnosis is evidence about that step. It does not by itself establish that planning, implementation, testing and rollout can all operate without intervention. Ownership and monitoring must extend along the chain.
“Precise execution does not prevent poor design.” An agent can carry out instructions accurately while the architecture or product direction remains wrong. Humans therefore retain strategy, architecture, prioritization and validation gates, including tightening guardrails and catching code drift. The intended division of work is for the team to set direction and risk parameters, then let agents execute within them.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Measure the trip to production, not the amount of generation
Token leaderboards reward consumption. Once token usage becomes a target, engineers have an incentive to increase it whether or not the additional spend produces better software. Palmie calls this Goodhart’s law in real time. Lines of code have the same problem: producing more code is an activity, while fixing a user’s problem is an outcome.
Factory’s preferred metrics examine different parts of the result:
- Signal-to-production time. Follow the journey from the incoming problem or request to deployed code.
- Human intervention count. Track how often the process requires a person to step in.
- Median time to repair. Track how quickly problems are repaired.
- Shelf life of code. Examine how long the produced code remains useful.
- Cost per PR. Relate spending to a unit of development work rather than rewarding raw token volume.
These measures bring the factory back to its purpose. Agent readiness, access controls, tooling and review systems support effective engineering regardless of who writes the code. The final instruction is deliberately blunt: “Don’t token max.” Users experience the resulting improvements in the product, so the factory should be judged by what its work changes for them.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
Related talks
- Beating RL With Reflection: GEPA and Optimize Anything
Explains a concrete approach to improving agent prompts, programs and repository skills from execution feedback, extending this talk’s discussion of correcting agents so future attempts improve.
Read the complete timestamped transcript
- 0:12
Hey, folks. I'm Davis Palmie, a member of technical staff here at Factory. Uh, before this, I co-founded Lumetric, which was an agent platform for investment teams that we took through Y Combinator and was acquired by Factory. And prior to that, I was a engineering lead for Slalom's Innovation Lab. So I've seen AI deployments both across large enterprises and more AI-native orgs. With that in mind,
- 0:42
today, I'm gonna be talking to you about the software factory and how you can move towards an autonomous system that translates input signals into production code. So first, we've moved through three distinct eras of AI engineering. First, we had tab autocomplete, next token prediction, and the engineer was still firmly in the driver's seat. Then
- 1:12
AI could generate entire files, and the software engineer became a bit of a discerning copy-paster. Um, but they still heavily had control over the output of the AI. And then you had agents that were more tightly coupled with the code. They could gather their own context, call tools, and debug their own output. Now, autocomplete used to feel magical, and now at best, it's, it's a quaint thing we remember. You wanna look at the speed at which AI
- 1:42
engineering is climbing the abstraction ladder from tokens to files to context now to entire systems. Every shift you have here is an adjustment, but each move up the abstraction ladder has made engineers more valuable, not less. And so now we're at a transformational point. Engineers are going from typing code to governing
- 2:11
agents, and eventually you'll organize those agents into a system. That's your software factory. So with this, the engineer's role shifts to keeping guardrails tight, catching code drift, and setting the high level priorities and direction. But meanwhile, the system is ingesting input signals like bug reports or user feedback and outputting production code.
- 2:43
So if you have AI systems that are producing more code than humans, there are some big questions that everyone is rightly asking. First, how can I prevent massive cost overruns? We've all seen the headlines about Uber's CTO blowing through the year's AI budget by April, or Microsoft clawing back Claude Code licenses to try and cut down on spend. At Factory, you're not locked into a single
- 3:13
model provider, and this means you unlock the entire spectrum of cost to performance with LLMs. You want the maximum, uh, cost efficiency and the minimum effectiveness necessary for each task. As our CEO likes to say, uh, "If your kid needs a algebra tutor, you can definitely find someone cheaper than Albert Einstein." Second point, how can I ensure the best access
- 3:44
to whatever model is top right now? Well, we think that having to back the right horse is a false choice, um, especially with the increasing effectiveness of open source models. We're strong proponents of being able to choose the right tool for the right job, or even now have it dynamically routed. Number three, what happens to software engineers? What does the future of software engineering look like? Well, the role is
- 4:14
obviously evolving and the line between product and engineering is blurring. Engineers will now build and maintain the system that's responsible for building your production code, and they'll set the high level direction for the software factory. And lastly, how does this gonna change the organization? Well, team boundaries are also starting to blur, and context needs to flow seamlessly across these boundaries.
- 4:44
Information can't be siloed. That was true for humans, but it's especially true for AI agents. Everyone across all your teams will have a shared responsibility for the pillars of the software factory. Now, coding was the easy part. This actually was never the bottleneck. Your context is laid out in natural language. I bet if you look at your real engineering metrics, what's the
- 5:14
median time to review a PR? How long does it take you to reproduce user bug reports or go through and debug code? Who keeps your docs current with every PR, and who keeps it current with every strategy meeting? Who maintains your test suites? Oftentimes, the testing teams and the engineering teams and the product teams are each on their own continent. If you pull your DORA metric, the time from code written to time in production, I would bet
- 5:44
dwarfs the time actually making those engineering changes. The real need you have today is cohesion across all of these surfaces. Right now, docs will go stale, dead code lingers, and there are so many layers of communications between your different teams. Again, this is hard for human engineers, but it is especially more pronounced for agents. The necessary
- 6:13
undertaking to build your factory is to make sure the agents can bridge all of these services and understand the nuance within your organization. So this brings us to the software factory, which is the system of agents that will take you from input signal to production deployment. Your agent's not just coding anymore. It's taking on customer support, product, engineering,
- 6:43
deployment ops, and many other roles across these pillars. Your agent should be triaging incidents, making plans, creating and updating docs, actually executing the code, reviewing it, testing it, rolling it out, and then monitoring that release. Now, that's obviously a ton of work, and you don't wanna jump into that all at once. For one, it's too heavy of an undertaking,
- 7:13
but for two, your teams will not have time to build trust in this system. You wanna start with something narrow and verifiable, like incident triage. Let your agent read a Sentry alert, pull the traces in other contexts, and then post in Slack. But keep it so your engineers can review this and decide what to do. And as more correct diagnoses roll in and in from the agent, your team can start to build trust and an
- 7:43
understanding of this pillar. So each step in the software factory needs careful ownership and monitoring before you even think about letting the, the pipeline loop. But once you can tie the output of this system back to the input signal, that's when you unlock organizational self-improvement.
- 8:07
Now, we feel pretty strongly that your software factory should have a few key principles. First, it should be model agnostic. Obviously, different models excel at different tasks and with very different cost efficiencies. We've found, for instance, that you can achieve the same performance on code review with GPT-5.2 as with the latest Opus models, but at half the price. And now if you look at the open
- 8:37
source alternatives, that can slash your cost down to 10 to 30X cheaper. So at Factory, again, you have the entire interplay of speed to task efficacy to cost. Not being locked into a single lab means you unlock the entire Pareto frontier of LLMs. Now, number two, your deployment model should be sovereign. You shouldn't have to compromise on your security
- 9:07
architecture or your policies to try and fit your AI deployment model in. Now, at Factory, you can choose your workspace and your deployment model. We support everything from fully managed to bring your own machine to entirely air-gapped. And number three, your software factory needs to be integrated across the SDLC. These are no longer separate, uncoupled concerns. Docs changing should inform your code
- 9:37
and vice versa. Planning goals should inform what you test for, not however many communication layers there are between strategy, product, engineering, and your testing teams. Each one of these pillars supports the software factory, and weakness in one means a weakness in the entire system. So tactically, in an AI native org, what does this look like?
- 10:08
Well, we typically see it as the fusion and the incremental automation of many sub-teams under the umbrella of one software factory. Some really common patterns we see are triaging incoming signals like Sentry alerts, having the agent make a diagnosis before the human can even answer the page, proactively testing security, scanning for secrets or vulnerabilities,
- 10:38
creating, maintaining your tests, monitoring the ones that are becoming brittle or unneeded. Again, keeping docs in sync with code, keeping docs in sync with your different meeting channels. And then the classic reviewing code, commenting on PRs, rollouts, monitoring releases. And with the agent doing all of this, the human engineer's role shifts to responding, to exercising judgment, and
- 11:08
correcting the AI in such a way that it can learn from these mistakes. Now, everything is transitive across sub-teams. An improvement to one group means an improvement to the entire system. So why should your enterprise care about a software factory? Well, the larger and the more fragmented you are, the more you feel the linear pain of managing teams. In
- 11:38
tech terms, your software factory should take this pain from O of N to O of one conservatively. You can define your guardrails, your agents, your docs, testing policies, really governance just one time, and then standardize it across the organization. And this standard should make sure that everyone is pulled to the highest bar, not the lowest common denominator
- 12:08
Now, also, if agents can move across these surfaces, then they can hold that quality bar across teams without you having to linearly scale your time to manage each one. And if agents are moving across all these services, then you have one place where context can be unified, and you can have executive visibility into how the entire SDLC is performing.
- 12:35
Now, just like model agnosticism removes lock-in, your agents should also be surface agnostic. They need to be highly available to your teams, regardless of what tools they use. Your teams should not have to compromise on their processes or their tooling to try and fit your AI provider in. At Factory, we have agents available everywhere, from remote machines to the CLI to Slack, any
- 13:05
tool your teams use. And even though humans and agents may use different services, everything should share the same underlying harness and context.
- 13:17
So the undertaking of building the factory will expose tribal knowledge, manual processes, and outdated info. Unblocking these is not only going to help your agents, but it's also gonna help your engineering teams. At Factory, we have the concept of agent readiness, and we evaluate this based on eight pillars. The first is validation. Do you have guardrails like
- 13:47
linters and formatters in place? Then your build system. Are your build commands in CI well documented? Feedback loops. Do you have unit and integration tests tight? And do you have docs like READMEs, AGENTS.md files?
- 14:06
Are your dev environments cleanly reproducible such that an agent can use them? Is your code modular with clear boundaries and rules against sprawl? Do you have observability? How long does it take you to find out why something went wrong? And are you proactively security scanning so that you don't have to worry about leaked secrets or vulnerabilities? Basically, the key takeaway here is that if your agents
- 14:36
and your human engineers are gonna walk the same roads, paving them well is doubly useful.
- 14:45
Now, you can't govern what you can't see, and that's true for agents and human engineers. You can start with this agent readiness assessment, but no matter what, your engineers must monitor agent performance and control behavior. At Factory, every action is auditable. It operates under your standards with role-based access control and least privilege, for example. Your teams that are building this software
- 15:15
factory need to first build trust and understanding. That's why you should roll this out incrementally. You wanna automate the individual pillars first before you turn this system into a loop.
- 15:31
And again, precise execution does not prevent poor design. It never has. You still heavily need humans in the loop for architecture, for strategy, and for setting direction. Engineers will not be replaced, but judgment and prioritization now beat mechanical implementation skill. Agents won't be able to handle everything. There will be guardrails to
- 16:01
tighten, code that drifts. Humans need to stand as validation gates in this factory. But just like we don't have to feed punch cards into the machine anymore, code does not have to be written by hand. Your team should be able to set the strategy and their risk parameters and then let agents execute. And knowing if your agents are executing well comes down to measuring the right things.
- 16:31
Basically, you wanna measure outcomes, not tokens. Lines of code and tokens generated are gameable metrics. These are not actually associated with outcomes. Token leaderboards at big companies just encourage profligate spending. It's Goodhart's law in real time. If you decide to measure token usage, engineers are going to optimize for token usage, and then it's no longer a good
- 17:01
metric. It's not about who spends the most on tokens. It's about who can get the highest leverage outcomes from that spend. Some metrics that we like at Factory are signal-to-production time, human intervention count, median time to repair, shelf life of code, and cost per PR. These are real, tangible outcomes.
- 17:30
So agent readiness is really a precursor to effective engineering, whether it's an agent or a human that ends up in that system. Your agents need the exact same careful governance as your engineers: tooling, access levels, systems for review. Don't just turn them loose in your code base. And don't token max. Measure what actually
- 18:00
matters. Your end users don't care about how much you're spending on AI. They care about feeling the dramatic improvements that will come from your software factory. Thank you.