AI Engineer World's Fair 2026
Harness Engineering: How to Build a Software Factory — Dru Knox, Tessl
Read the talk
Harness Engineering: How to Build a Software Factory
Dru Knox explains how cheap development checks, deeper PR review and feedback-driven improvements can turn coding agents into a system engineers build and maintain—one workflow at a time.
From a talk by Dru Knox
At a glance
Ideas worth remembering
Autonomy measures how much correction an agent needs; automation measures how much work can proceed without manual verification. Improve both while tracking product quality.
Run cheap checks during development, expensive checks on PRs, and a meta loop across logs, reviews and user feedback to improve future attempts.
Shared issues, sandboxed agent runs and PR comments make human corrections available for later analysis; reusable skills distribute what the team learns.
A recurring task such as a flaky-test hunt can become a skill and an automated workflow, reducing the separate effort required to start each run.
When agents write the product, engineers build the factory
A coding agent can finish a task correctly and still require a human to inspect every line. That distinction is the starting point for Dru Knox, head of product and design at Tessl. His approach to harness engineering assumes a team already uses coding agents, often in several sessions at once, and gets reasonably good results on easy and medium-complexity tasks. The next problem is making that success repeatable without constant intervention.
Knox uses “software factory” for an agentic system in which agents create everything shipped to users, while engineers build the system that produces it. This is a working definition, rather than a settled industry term. Engineering work shifts toward internal tools, workflows and checks that make the factory more autonomous, more automated and better at producing useful software.
Three dimensions explain the progression:
- Autonomy: How often must a person correct the code or nudge the agent toward the right approach?
- Automation: How much work can proceed without a person manually reviewing and verifying the result? An agent that regularly solves tasks in one attempt can still have low automation if nobody trusts its output yet.
- Quality: How good is the resulting product? User analytics, test quality and test coverage remain relevant even as the method of producing software changes.
The intended order is to improve autonomy, then automation, while holding quality steady. The eventual payoff, in Tessl’s view, is higher quality: additional capacity can go toward neglected bug fixes, better tests, architectural refactors and experiments. That is an expectation about what a factory enables, rather than a demonstrated quality gain in this recording. It depends on using the new capacity for those improvements, rather than spending it entirely on more features.
The same division of work can broaden participation. People outside technical roles can contribute ideas and explore them through the system, while engineers concentrate on the machinery everyone uses to ship. The factory therefore changes collaboration as well as coding speed.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Put checks where their cost makes sense
Harness engineering builds the loops that automate work and improve the factory. Knox also notes the emerging name “loop engineering.” The useful distinction is where each loop runs: during development, after a pull request appears, or across the history of many attempts.
- Inner loop—correct while building: Before the agent opens a PR, checks should be fast and cheap enough to run repeatedly. They catch problems and guide corrections without waiting for a human. This primarily improves autonomy.
- Outer loop—build confidence in the PR: Once a PR exists, more expensive and exhaustive checks can replace work a reviewer would otherwise do. Agentic QA and mutation testing to examine test-suite quality are Knox’s examples. The agent then iterates on their results.
- Meta loop—improve future attempts: Outside an individual development task, inspect agent logs, PRs, issues and user feedback. Find mistakes that reached users or required human correction, then change the inner and outer loops to catch them earlier.
Where does a correction go after the current PR is fixed? The diagram separates feedback that repairs this attempt from feedback that changes future attempts. Inner and outer checks help the agent finish the current work; the meta loop uses the accumulated evidence to improve those checks. Without that return path, the team can keep correcting the same failure.
The agent builds and runs cheap, frequent checks before opening a PR.
Cheap checks operate during development; deeper checks operate on PRs. Evidence from both, together with user feedback, guides changes to future checks.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The improvement work competes with the feature deadline
These loops are easy to name and difficult to maintain. Practices change quickly enough that keeping up can become a research job: read papers and blogs, adopt an approach, then discover that it has become an anti-pattern. A team needs time for revisiting its methods, rather than treating its first harness as finished.
A more persistent problem is scheduling. Agent failures emerge while a feature is underway, and neither their form nor the time needed to fix them is predictable. Improving the harness competes directly with shipping that feature. Skip the improvement and the agent stays stuck; invest in it and the deadline may slip. This explains why a useful correction can remain a one-off human intervention instead of becoming a reusable workflow.
Even a team that makes time may lack the evidence it needs. The failure history sits in local agent logs, on individual machines or in someone’s memory. Improvement loops require workflows that save those interactions in places other tools can inspect.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Make agent work visible before trying to optimize it
Knox organizes the implementation into three layers: a control plane, agent-ready infrastructure and improvement loops. These are his proposed categories and ordering. The control plane comes first because the rest of the system needs a readable history of what agents and humans did.
At Tessl, work begins as an issue, goes to a headless agent running in a sandbox, and becomes a PR. Engineers correct the agent through PR comments. Human involvement remains, but it now leaves a record: the original request, the proposed implementation and the corrections are available for later analysis. Moving the conversation changes what the meta loop can learn.
The control plane needs three capabilities:
- Start work: Track issues and use them to initiate agent tasks.
- Review work: Give engineers a shared review interface; GitHub PR review is the straightforward example.
- Distribute improvements: Store and share workflows or playbooks through a skills registry or a shared GitHub repository, so a correction can help agents beyond the original session.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Give the agent the environment the work actually requires
The second layer is “agent IT.” Moving work from a developer’s machine into an agent environment exposes dependencies that local setup has concealed. A CLI may exist without usable credentials. An internal service may lack appropriate API access. The agent may need to click through the product or execute the code it changes. Knox’s warning is blunt: “it’s worse than you think.”
The work spans documented company workflows, access to internal services, production logs and a suitable execution environment. Each brings its own decisions: service access requires governance, while production-log access raises compliance questions. These requirements are company-specific and become visible as agents try to produce code without someone stepping in.
Once that foundation works, improvement loops become the ongoing investment. Scheduled repository sweeps can find problems; playbooks can encode recurring development practices; repeated tasks can become automations. The important return path is from observed output to a changed workflow. If an agent makes a mistake while adding a CLI feature, update the CLI-feature playbook so later attempts receive the correction before they repeat it.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Turn a recurring flaky-test hunt into an automated workflow
Tessl’s product approach is to make this transition through small, repeatable improvements. Knox’s image for the alternative is “one massive snake eating the moose all at once.” Tessl aims to carry the burden of keeping harness practices current while letting teams choose components for their own stack. The techniques described earlier do not require Tessl.
For the control plane, Tessl provides a skills registry that publishes and versions workflows, with security and quality reviews and controls over who can publish or update them. At presentation time, its issue-tracker integration focused on Linear tickets triggering GitHub workflows. Code-review tools support both general standards and targeted checks; Knox describes a registry workflow being assessed for security, quality and its effect on agent output.
Tessl Agent addresses the scheduling problem by inspecting PRs and issues for repeated work. The concrete example is a weekly hunt for flaky tests. Instead of requiring someone to notice the repetition and separately make time to automate it, the agent identifies the pattern, turns the workflow into a skill and sets it up in a GitHub Action. The observable change in this example is how the task starts: a recurring manual activity becomes an automated workflow. It is an illustrative workflow, without a reported before-and-after reliability result.
How does that discovered task become an agent run that can produce a reviewable change? The diagram follows the workflow from historical evidence to a skill, then to execution and a PR. The skill carries the procedure; the automation supplies the recurring execution; the sandbox and GitHub access let the agent perform the work and discuss its result.
Tessl Launch runs a skill as an automated workflow with a selected coding agent. Knox names Codex, Claude Code and Gemini as options for coding tasks, alongside Tessl Agent. Execution happens in a sandbox with appropriate permissions, can be long-running, and uses GitHub tokens to open PRs and respond to comments. This connects the reusable procedure to the same shared review process that made agent work inspectable in the first place.
Prebuilt maintenance workflows cover architecture quality, code duplication, test-suite quality and security vulnerabilities. Their role is to lower the effort needed to begin regular scans and receive proposed improvements. The adoption unit stays small: find an improvement, codify one workflow, distribute it, then run it repeatedly.
Tessl Agent finds the recurring weekly flaky-test hunt.
The flaky-test example illustrates discovering a recurring task and packaging it as a skill; Launch supplies agent execution and PR interaction.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Measure less intervention alongside product quality
The closing metrics turn the factory idea into something a team can track:
- Fewer manual takeovers: Agents need less rescue to complete the work.
- Fewer human PR comments: Less corrective review is needed before accepting changes.
- More PRs initiated without human input: The system starts more work itself, indicating greater automation.
- Steady, then improving quality: The product must continue to meet its quality standard as human involvement falls, with improvement as the subsequent goal.
Read these measures together. Fewer comments can indicate progress when checks provide the confidence reviewers previously supplied; quality measures determine whether that confidence is justified. The practical path is one workflow at a time, with the saved history of each attempt helping improve the next.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
Related talks
- Beating RL With Reflection: GEPA and Optimize Anything
Explains a complementary mechanism for the meta loop: use execution traces and evaluator feedback to improve prompts, agent programs and repository skills.
Read the complete timestamped transcript
- 0:12
Quick introduction, my name is Dru Knox. I'm the head of product and design at Tessl. Tessl is an agent enablement platform. We help you go from scaling skills to building your own software factory. So today, I'm gonna talk about a few things. The core is gonna focus on harness engineering, which we see as sort of the new discipline of agentic development, and how you can use it to ladder up to a software factory.
- 0:42
So before we get started, a few prerequisites. This is a bit of an advanced talk in the sense that, like, if you aren't already using coding agents, what I'm about to walk through is probably not where you wanna get started. So I'm sort of assuming that if you're working with coding agents, you're used to having a couple sessions at a time. Uh, you've gotten to a place where coding agents will, for, like, easy tasks and maybe medium complexity tasks,
- 1:12
frequently do the right thing. This-- you're sorta at the sweet spot for starting to think about harness engineering, starting to think about a software factory. What I'm gonna say here is not exclusive to that, but that's where you'll get your best results.
- 1:27
So I'm gonna walk through a few things. First, I wanna define software factory and talk a little bit about why you might want to build towards one. Second, we're gonna then focus a lot on harness engineering, which is kind of the core practice or skill that you'll use to reach a software factory. We'll talk about what are gonna be the components you work on while you're doing harness engineering, and finally, I'll walk through how the Tessl product, uh, briefly can help you along that journey. But most of the
- 1:57
topics I'm gonna talk through here today work with any stack. They're mostly techniques. Tessl makes them easy, but they're not exclusive to us. Okay. Part one, software factories. What are they? There's no strict definition. I'm sure you're all probably as frustrated as I am with how quickly terminology changes and migrates in the, in the industry today. But roughly speaking, a software factory is any agentic system
- 2:27
where all of the end product of what you're building and shipping to users is created by agents, and your software engineering teams effectively fo-focus on building the factory. So they work on making the factory more autonomous, more automated, and raising the quality of what it produces. I'll go into each of those terms in just a second. That's kind of the core idea of a factory, though, is everyone on your team effectively becomes internal tool builders, and then you're building a system that produces your
- 2:56
product.
- 3:00
So I mentioned three key metrics. These are sorta like the pillars you need to think about when building towards a software factory, and they sort of go in this order. So autonomy is effectively how much human intervention do you need to get to the right answer. So how many times you need to correct the code or provide a nudge on how the agent should be building. Automation-- Uh, oh, sorry, I've gotten them flipped here. Uh, autonomy is the first one. It's how much do you have to,
- 3:30
um, do humans have to correct. Automation is how much can you let them run, right? How much do you need to review or build trust in the solution before you accept it? They sound quite similar, but there is a difference. You can have high autonomy in the sense that agents frequently are one-shotting the problems you're giving them, but low automation because you don't trust it yet, and so you're reviewing all of the code directly. You are manually verifying everything. So they are distinct things
- 3:59
that obviously have a connection, right? You have to build up autonomy first before you can move to automation. The final quality, probably the most self-explanatory, how good is the product that you're actually producing for your users? These metrics are like your usual user analytics, your, uh, test quality, your test coverage, et cetera. As you work towards a factory, you're gonna go improving au-autonomy, then improving automation,
- 4:29
all while keeping quality constant, and then the sort of payoff at the end, why would you want a software factory, is that ultimately we think you can raise quality once you have a factory. So talk about this a bit. I think most folks look at factories as a way to ship faster. They'll say, like, "Oh yeah, you'll get a little bit of slop, but my God, you'll, you'll launch so much more features, like the cost to return, the ROI will be worth it." I think that's a very short-term phenomenon,
- 4:59
right, as agents are sort of coming online and getting better. In reality, yes, you will move to higher velocity with a factory, but every team has backlogs of bug fixes or improvements that they wish they could be making, that they just don't have time to do, or explorations that they wish they were, they were doing. With a software factory, the idea of a backlog kind of goes away, and so you actually just have much more capacity to work on test quality
- 5:29
improvements, architectural refactors. So at Tessl, we believe that yes, at first as you transition, you will wanna keep quality constant or maybe like a very small dip, but that the ultimate payoff is you should see better quality, uh, in the code that you're producing. Also, it's great for your overall team collaboration style. So once you've moved to a software factory, it's much easier for people who aren't in technical roles to contribute ideas, so it's easier to explore more and to have more
- 5:59
perspectives coming in and contributing while your engineering team is much more focused on- Building the underlying system that everyone is using to push features out to your users. So I think ultimately it leads to a more dynamic and inclusive, uh, development experience.
- 6:16
So how do you get to a software factory? Harness engineering, uh, as a sign of how fast terms evolve, it's also be- started to be called loop engineering over the last couple weeks. Um, this is really the core discipline that you're gonna take to get towards a software factory. At its core, harness engineering is... You know, I mentioned the software factory is building your product. You are building the loops that automate and improve the quality of your
- 6:46
factory. So what do I mean by loops? What are the loops that you might be working on? There are three core concepts, each at-- maps to a certain phase of development. So you have your inner loop, which is as the coding agent is working on a PR before it has put it up, right? So these are very fast iterative loops. You want them to be cheap. You're expecting the agent to run them all the time. Improvements here will drive better autonomy.
- 7:16
The-- They help catch things and correct the agent without your, uh, coming in and human intervention.
- 7:23
Once the PR is up, you move to the outer loop. So this is where you're gonna put more expensive, exhaustive checks that don't make sense to run over and over again as the agent is iterating on features. But they do encode things that otherwise a human would have to verify to build confidence in the system. And so this is because they're more expensive, but they are taking away human review time, they can be worth it, right? So this is where you might put something like agentic QA or where you might, say, like,
- 7:54
run mutation testing to check the quality of our test suite. This is for expensive things that you wanna run once when the PR goes up and then have the agent iterate on, on the results.
- 8:06
Finally, you have the meta loop. This is where you'll drive quality. So the meta loop sits outside of the development process. It observes your coding agent logs. It looks at PRs, your issue tracker, user feedback, and it's basically finding mistakes that are making their way through the pipeline to your users or that you had to correct as a human to stop that from happening. And it feeds back into the inner and the
- 8:35
outer loop to make sure that that mistake is not made again. So the outer loop is what really kind of... Once you have the infrastructure for each of these loops, the outer loop-- or sorry, the meta loop is where you're gonna be actually driving from a certain percentage of, like, how AI native am I? You're gonna be driving it up by investing in your meta loop.
- 8:59
Before we get into, like, what are the components, what do you do to actually accomplish this, I think full disclosure upfront, harness engineering, at least, uh, raw unassisted harness engineering is hard for a few specific reasons that basically come down to human psychology. So the first, hopefully this one will go away over time, but it's a new discipline, and it's changing very fast. And so a lot of teams that get into harness engineering, you
- 9:29
basically become an AI researcher just to try and keep up. You're reading papers, you're watching blogs. You're coming up with best practices that go out of date the next week and then two weeks after that. And actually, at some point, they become anti-patterns, and you don't wanna do them anymore. Uh, and so there's just, like, a lot of time and space you'll devote to keeping up with how to do harness engineering. And so if you wanna get into this discipline, you need to have a sort of upfront approach to how you're gonna solve this problem. How are you gonna make
- 9:59
time and space to keep up with the knowledge that's changing every week? The second, and this one is more durable in my opinion, harness engineering is fundamentally unplanned work, so you will start shipping a feature. You have no way to anticipate will agents fail, how will they fail, how long is it gonna take me to fix it? And all of that work that you're gonna do to try and make the agent better competes with shipping, right? So it will slow you down from getting the features out, and we all know that is
- 10:29
just a perpetually hard trade-off to make. You either don't do it, and you get stuck in the local maxima of the agent never improves, or you do it, and you miss your deadline for the feature, and you get in trouble, and so no one ever does it. The final piece is that once you've solved these re-- initial problems, things like make time and space for harness engineering, you'll find that a lot of the information you want access to is not available. It's hidden away in local coding agent logs, or it is on someone's
- 10:58
machine somewhere. It's in someone's head. A lot of the signals that you wanna see to make agents better, you need to do some work to move your workflows onto surfaces where everything is saved and publicly available for later optimization loops.
- 11:16
Okay. With all that aside, then what are the pieces that you should be working on as you're doing harness engineering?
- 11:23
I've broken it into three layers. These aren't by any means standard terms, though if you wanna help me make them so, I would be g-- forever grateful. The first layer, and these are sort of like in the order you need to do them as far as Tessl is concerned, is you have to get your control plane right. And so I mentioned how most of the data that you wanna have access to to make agents better and move towards a software factory is not legible. If you move your workflows into, for example,
- 11:53
at Tessl, we have all issue-- all work starts as an issue on the issue tracker. It gets sent to a headless agent running in a sandbox. The agent puts up a PR, and then engineers engage with comments on the PR. So it still has manual correction. It still has manual interface, but now all of your touch points with the agent Are legible. So we can-- you can then point tools at them to pull that information down and make things better. Typically, the three things that
- 12:22
everybody comes to when they're trying to set up a control plane is you have some way of tracking issues and kicking off agent work from that. You're gonna have some way to review, GitHub PR review is probably, like, the easiest. Uh, and then you're gonna have some way to, uh, standardize and distribute workflows, playbooks, things that make the agents better. Typically, that ends up being something like a skills registry or a shared GitHub repo where you push skills that improve agent quality.
- 12:51
Next is a massive grab bag that I call agent IT, which is, uh, if anyone's ever tried to move, like, a project that had a lot of local config into something that someone else can set up, and you realize with horror all of the dependencies that you thought were cleanly isolated and are not, you basically are gonna have to go through the same thing with agents. There's gonna be all these things that you thought were easily accessible by an agent, right? Like CLI, API access, giving a way for the agent to click through your product to build
- 13:21
it. I, I promise it's worse than you think. Like, there's no one who has come in thinking like, "Yeah, yeah, it's not gonna be that hard for us," has left saying that at the end. And so there's just, like, a lot of pieces you'll work, work through, and this is probably the most unplannable and most, like, spend a week, spend two weeks, and just get it done. But it will come through things like writing, writing down in your company brain how you get work done, what are your workflows, et cetera.
- 13:52
You're gonna have to give access to all your internal services, which will require governance and, like, API access questions. You're gonna wanna give a way to access production logs, uh, and that's gonna have its whole own question of, like, what's your compliance stance, things like that. Uh, and then you'll need an environment where the agent can execute the code that it is running. There's infinite more, but they are generally pretty company specific, and you'll just work through as you try to get agents to put code up without you intervening. You'll quickly learn what are the problems.
- 14:23
Finally, this is where you'll end up spending most of your time. These are the improvement loops. So these are the things I mentioned before, the, like, inner, outer meta loop. These are components that tend to go into your meta loop. So repo maintenance, you're gonna want things that sweep your code base every day, every night, every week, something like that, to look for problems. You're gonna want playbooks for common development practices, so how to add a feature to your CLI. You're gonna wanna find a way to identify repeated
- 14:53
tasks and automate them, uh, and ideally as quickly and easily as possible, because otherwise people won't do it. And then you're gonna want something that looks at the output quality just on a consistent basis and brings back improvements and learnings to the rest of your code base. So say things like, "Oh, I saw the agent made this mistake. Let's update a skill on add-- like, the playbook you have for adding features to the CLI. Let's update that so it doesn't make this mistake again in the future."
- 15:21
So those three control planes, Tessl sort of exists to try and help you solve them. As I mentioned, everything I've said before here are just techniques. You could go do them all yourself if you wanted. There's plenty of tools that do that. So I'm gonna talk a little bit through how Tessl will help you if you wanna use us. To start, why would you wanna use us? So we focus really hard on making it relatively iterative and sustainable to get to the cutting edge, so we try to make it, like, a bunch of
- 15:51
small lifts rather than one massive snake eating the moose all at once. We're batteries included, so we solve the knowledge gap by saying, "We will keep up, and you will have an agentic experience that basically knows the best practices of harness engineering, is kept up to date on your behalf." We have a real commitment to being modular and open, so we don't think every component of a factory is gonna be best in class with a single company. We wanna make sure you can pick and choose what is the most important
- 16:21
piece for your stack. And then, like I said, we focus on making it easy. So we have automated loops for you that you can install and just react to the changes that are proposed, and we really try to make sure it's just incremental steps so that you, six months later, are like, "Oh, wow, we're, like, forty percent there to a software factory. We never had to stop and delay shipping or anything like that."
- 16:43
So the first place that Tessl helps is in setting up your control plane. We have a skills registry where you can publish and version the workflows and the automations that you're working on with built-in governance, things like security reviews, quality reviews, controls for who can publish and update what. We also have easy connectors for issue trackers to GitHub. Right now we're focused on Linear to GitHub, so you can basically file a Linear ticket and get it to run a GitHub workflow, uh, with more connectors
- 17:13
coming soon. Uh, and then we also have a suite of tools for code review. This will make it easy to set up your standards for agentic code review as well as more targeted checks. See here an example of something that ha-- like a workflow that has been posted to the registry being scanned for security, quality, uh, and how much it actually improves the agent's output.
- 17:38
After this, I'm gonna play a video while we talk, see how this goes. So, uh, we have an agentic experience called Tessl Agent, which will help you with creating and then maintaining your improvement loops over time. So the first thing that we do is we will help you with that change management of, like, how do I actually make the time to find repeated tasks and set them up in automation? So the agent will mine through PRs, issues, things like that,
- 18:08
and find, like, every week you do a hunt for flaky tests. Why don't we just set that up as a skill for the workflow, and then we'll put it in automation on a GitHub action? So we help you actually make time to do the automations, just, like, one workflow at a time. The second is that we come with a lot of out-of-the-box maintenance tasks. So I mentioned you wanna have weekly scans for things like architecture quality, code duplication, uh, how good are your test suites, do you have any security vulnerabilities? So
- 18:38
if you use Tessl, you can just sort of like one-click install a bunch of those things and just get weekly improvements to your code base without any other effort. And then finally, we offer a way to turn any skill into an automated workflow. So it's through a command called Tessl Launch, where you effectively take a sk- a skill that codifies a workflow. You can then pick a s- particular coding agent that you want to handle it. Tessl Agent can be one of them, but for most coding tasks, you probably wanna use something like Codex, Claude Code, Gemini. We have access to
- 19:08
all of them, and we will run them in a sandbox that has the appropriate permissions. It can be long-running. Uh, we'll have GitHub tokens so that it can put up a PR and respond to comments as you leave them. Uh, so it just makes it really easy to automate workflows like that. And so with these things together, we just make it really easy to get started and just one at a time find an improvement, take a workflow, turn it into a skill, ship it to everyone, and then get your improvement loops going.
- 19:39
The what... what you're gonna wanna focus on as you're doing this work is driving down manual takeovers, driving down m- human PR comments. So these are like metrics that you can track to see, like, how far am I towards software factory. And then over time, you wanna see more PRs initiated without human input. That's a sign that you're automating more. And then, of course, you want to first hold quality constant and then s- aim to drive it up after that. That's it. If you wanna learn more, Tessl booth is just that way. Please
- 20:09
give a scan, come over, see us. We'd love to chat. Um, thank you all for your time.