AI Engineer World's Fair 2026
No, That's Not a Software Factory — Ryan Cooke, WorkOS
Read the talk
No, That's Not a Software Factory
WorkOS’s first sandbox could produce pull requests, but Ryan Cooke found little difference from engineers using coding agents locally. TARS and Horizon extend automation into planning, review, project coordination and context gathering—the work that turns code into a product.
From a talk by Ryan Cooke
At a glance
Ideas worth remembering
A sandbox can generate code without improving product delivery. WorkOS’s initial experience pushed it toward automating planning and coordination alongside implementation.
Ticket dependencies and completion webhooks let TARS continue a project without a fresh prompt at every handoff; requested reevaluation helps expose missing work.
Draft Hilltop documents reduce setup work and give engineers something to refine, but scope control remains an important human contribution.
An internal MCP gateway becomes more useful when it explains where information lives and when to query it. WorkOS reuses that guidance for coding, data analysis and customer analysis.
Evaluate customer delivery alongside defect rate, recovery time and voluntary adoption. Company-wide memory and session-driven improvement are future directions; authorization remains unresolved.
A pull request is an output. What did it change?
The familiar software-factory recipe is straightforward: put a repository in a sandbox, add a coding agent, give it a prompt, and merge the resulting pull request. Ryan Cooke, an engineer at WorkOS, opens with that recipe because it leaves an important question unanswered: does the organization actually deliver more useful software?
Counts of pull requests, percentages attributed to AI, and quantities of generated code describe production activity. They do not establish that the activity advances a customer outcome. More pull requests can make an automation system look successful while leaving the harder question—whether features reach customers sooner—unresolved. That distinction matters when building the factory itself consumes engineering time and potentially a dedicated team.
WorkOS’s ambition is to give each engineer something like a small engineering team. The consequence should be observable in delivery: more features shipped, with complex work completed faster. Cooke’s account offers qualitative experience and proposed success measures rather than numerical delivery or reliability results; it explains the design and its intended benefits without establishing their magnitude.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The sandbox worked; the surrounding process needed work
WorkOS began with a sandbox built on Cloudflare, a model router, and prompts. The disappointing result was how ordinary it felt. Cooke found the outcomes pretty indistinguishable from engineers driving coding agents on their laptops. Moving execution into a sandbox had not, by itself, removed enough of the surrounding work to change product delivery.
The next design question was how to put WorkOS’s engineering processes inside the factory. Engineers do more than write code: they organize projects and carry work through the steps needed to produce a product. Automating those steps required the system to participate in the tools where that work already happened.
That led to two parts with different responsibilities:
- TARS: The interface to the coding agent lives in Slack, Linear and GitHub. It also subscribes to webhooks so it can track project progress alongside its code work.
- Horizon: The infrastructure orchestration layer sits in front of an internal MCP gateway. It provides the execution infrastructure, while the gateway connects the agent to organizational tools and context.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Webhooks turn ticket completion into the next action
A dependency between Linear tickets gives the factory a plan it can follow. One ticket blocks another; completing the first changes which work can proceed. Because TARS receives ticket-completion webhooks, it can pick up the next ticket automatically. Planning therefore supplies both the units of work and their execution order, reducing the need for a person to issue a fresh prompt at every handoff.
The plan can also change as implementation reveals gaps. Between completed tickets, an engineer can ask TARS to reevaluate the Linear project and identify missing work. This is a separate operation from advancing to the next ticket: advancement follows the existing plan, while reevaluation checks whether the plan still covers the project’s goal.
What connects a finished ticket to continued execution, and where can the plan be reconsidered? The diagram separates those paths. The completion event drives the next handoff; an intervening request can send TARS back to the project plan to look for gaps. The useful autonomy comes from reacting to project state, rather than treating each coding session as an isolated task.
Tickets describe smaller units of work and their dependencies.
Ticket completion can trigger the next unit of work. A requested reevaluation checks for missing tickets as implementation exposes gaps.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
From a Slack brief to a reviewed Hilltop
At WorkOS, engineers carry many product responsibilities; Cooke says teams do not have product managers. Their planning ritual is a Hilltop document, a product requirements document that gathers the project’s purpose, customer needs, competitive analysis, early designs and major milestones. Consolidating those inputs gives an agent a coherent resource from which to break the project into units of work.
The PM agent automates the first pass through that ritual. A brief specification becomes a draft Hilltop; the agent adds context, reads human reviews, and breaks implementation into tickets. Humans can comment on generated tickets and refine the work throughout the process. Approval is also an event the factory understands: once the Hilltop has been reviewed and approved through Linear, the system can begin the project automatically.
The concrete example is a new API for Vaults, a product Cooke leads. He starts the project with a Slack command and a few sentences describing the goal. TARS creates the Linear project, a draft Notion document, and first passes at decision logs and open questions. The observable change is from a short request to an organized set of project resources that the lead engineer can immediately edit.
Cooke then adds specifications and fills in poorly defined areas using his knowledge of the project. The example is still awaiting its Hilltop review: completing that review ticket will let TARS detect the change and move to the next implementation stage. The review is therefore doing two jobs. It improves the product definition, and its recorded completion tells the automation when work may advance.
The benefit is partly the familiar “blank page problem.” Cooke reports that teams are doing more project-brief and Hilltop work because the agent seeds the information. The drafts can grossly overestimate a project’s scope, requiring engineers to cut substantial portions. The tradeoff is deliberate: reviewing and trimming an existing draft can be easier than creating every project primitive from scratch. Work then continues through Slack and Linear toward GitHub pull requests.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Extend the process to triage, then improve its infrastructure
The same event-driven approach extends beyond planned features:
- Bug requests: TARS listens through webhooks to requests arriving in Slack, takes a first pass at triage, and can implement fixes and open pull requests.
- Customer support: Shared customer Slack channels provide support requests. TARS can inspect the code to help identify where a customer may be encountering a product problem.
WorkOS is also using TARS to build its own sandbox infrastructure. The motivation for moving away from some sandbox services is deeper control over session information and the ability to move workloads among parts of its infrastructure. Owning that layer becomes relevant when the factory needs to inspect how work happens, rather than simply launch an agent and collect its output.
The next planned component is a company-wide memory layer: continually useful context about what people are working on, their teams, the products they own, and how WorkOS does its work. The intended consumers include both the factory and other AI tools. This is a proposed expansion of shared organizational context, with no completed memory implementation described in the talk.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Measure delivery, reliability and the choice to use it
The closing measurement discussion returns to customer impact. Shipping faster is valuable only if the products remain dependable, and a tool that engineers voluntarily choose offers a different signal from one that merely generates activity. Cooke identifies three areas to watch:
- Delivery and customer impact: Does the factory help useful work reach customers?
- Reliability: Track defect rate and time to recovery to watch for instability introduced by the new workflow.
- Voluntary adoption: Look for engineers choosing TARS and cloud sandboxes over their local harnesses. Cooke treats this as an anecdotal indication that the system helps them work.
Infrastructure visibility is meant to feed a further improvement loop. Agents could examine sessions to find gaps in how engineers use the factory, identify mistakes that call for a new skill, and recognize skills that have become obsolete as the code changes. A skill written six months earlier may no longer fit the current code. Better code generation and better engineering practice are goals of this loop, rather than demonstrated self-improvement results.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Read the complete timestamped transcript
- 0:12
Well, thank you for joining
- 0:14
the final day of the conference here. I
- 0:15
know it's been a long week.
- 0:18
But I'm excited to talk about a topic
- 0:21
here that is very near and dear to me.
- 0:24
Software factories.
- 0:26
Um
- 0:27
You this is probably not a new concept
- 0:30
to any of you. I'm guessing if you're
- 0:31
all here, you're very familiar with
- 0:33
software factories. You know, Ramp
- 0:35
blogged about their inspects
- 0:38
several months ago. I think maybe
- 0:40
December last year. I forget exactly
- 0:41
when. It's been a little while. And that
- 0:43
really kind of like excited the entire
- 0:45
industry of thinking like, oh, we can
- 0:47
actually build software very differently
- 0:49
now that AI is extremely capable of
- 0:52
writing code. Um a lot of other
- 0:55
companies have followed suit. Um and
- 0:58
there's kind of like a standard set up
- 1:00
here. Uh you have a sandbox that you put
- 1:02
your code into. Um you put a an AI agent
- 1:06
in there. Maybe a Claude code section or
- 1:09
an open code.
- 1:11
Um and with a prompt, it generates a PR
- 1:13
and you merge it.
- 1:15
Um
- 1:15
You might be building this already
- 1:17
[clears throat] yourself.
- 1:18
We have taken a slightly different
- 1:21
approach at WorkOS.
- 1:23
Um and the motivation for what we are
- 1:26
building is that we see a lot of the
- 1:30
kind of success metrics um really
- 1:33
focusing on output, right? Like you'll
- 1:35
see posts on X and blogs
- 1:38
talking about the percentage of PRs or
- 1:41
the number of pull requests or how much
- 1:43
code uh is being generated by AI
- 1:47
moving into production. Um and this is
- 1:49
just really output metrics that
- 1:53
um
- 1:53
you know, this can actually disguise how
- 1:56
well these systems are working.
- 1:59
Percentage of PRs or number of PRs
- 2:02
may be lying in overall increase in PRs.
- 2:04
It's very hard to distinguish
- 2:06
whether the output is actually driving
- 2:09
outcomes.
- 2:11
And so that's kind of where we focus our
- 2:14
efforts when we're thinking about a
- 2:15
software factory. It's like if we're
- 2:17
going to go put engineering time into
- 2:20
building this level of automations and
- 2:22
there's a
- 2:23
bit of work that needs to go into this.
- 2:25
Full-time teams dedicated to this.
- 2:27
How do we want to think about the
- 2:30
success of this and whether this is
- 2:32
actually driving value for our
- 2:33
organization, our engineering
- 2:35
organization?
- 2:36
And so we measure instead of just code
- 2:40
output, we're looking at outcome
- 2:42
metrics.
- 2:44
And that is really about like are we
- 2:46
accelerating our ability to deliver
- 2:48
features?
- 2:50
We think that the dream of the software
- 2:51
factory is that it gives each of our
- 2:54
engineers a small engineering team for
- 2:56
themselves. And so the consequence of
- 2:58
that really needs to be that we're
- 3:00
building more and we're shipping more
- 3:02
and that as we embark on complex
- 3:04
features, those can get built a lot more
- 3:07
a lot more quickly.
- 3:11
So what does this actually look like?
- 3:12
Well,
- 3:14
we kind of started with the same thing
- 3:16
you've seen in other other
- 3:19
factories.
- 3:21
We start with a sandbox we built on top
- 3:22
of Cloudflare. We put an open code model
- 3:26
router in there. We're feeding it
- 3:27
prompts.
- 3:29
Um,
- 3:29
and we kind of very quickly kind of ran
- 3:32
into this situation where this wasn't
- 3:34
driving more outcomes. This was not an
- 3:37
incremental increase or an exponential
- 3:39
increase over engineers just driving
- 3:41
cloud code on their laptops. It was
- 3:43
actually pretty indistinguishable for
- 3:45
us.
- 3:47
And so we really kind of started to
- 3:48
think about, well,
- 3:49
how can we embed our engineering
- 3:52
processes into the factory itself? It's
- 3:56
not sufficient that the factory is
- 3:58
producing code. We actually want it to
- 4:00
take a lot of the other work that our
- 4:02
engineers do to produce products and
- 4:05
automate that as well.
- 4:08
So, we kind of split this into two
- 4:09
parts.
- 4:10
One is we have the system we call TARS.
- 4:13
This is a way for users to interact with
- 4:18
the coding agent. And it is embedded
- 4:21
into the tools that we use. So, not just
- 4:24
Slack, but also Linear and GitHub.
- 4:27
We subscribe to webhooks through TARS so
- 4:29
that TARS can actually track the
- 4:31
progress of projects
- 4:34
in addition to generating the code
- 4:36
outputs.
- 4:38
And then we built a separate system
- 4:39
called Horizon.
- 4:42
And this is the infrastructure
- 4:43
orchestration layer. This is a lot more
- 4:45
similar to
- 4:47
Inspects and Minions and and some of the
- 4:49
other systems that other companies have
- 4:51
built.
- 4:52
But, it sits in front of an MCP gateway.
- 4:56
And I'm going to talk a little bit more
- 4:57
about why that MCP gateway has been
- 5:01
really transformational for us.
- 5:07
So, by
- 5:08
feeding our webhooks, activities that
- 5:12
are happening in other uh source code
- 5:14
systems and project tracking systems,
- 5:17
we're starting to get to that level of
- 5:18
autonomy where
- 5:20
our
- 5:21
factory is performing product
- 5:23
engineering work, not just code work.
- 5:25
So, what's an example of this?
- 5:27
In Linear, we can define tickets that
- 5:30
have dependencies. So, one ticket blocks
- 5:32
another ticket. Pretty common for, you
- 5:34
know, breaking down uh large units of
- 5:36
work into smaller ones.
- 5:38
Because
- 5:40
uh TARS is getting webhooks on ticket
- 5:42
completion, it can automatically pick up
- 5:44
the next ticket in a cycle. And so, we
- 5:47
can do through our planning process
- 5:50
construct like a map of how this plan we
- 5:53
think this plan is going to get executed
- 5:55
and TARS can start to execute on that
- 5:57
autonomously. The other thing that we
- 5:59
have is in between those steps when a
- 6:01
ticket is completed, we can ask TARS
- 6:04
"Can you reevaluate the linear project
- 6:06
and let me know if it's missing now
- 6:07
tickets?" Because as you do work, you're
- 6:10
learning about where the gaps in your
- 6:12
plan. And we want to continuously keep
- 6:16
our our plan fresh by using the agent
- 6:19
itself. Because the agent is determining
- 6:22
that there may be missing pieces to what
- 6:25
we initially planned for the project and
- 6:27
the goal that we have in mind.
- 6:31
So, we run this product engineering
- 6:33
culture at WorkOS. This is where
- 6:35
engineers are responsible for a lot of
- 6:37
the product functions. We don't have
- 6:39
product managers on teams today.
- 6:41
Um and one of the rituals as part of
- 6:45
this product engineering process is we
- 6:46
create a hilltop document. It's a PRD
- 6:49
and it is intended to
- 6:52
both define what is the purpose of this
- 6:54
project, but also incorporate
- 6:58
um what are customers talking about?
- 7:00
Like where are we seeing the need for
- 7:01
this unit of work? Uh it looks at
- 7:04
competitive analysis. Like are there
- 7:06
similar products out in the market today
- 7:08
that we can draw inspiration from? We
- 7:10
start to bring design early design
- 7:12
screens into this.
- 7:13
Um and we outline like what are the
- 7:15
major milestones. And so, by kind of
- 7:18
consolidating and canonicalizing this
- 7:21
information into a document, this is
- 7:23
something our product engineers have
- 7:24
been doing uh throughout the entire uh
- 7:27
company.
- 7:28
Um
- 7:29
now we can ask give this resource to an
- 7:31
agent and an agent can break this into
- 7:33
units of work.
- 7:34
Uh and so, this has been a really uh
- 7:37
powerful way for us to take existing
- 7:39
processes
- 7:40
and encode them into the factory
- 7:42
themselves. Like I mentioned that it's
- 7:45
listening for web hooks. I'll show an
- 7:46
example of what this looks like, but it
- 7:48
it can see that the hilltop document
- 7:51
through linear tickets has been reviewed
- 7:54
and approved and by virtue of it being
- 7:57
approved can start work on this project
- 7:59
automatically. So we don't need a human
- 8:01
to be shepherding this agent through
- 8:03
every single step of the life cycle.
- 8:09
So we created an agent specifically for
- 8:12
this part of the process. We call it the
- 8:13
PM.
- 8:14
And it is doing the first draft of the
- 8:17
hilltop based on the brief specification
- 8:20
that we give it.
- 8:21
It's adding context to that in addition
- 8:25
to reading the human reviews and then it
- 8:27
picks up the implementation and breaking
- 8:29
that into tickets.
- 8:31
And there's an opportunity for humans to
- 8:33
stay in the loop for each part of this
- 8:35
process. I mean it's very common for a
- 8:38
human to intervene or comment on a
- 8:39
ticket that's generated by AI, give it
- 8:41
further guidance or refinements.
- 8:44
Um but it's really about that like cold
- 8:46
start problem where we can kind of get
- 8:49
over the hump of creating all these
- 8:50
different resources. Okay, so what does
- 8:52
this actually look like?
- 8:55
So here's a project I kicked off last
- 8:57
month two months ago.
- 8:59
Um we're adding a new API to products I
- 9:01
lead called Vaults.
- 9:03
And with a command here in
- 9:07
um
- 9:08
in Slack
- 9:09
I can kick off this project with a
- 9:11
simple description, usually a few
- 9:12
sentences of what I'm trying to
- 9:14
accomplish.
- 9:16
Tarsin goes in and creates all these
- 9:18
resources for me. I don't need to go and
- 9:20
create the project in linear. I don't
- 9:21
need to create a draft of the notion
- 9:23
document.
- 9:25
Um
- 9:25
all of the decision logs and the open
- 9:28
questions.
- 9:29
Um Tars is creating a first pass at
- 9:31
that. And so this then gives me an easy
- 9:34
framework for me to step into as the
- 9:37
lead product engineer and start giving
- 9:39
more specification, rounding out areas
- 9:42
that aren't well defined with knowledge
- 9:45
that I had know about what we're trying
- 9:46
to accomplish with the project, and then
- 9:49
hand it back to Tars uh for execution.
- 9:53
So, here's an example of uh the project.
- 9:57
Um and it's like I said, it sets up
- 9:58
these first
- 10:00
milestones around our product
- 10:02
engineering process. So, we haven't done
- 10:05
the hilltop review yet here. When that's
- 10:07
done, we'll uh
- 10:09
we'll we'll mark this ticket as complete
- 10:12
and Tars will pick that up and move on
- 10:14
to the next stage of project
- 10:15
implementation.
- 10:17
So,
- 10:18
what we found is like a lot of
- 10:21
teams are now doing more project brief
- 10:24
and hilltop work because our agent can
- 10:27
seed all that information. It's that
- 10:29
blank page problem with writing. If you
- 10:32
can come into a document that already
- 10:33
has information, and you know, these
- 10:35
agents aren't perfect. There are times
- 10:37
where it grossly overestimates what
- 10:40
we're trying to accomplish with this
- 10:42
project, and we have to cut out a lot of
- 10:43
the scope that it comes up with. But,
- 10:45
that's fine. That's a lot simpler for an
- 10:47
engineer to add input into into rather
- 10:51
than them spending time to set up all
- 10:52
these primitives themselves.
- 10:54
And then we can continue to drive this
- 10:56
work, you know, through Slack, through
- 10:58
Linear,
- 10:59
um and ultimately uh results in PRs in
- 11:02
GitHub.
- 11:03
The other thing that this lets us do is
- 11:05
we don't necessarily always need to use
- 11:07
our coding agent to implement uh pieces
- 11:11
of work. Uh Devin is very popular. We
- 11:13
have a lot of folks that really like
- 11:14
using Devin. They can point Devin to
- 11:17
these tickets and the same documentation
- 11:20
to give Devin the context it needs to go
- 11:22
and implement different parts of the
- 11:23
work. Uh same with the local cloud code
- 11:26
if you're just using Opus in a local
- 11:28
harness.
- 11:29
Um it can grab through MCP all of these
- 11:33
uh
- 11:34
all of these pieces of documentation and
- 11:38
use that as context for its work.
- 11:41
And then, of course, we have people that
- 11:43
collaborate in these project channels as
- 11:44
well. You'll notice our security team
- 11:46
will come in and look take a look at new
- 11:48
projects, weigh in on the security
- 11:50
implications, and so forth.
- 11:59
So, I alluded to we we kind of early on
- 12:03
in the the development of our factory is
- 12:05
created our own MCP gateway.
- 12:07
Um, we call it our context engine.
- 12:10
Um, this connects into all of our
- 12:11
internal systems, but it also builds
- 12:15
system prompts and context around what
- 12:18
how to navigate these tools and when to
- 12:20
use these tools.
- 12:21
So, it is connected into Snowflake,
- 12:25
which is our main data lake. We have
- 12:28
semantic tables that we've built in
- 12:30
Snowflake that describe product
- 12:32
utilization or customer conversations.
- 12:35
And then, we can provide an agent
- 12:38
through our MCP server in the tool
- 12:40
description
- 12:42
list of here are the tables, this is the
- 12:44
content they
- 12:45
contain. If you're trying to answer
- 12:47
questions about this type of content,
- 12:50
write a query for these tables. And so,
- 12:52
it gives us a little bit of orientation
- 12:54
to an agent to understand how to use how
- 12:57
we use our tools, how we've organized
- 12:59
linear in Snowflake, and give some
- 13:03
guidance on where to find information.
- 13:06
What has been really
- 13:08
surprising about this is we actually now
- 13:10
use this MCP server across a lot of
- 13:12
other different internal tools.
- 13:15
We open it up for people to query
- 13:17
directly from Slack
- 13:19
to do data or customer analysis. And so,
- 13:21
while we initially kind of envisioned
- 13:23
that this would just be a way to connect
- 13:25
our agent or coding agent to all of our
- 13:27
systems. Uh this has been an incredible
- 13:30
piece of leverage for a lot of our
- 13:32
internal teams, and we're building a lot
- 13:34
of other tools on top of this MCP
- 13:37
gateway. So, if you're just getting
- 13:38
started with thinking about a software
- 13:40
factory, it's well worth your time to
- 13:43
invest in an internal MCP gateway server
- 13:46
that both connects all your tools,
- 13:49
has descriptions that can tell the agent
- 13:51
how to use those tools and how you
- 13:53
specifically organize the information in
- 13:55
them, and I think you'll find that that
- 13:58
ends up being useful in a lot of other
- 14:00
use cases.
- 14:06
Like I mentioned, we try to automate
- 14:09
other parts of our software development.
- 14:11
Um we have like bug
- 14:14
bug requests that come in through Slack.
- 14:17
Uh so, TARS through webhooks can listen
- 14:18
to those and take a first pass at
- 14:21
triaging and implementing opening PRs to
- 14:24
fix bugs.
- 14:25
Uh we have a lot of our customers uh in
- 14:28
shared channels in Slack, as well.
- 14:30
Uh we found TARS to be pretty useful to
- 14:33
uh triage
- 14:34
support requests that we're getting
- 14:36
through them, because it can look at the
- 14:38
code, it can often understand like where
- 14:40
the customer might be running into
- 14:42
problems with our product uh from the
- 14:44
code perspective.
- 14:46
Um I showed you an example of how we
- 14:48
kind of organize major features around
- 14:50
this.
- 14:51
Uh and then the last place I think is
- 14:53
really exciting, we're all trying to get
- 14:55
to the self-improving software, or the
- 14:57
self-driving software. So, we're using
- 14:59
uh TARS to actually build out our own uh
- 15:02
sandbox infrastructure.
- 15:05
So, we want to move off of some of the
- 15:07
sandboxes as a service and own that
- 15:09
infrastructure layer ourselves. Um
- 15:13
Some of the reasons we have for that is
- 15:15
we want really deep control over the
- 15:18
session information, and
- 15:20
uh be able to move workloads around
- 15:23
different parts of our infrastructure.
- 15:25
Uh so what we're building out next is
- 15:27
our memory layer and we want this to be
- 15:29
an evergreen context uh about what every
- 15:34
person in the company is working on,
- 15:36
what team do they sit in, what products
- 15:38
are they responsible for, and then at an
- 15:40
organization level, what is the
- 15:42
semantics about Work OS and how we do
- 15:44
our work. And this we want to be able to
- 15:46
plug into both our software factory, but
- 15:49
lift that context out of our factory and
- 15:51
into our other AI tools as well.
- 15:57
So, I think this is kind of our our
- 16:00
major takeaway of where we've seen a lot
- 16:03
of folks invest effort just around
- 16:05
outputs and where we think we can
- 16:06
actually get that human exponential
- 16:09
value out of our factory.
- 16:11
Can we be actually delivering values to
- 16:14
our customers? We think a lot about
- 16:16
shipping and customer impact and so we
- 16:18
want to encode that into their our
- 16:20
software factory itself. Um certainly
- 16:24
worried about, you know, introducing
- 16:26
instability into our products. So, all
- 16:28
right, can we measure
- 16:30
uh defect rate and other
- 16:32
um
- 16:33
you know, time to recovery uh metrics?
- 16:36
Um and then a lot of anecdotal, like we
- 16:38
want to see our engineers using TARS as
- 16:42
a sign that this is giving them value
- 16:44
and accelerating their work. We want to
- 16:47
see that they're moving off electing to
- 16:48
move out of their local harness and into
- 16:51
a cloud sandbox.
- 16:58
And then the real goal is we can use
- 17:00
this information because we own the
- 17:01
infrastructure, we can see what's
- 17:03
happening in the infrastructure. We use
- 17:05
this information to self-improve the
- 17:06
factory.
- 17:08
We want it to learn, we want it to get
- 17:09
better about how it writes code.
- 17:13
We want to see where um our engineers
- 17:17
can improve their skills. There's so
- 17:19
much moving quickly in AI. It's kind of
- 17:22
like a constant race to keep up with the
- 17:24
latest
- 17:25
um
- 17:26
you know, tips and techniques.
- 17:28
We can actually point agents at sessions
- 17:30
and see where are the gaps and how
- 17:32
people are using
- 17:34
our factory. Where is there a skill that
- 17:36
we should be building because the agent
- 17:38
made a mistake. What skills are no
- 17:40
longer relevant? Maybe the code has
- 17:43
changed to a degree where a skill that
- 17:45
we wrote six months ago is now obsolete.
- 17:48
So, that continues verification.
- 17:50
Semi-online
- 17:52
is where we kind of see a big advantage
- 17:54
to actually owning all the
- 17:56
infrastructure ourselves.
- 18:02
So,
- 18:04
say it for the third time, it's worth
- 18:05
repeating.
- 18:07
Sandboxes are great.
- 18:09
It's cool to run an agent and have it
- 18:10
open a PR.
- 18:12
We really think about our software
- 18:13
engineering practices and processes and
- 18:15
we want to encode those in automation.
- 18:21
I'm Ryan. I'm one of the engineers here
- 18:22
at WorkOS. We have a booth right
- 18:24
opposite the speaking area. If you are
- 18:26
building software factories, I would
- 18:27
love to talk to you. I would really love
- 18:29
to hear how you're handling
- 18:30
authorization. This is something I don't
- 18:32
talk about because we haven't figured it
- 18:34
out yet for ourselves, but maybe you all
- 18:37
you all have some insights. So, please
- 18:38
come by the booth, talk to me, would
- 18:40
love to chat. Thank you.
- 18:55
>> [music]