AI Engineer World's Fair 2026

From 36% to 100%: How Self-Improving Agents Write Their Own Skills — Rafal Wilinski, Runlayer

Read the talk

From 36% to 100%: How Self-Improving Agents Write Their Own Skills

Rafal Wilinski explains how Runlayer turns agent traces into reusable playbooks, distributes them through MCP, and uses the resulting feedback loop to preserve company knowledge across sessions and models.

From a talk by Rafal Wilinski

At a glance

Ideas worth remembering

  • Skills use progressive disclosure: a short description helps an agent find relevant knowledge, then the full procedure enters its context to guide the attempt.

  • MCP can distribute skills from a central server that selects procedures by query and role. Runlayer uses tools because MCP resource support is uneven.

  • Successful traces supply reusable routes; failed traces expose missing guardrails, dependencies and edge cases that can improve the procedure.

  • In the reported Chromium-on-AWS-Lambda evaluation using Sonnet 4.6, success rose from 9/25 runs (36%) without a skill to 25/25 (100%) with the distilled skill.

  • Company-wide learning requires more than distillation: candidate skills pass checks, enter a shared library and reach future agents through the MCP gateway.

Keep the solution after the session ends

A hard problem gets solved, the session closes, and the next person starts from scratch. The missing artifact is often the useful part: the debugging steps, rejected approaches and reasoning that made the solution possible. Rafal Wilinski, founding engineer at Runlayer and previously a leader of Zapier’s agent work, opens with a practical ambition: make every successful investigation leave a playbook the rest of the company can use.

That ambition extends beyond one agent remembering its own work. An engineer, salesperson or lawyer who cracks a difficult problem can improve the next attempt by sharing the procedure. Automating that handoff creates the possibility of company-wide learning: one agent’s discovery becomes another agent’s starting point.

0:120:35
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:12 · section reference included

Longer tasks make the starting trajectory more expensive

A skill is a playbook an agent reads when its knowledge becomes relevant. It uses progressive disclosure: the client or harness initially sees a short description; when the agent identifies a useful match, it reads the full skill into its working context. The procedure can then change which actions the agent takes. Better models, prompts and tools remain other ways to improve an agent, but skills offer a direct way to supply knowledge at the moment it is needed.

The stakes grow as models become capable of longer, multistep work. Wilinski points to a task-horizon benchmark measuring how long agents can work independently while retaining a chance of completing the task. That word—chance—matters. A longer reasoning window also lets an agent spend hours and hundreds of tool calls pursuing a wrong hypothesis. Without a useful starting procedure, greater persistence can make a mistake more expensive.

The most valuable skills often contain what Wilinski calls “the rituals that nobody wants rediscovered.” These are local details that general model capability does not automatically supply:

  • Production rollback: the sequence of scripts and feature flags needed to reverse a change safely.
  • Customer retention: the wording and approach appropriate for an enterprise customer close to churning.

Both examples depend on how a particular company works. Asking an agent to reconstruct that knowledge on every attempt adds exploration precisely where mistakes can be costly.

The slide presents a short skill description beside a fuller skill procedure.Open full source frame
The slide presents a short skill description beside a fuller skill procedure.
2:002:30
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

2:00 · section reference included

Useful playbooks still need authors, access and trust

Skills have three separate adoption problems. Solving one does not automatically solve the others:

  • Developer-oriented packaging: Markdown, JSON, CLI commands and Git repositories are familiar to developers. They make creating and sharing skills harder for people in marketing, HR, legal or finance.
  • The cost of writing: everyone likes finding an existing playbook, but producing one requires time and deliberate reflection after the difficult work is finished.
  • Uneven client support: a useful skill cannot guide an agent if the employee’s LLM client has no way to discover or load it.

Installation introduces another practical concern. Wilinski jokes that a skill installer can bring along prompt injections, a script requiring root permissions and somebody else’s API key. The joke comes with a concrete observation: Runlayer scans thousands of skills every day and encounters suspicious content. A playbook influences the agent’s decisions, and its package may also include executable material. Making installation easy therefore increases the importance of checking what is being installed.

5:015:31
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

5:01 · section reference included

Use MCP to deliver knowledge as well as capabilities

Skills and MCP serve complementary purposes. A skill explains how to do something; a tool supplies the capability to do it. Putting a command or API key in a Markdown file does not remove the need to connect agents to systems. MCP can also carry the playbook itself, giving knowledge a shared distribution channel alongside tools.

In this design, clients connect to a remote MCP server and search for relevant skills instead of installing their own local copies. The server selects what to return using the query and policy, including the requester’s assigned role. A central source can supply the appropriate version of a procedure and give an AI adoption lead one place to manage the knowledge being shared. Wilinski reports that Runlayer had used this approach internally for six months.

The implementation choice is about client reach. The Skills over MCP working group was developing an API, with the specification direction leaning toward MCP resources. Runlayer chose tools because resource support was mixed. The API was not final at the time of the recording, and practical reach still depends on the client’s MCP support. Within that constraint, distributing skills through tools lets more employees retrieve procedures through the clients they already use.

Central distribution addresses access and governance. It still leaves the authoring problem: somebody has to turn an investigation into a useful skill. The next step is to let completed agent work supply the material.

The diagram places an MCP gateway between client tools and several skills.Open full source frame
The diagram places an MCP gateway between client tools and several skills.
7:277:58
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

7:27 · section reference included

Turn exploration into reusable procedural knowledge

Voyager, the 2023 Minecraft agent work, provides the model for this kind of learning. In Wilinski’s account, the agent explored and pursued goals while accumulating a skill library, without changing the underlying model weights. When it discovered how to do something, it saved a code recipe containing the tool calls that accomplished it.

A later goal—making a crafting table or a diamond pickaxe—could then retrieve the recipe rather than reconstruct the procedure. The accumulating library changes what the agent has available for its next attempt. This is the idea Runlayer carries into workplace tasks: exploration should produce reusable procedural knowledge instead of ending as a discarded trace.

Runlayer lets agents perform real work, then takes an interesting successful run containing multiple tool calls and asks a frontier LLM to distill a skill from its trace. The saved artifact is a procedure for future work. Authoring happens after the run, asynchronously, rather than depending on someone remembering to write documentation.

Failures supply a different kind of information. A successful run demonstrates a route that worked; a failed run can expose a missing guardrail, an unavailable library or an overlooked edge case. Both feed the distiller. Successes support the existing procedure, while failures provide reasons to revise it. Wilinski describes this update loop as autonomous, without a human in the loop.

The slide shows a skill library alongside a diagram of exploration and reusable knowledge.Open full source frame
The slide shows a skill library alongside a diagram of exploration and reusable knowledge.
9:4710:16
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

9:47 · section reference included

A Chromium-on-Lambda task moves from rediscovery to reuse

The concrete test was getting a Chromium fork to run inside AWS Lambda using Sonnet 4.6. The agent sometimes completed the task, but its initial success rate was 36%. Runlayer enabled its self-improvement mechanism. When one invocation solved the problem, the distiller generated a skill from that run and included it in all subsequent runs. Wilinski reports that the success rate rose to 100%. The slide below answers the immediate question: how many attempts sit behind those percentages? It labels the comparison a 25-run evaluation, showing 9/25 successes without the skill and 25/25 with it. The recording does not give the full evaluation protocol or contents of the skill, so this result describes the reported test rather than a general guarantee of reliability.

The comparison makes the observable change clear: occasional success became consistent success within the reported evaluation. Initially, separate invocations had to find a working route. One successful invocation supplied a trace; the distiller converted that trace into reusable instructions; later invocations received those instructions before attempting the task. The successful run stopped being an isolated event and became input to future behavior—“a developer with better attitude towards documentation,” as Wilinski puts it.

Why might the same change also reduce cost? A good skill narrows the search space. It steers the agent toward a known procedure, reducing the branches it explores and the tool calls it spends on rediscovery. Wilinski presents fewer tokens and lower spending as consequences of that mechanism; the numerical result here measures task success.

The slide compares 36% success without a skill with 100% with a skill.Open full source frame
The slide compares 36% success without a skill with 100% with a skill.
12:3613:06
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

12:36 · section reference included

Share the learning loop across the company

There is another reason to retain procedures: access to frontier intelligence is rented. A provider can deprecate a model, and users may experience changes in its behavior. Wilinski’s response is to use capable models while they are available to discover and persist useful procedures. A catalog of how work gets done keeps something the company can carry forward when its model changes.

The organizational version combines three pieces already introduced: skills retain procedures, a central library stores them, and MCP distributes them. Runlayer groups similar runs, distills a candidate skill, and passes it through checks that include prompt injection and personally identifiable information. A skill that passes goes into the library and becomes available through the MCP gateway to clients and agents across the company.

Runs → candidate skill → safety checks → shared library → MCP clients. Each new attempt contributes traces, including failures that can reveal where the procedure needs to change. Checks happen before a skill becomes available to other agents.

The shared library is what extends learning beyond the original agent. When another employee’s agent encounters the same database restart or production rollback problem, it can retrieve the accumulated playbook. As more runs expose useful procedures and edge cases, the library can improve the starting point for subsequent work throughout the organization.

13:5914:29
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

13:59 · section reference included

The advantage is the knowledge accumulated in doing the work

The ending turns to competitive advantage. If competing companies can access the same models and reproduce similar product interfaces, where does differentiation come from? Wilinski proposes the accumulated learning loop: governed procedures drawn from a company’s own hard-won experience. That is a strategic argument about the value of internal knowledge, rather than a demonstrated barrier to competition.

The proposed advantage has a practical form. Each solved problem can raise the starting point for the next person; each useful failure can add a missing restriction or edge case. Retained procedures can reduce repeated exploration without changing the model and preserve useful knowledge through model transitions. The company’s particular way of working—its internal culture and identity—becomes something agents can retrieve and apply, rather than something lost when a session closes.

16:1216:42
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

16:12 · section reference included

Resources

Read the complete timestamped transcript
  1. 0:12

    Hi, everyone. I'm Rafal. Today, I wanna talk about self-improving agents, but not in a single player mode where just one agent gets smarter. I wanna talk about how we can apply that concept and mix it with network effects to get a self-improving company or maybe even, uh, an enterprise that self-improves while you sleep.

  2. 0:35

    Um, imagine the best engineer you ever worked with, or a salesperson, or a lawyer. Um, the one that every time they cracked a genuinely hard problem, they wrote down exactly how they did that and handed over the playbook to the team. So whenever they face a similar problem, they don't have to figure it out from the scratch, right? And imagine that happening for every hard problem across your whole company, and that happening automatically, uh, each time. That's a dream, right?

  3. 1:06

    But it almost never works that way. Um, even when you crack a really hard problem, um, oftentimes the knowledge, um, all that debugging, um, all the, all the reasoning behind that is just gone. Oh, sorry. Um, all that, uh, all that reasoning is just gone when you move on to the next, the next task or close the session. But I wanted to convince you that this is actually possible, and we already have all the primitives in place to make it happen.

  4. 1:37

    If it sounds interesting, let me intro- introduce myself once again. Hi, I'm Rafal. I'm a founding engineer at Runlayer. We are making the golden path for AI at your company. Um, before that, I was leading Zapier agents and different AI teams, but generally speaking, my previous work was about making agents not just flashy demos, but reliable things in production.

  5. 2:00

    Um, before we go into interesting parts, I wanna break down my topic into three separate pieces and start by just improving agents. There are many ways to do that. We can use better models. We can use better prompting tools, harnesses and whatnot. But I wanna focus today only just on skills because they are super accessible, right? At least in theory. Um, to make sure that we are on the same page, a skill to me is a playbook that essentially agent can read, um, when it thinks that the knowledge is

  6. 2:30

    relevant. From the technical perspective, the, the term is called progressive disclosure, and it means that at the very beginning, the harness or a client sees just a summary of the skill or a short description of the skill. But when it thinks that the knowledge inside the skill can be useful, it can read the whole of the skill, inject it to the main context, and based on that knowledge, it can change the trajectory of the agent. Why is this so important? Well, if you haven't been living under a rock for the past

  7. 3:00

    six or twelve months, your workload probably has changed a lot, uh, and that's because the agents got so much better. Actually, the models got so much better. And, and it's not just vibes. It's measurable. Um, here you can see a chart that probably you have seen too many times already. For those who don't know, it's METER benchmark. It measures for how long agents can work independently while still having a decent chance of completing a task. And the word chance is essential here. Um, we are

  8. 3:30

    going to go back to it in a moment. But the trend is clear. The frontier capabilities, the frontier intelligence is getting better and better. They are better at multi-step deep work. There's only one caveat if that frontier intelligence is not being taken from you, and I'm as fabled probably just like you.

  9. 3:50

    But there is a catch. With great power or long task horizon comes great responsibility. If an agent can reason for much longer, it also means that it can dig a really, really, uh, deep rabbit hole. Um, it can spend hundreds of tool calls trying to defend a wrong thesis, reason for hours. The next agent can do exactly the same thing or maybe even do something wrong, um, worse. Um, so the initial trajectory matters a lot, and that's why skills are so

  10. 4:20

    important because they guide the agent towards the right target. Otherwise, you'll waste time, you'll waste tokens, and most importantly, you'll also waste money.

  11. 4:31

    Um, valuable skills are the rituals that nobody wants rediscovered. That knowledge is often not glamorous. It's probably deeply embedded within your company. And if agents would try to rediscover that knowledge from scratch, like for instance, the sequence of scripts and feature flags in order to make a nice rollback, or what's the exact wording and language to use when dealing with that pesky es- uh, enterprise customer that's about to churn. Um, yeah, if agents would try to figure out that from

  12. 5:01

    scratch, they are likely going to fail, and oftentimes that failure can be really costly. That's why I believe so many AIPs or POCs are failing because you're not giving agents the relevant context and the playbooks how to act within your company. So skills are great, but I feel like they are not problem-free. Um, there are three most important problems, and the first one is that they are local first and dev-centric. What I mean by that is that, um, when you

  13. 5:31

    look at the skills ecosystem, you're going to find, um, Markdown files, JSON files, CLI commands, Git repositories, and if you're a developer, that's great. But if you're, um, ado- uh, VP of AI or AI adoption lead, you need to think more empathetically about how you can roll out the AI strategy across your whole company, which includes non-technical people, right? And the challenge here is how non-technical people can actually create and share, uh, the

  14. 6:01

    skills. Uh, I love this meme. Um, this is the junior developer or a non-technical person discovering command ex- npx skills, which is the most important, but the most popular command to install them. That's the only command I know that installs just, uh, not just the skill, but also free prompt injections, um, a wild script that requires root permissions, and someone else's API key all in one run. Trust me, at Runlayer, we are scanning thousands of skills every day, and

  15. 6:31

    there is really weird stuff inside. Okay. Second problem is that I think we all like-- we all love the situation where we are facing a hard problem, and there's already playbook on how to deal with that situation, right? But oftentimes there's, there's no playbook, right? Because it time- it takes time, and effort, and intention, and energy to reflect on a problem and leave something useful for posterity behind. And the third problem is that not every

  16. 7:01

    LLM client supports skills, right? If you're in cloud ecosystem, it's probably great. If you're a developer, you're probably going to manage, even if the skills support is weird, depending on which client you're using. Um, and once again, if you think about marketing, HR, legal, finance, those departments are probably still using ChatGPT or some other weird LLM client, and skills might be not supported there.

  17. 7:27

    And some time ago, there was this big heated debate when the skills were introduced. Um, some people proclaimed that, "Thank God MCP is finally dead because I'm not interested in dealing with OAuth. I can now hard code a bunch of CLI commands, um, hard code my API key inside Markdown file and call it a progress," right? Um, but I think everyone missed the point. Skills can be delivering, um, the knowledge how to do something, and tools gives us capabilities. Tools from MCPs give us capability to do those things. But

  18. 7:58

    MCP can be also used as a protocol to deliver the knowledge, to deliver the skills. We can use MCP as the distribution layer. So instead of installing skills locally, you can create a remote MCP server, which can become the source of truth across your whole company. Clients can connect to that. They can ser- they can search for the right skills. Uh, server can enforce the policy and make sure that everyone gets the best version of the skill possible, depending on the query and, for instance, the role that they

  19. 8:28

    have assigned. We've been doing that internally at Runlayer for the past six months, and it's been working quite well. Um, but it's no longer just an internal thing. We see that the broader MCP ecosystem is also going that direction. Um, there's already a working group called Skills over MCP, and while the exact API is not final, direction is clear. MCP can be used to deliver the skills, to deliver the knowledge across your whole company.

  20. 8:57

    Um, yeah, and while the official spec is leaning towards MCP resources, the support for resources is also mixed, so we are going with tools. Um, and yeah, if you remember the problem, the first problem I mentioned and the third problem I mentioned, which is, um, the problem that they are dev-centric and the lack of, uh, support for clients, MCP as a distribution layer fixes that because it's unified and on-- and works in almost every single client, not just developer

  21. 9:27

    clients. So the problem of devel- of distributing skills to non-technical people is essentially solved, right? Plus, you have now a centralized API, so as a VP of AI or, like, AI adoption lead, you have one place to govern and manage what's the knowledge that's getting shared and distributed across your company.

  22. 9:47

    Okay, so now we know how we can improve agents because we have, um, skills. We know that we can use MCP to distribute those skills across all the clients and agents. But let's talk about something more interesting, which is how agents can improve themselves. And whenever I, uh, see a new paper or a framework speaking about this topic, I always go back to, um, Voyager paper. Unfortunately, you cannot see the video, but it's Minecraft gameplay in, in background. So in, in twenty

  23. 10:16

    twenty-three, right after GPT-4 was released, so ages ago, the question was: For how long multiple agents or harnesses can explore Minecraft space, right? Um, they dropped the agent into Minecraft with no instructions. They set its own goals. Agent just kept getting better without even touching the weights or, or model and the prompts. And in twenty twenty-three, I think that was pretty groundbreaking. How did they pull it off? Um, the key thing to

  24. 10:46

    notice here is the skill library. So Voyager did not just explore the Minecraft space care- carefully and throw the trace away. When the agent figured out something, um, it saved the recipe as a code that, uh, that consisted of the tool calls that lead it-- that lead the agent to the discovery. So later on, when the agent decided that, hey, I need to create a, a crafting table or a diamond pickaxe, it didn't have to figure it out this whole knowledge from scratch. It could

  25. 11:16

    simply go to the skill, uh, library and retrieve it as a code, as a recipe, um, that can be reused. And this is really important. This is the core idea that I wanna carry forward, is that, um, exploration should be turned into reusable procedural knowledge.

  26. 11:36

    And we started doing something very similar at Runlayer. Um, we let the agents do the real work, and when the run succeeds and when there is something interesting in this run, there's multiple tool calls, um, we just take the trace, and we ask Frontier LLM to simply distill a skill, right? And while it's super simple to distill a skill from a single run, um, I think we humans and agents learn the most from mistakes, right? The real

  27. 12:06

    sharpening happens when the agent fails. Um- So a run that would go sideways is actually the richest signal that you can get because it exposes the missing guardrail, the missing library, some kind of edge case that you haven't thought about upfront. Um, and every run or, or pass or fail just feeds our distiller mechanism, constantly making our skill better. Um, the successes are making the skill more trustworthy because if the agent achieved the goal, it

  28. 12:36

    means that, hey, this skill works, and if it doesn't, it actually makes skill the more robust because now we know what's missing in here, and it all happens completely autonomously, asynchronously without human in loop. And the initial results were great. Uh, we had, uh, this task of trying to run some Chromium fork inside AWS Lambda, and we were trying to do that using Sonnet four point six, so not the most intelligent model, but sometimes it succeeded. Um, but

  29. 13:06

    it was just, you know, it's not production ready to have thirty-six percent success rate. Um, so what we did is we flipped the flag for self-improving agents, um, and once in one run, the... one of the agent invocations cracked the code, our skill distiller create a reasonable skill. It was, uh, then included in all the future runs, and the success ratio went to one hundred percent. So yeah, that's great. It's like a developer with better attitude towards documentation. Um,

  30. 13:36

    the success rate is the easiest number to show, but good skill is narrowing the search space. It makes the agent... Uh, it narrows the search space, and agents starts doing, um, stops doing tool calls, and it's exploring less branches. It's wa- it's wasting less tokens and, and less money.

  31. 13:59

    But putting savings and accuracy aside, I think there's also a deeper reason why you should be distilling skills. Um, the problem is that frontier intelligence is rented or borrowed. It can disappear. I think, yeah, we all miss Fable. It was a great model, right? It can deprecate it, or it can get dumber. Like, we complain every single day on Twitter, on Slack that Opus is really dumb today, right? So while we have the access to the best version of that intelligence, we should be using that

  32. 14:29

    to discover and persist those use- useful procedures. We should be building a catalog of how the works gets done inside our company as a collective subconscious.

  33. 14:42

    And how we do it? Well, I believe we already have all the pieces of the puzzle here. Um, we have skills as a way to retain procedures and improve agents and guide models. We have MCP as a unified distribution layer for those skills, and you also have a central repository, the knowledge base for all of those skills, which is constantly being updated and fed with new data. Um, yeah, we also have the continuous learning flywheel for one

  34. 15:12

    agent, but what's interesting is what I promised you at the very beginning, the network effect. Um, so what happens when you multiply that by every agent and every single client in the entire company? At Runlayer, we call it self-improving organizational flywheel. Um, we group a bunch of similar runs and try to distill a skill. It passes through our various series of checks, like whether it's not containing some kind of prompt injection or maybe PII, um, and bunch of other

  35. 15:42

    stuff. Once it's good, it's landing in our skill library, and thanks to our MCP gateway, it's being served to each and every client and agent within the company. What we end up having is a constantly evolving knowledge base, which is always up to date, continuously updated with new alpha, fed with new edge, uh, edge cases. So once one agent figure out a really hard problem, it's getting dist- it, it's being distilled into a skill persistent in all organ- organizational library, library.

  36. 16:12

    So whenever someone else, um, has the same problem in the future, they don't have to figure out the solution from scratch. They don't have to know, uh, what's the exact sequence of this magical commands that are going to restart the database or roll back the production. They can simply, um, reach out to the playbook they have. And why is this so important? Because I believe, like, fifty percent of us probably already had an existential crisis thanks to the models, and

  37. 16:42

    the thinking goes along the lines of, uh, okay, so models are becoming commodity. Everyone here has access to them. Um, cost of software is going to zero. Everyone is... can be rebuilding my product, um, at least on the surface. In fact, so many people are working today across the expo, and they are thinking, "Hey, everyone's building the same thing, right? So what is my moat?" Um, and I believe the moat can be in building this flywheel

  38. 17:13

    and distilling the knowledge to this flywheel because it's like your own hard-won lessons distilled, governed of how your company, uh, does things, you know. And it's captured, compounding, raising the floor for everyone across the company, giving them access to the best knowledge available. Um, it is also a way to be resistant against model intelligence deprecation, and it's a great way to lower your costs without ever changing a

  39. 17:43

    model. So yeah, that's probably one thing that competitor can't copy, your internal culture and your identity.

  40. 17:53

    And yeah, that's pretty much it. Uh, thank you for coming. If it resonated with you, uh, I'm Rafal. I'm working at Runlayer. We are figuring out the AI golden path for you. I'm open to talk about skills, MCP agents, whatnot, and this is exactly what we are building at Runlayer. Thank you.