Turning Agent Memory Into Skills That Work — Will Lyon, Neo4j
Read the talk
Turning Agent Memory Into Skills That Work
Will Lyon explains how Neo4j connects conversations, entities and decision traces, then distills that memory into reusable procedures whose supporting evidence can be checked as it changes.
From a talk by William Lyon
At a glance
Ideas worth remembering
Entity resolution and a shared ontology turn retrieved mentions into connected domain knowledge with canonical identities and explicit types.
Decision traces preserve evidence, policies, tool calls, results and execution costs so experience can be shared across agents.
Skill distillation needs grounding, coverage and coherence: supporting some claims is insufficient if other steps lack evidence or the input combines unrelated topics.
Connections from skill components to their supporting memory let curation detect when changed, removed or contradictory data makes a procedure stale.
The patient-intake demonstration turns an earlier conversation into steps through charting, grounded in tool calls and entities and packaged with SKILL.md and references.
A successful run can still leave the next agent starting over
An agent reasons, acts and successfully completes a task. Then it forgets what it learned. Will Lyon, a product manager at Neo4j, opens with this “amnesic loop”: finishing an attempt does not necessarily make the next attempt better. A memory system may preserve the conversation without preserving a usable account of the people, actions and decisions involved.
Embedding-based memory addresses part of that problem. It stores text, finds similar chunks when a new request arrives, and inserts those chunks into the context window. The agent must then work out what they mean and how to use them. Similarity retrieves relevant wording; it does not itself establish that several descriptions refer to the same thing.
The opening example is a provider described in several ways: by name, by an abbreviated name and as a cardiologist. Retrieving those mentions leaves an identity question for the agent to solve. A canonical representation gives the mentions a common referent. That is the first requirement behind Lyon’s formulation that agent knowledge should be “connected, typed, and traceable.” Retrieval remains useful, but the memory must also describe what its contents refer to and how they fit together.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Turn mentions into entities with a shared domain model
Neo4j’s context graph connects three kinds of memory: short-term, long-term and reasoning memory. The first transformation starts with user and assistant messages. Entity extraction identifies the things discussed and their relationships; the graph records those relationships with explicit types. The result carries more structure than a collection of text passages.
Two distinct mechanisms make that structure useful:
- Entity resolution: Different mentions must resolve to a canonical entity. For the provider example, the graph can associate the person with a name and role rather than leaving each mention as an unrelated fragment.
- Shared ontology: A domain model describes the kinds of entities and relationships the system recognizes. In the later healthcare demonstration, that vocabulary includes providers and encounters. It gives extraction a common structure to work toward.
The distinction matters: identifying a provider-shaped mention and deciding which provider it denotes are separate jobs. The ontology describes the kind of thing; resolution establishes its identity. Lyon places particular weight on resolution because the connected graph depends on getting that common identity right.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Preserve the decisions and tool results behind the outcome
Messages and extracted entities describe what the conversation contained. Reasoning memory adds the actions taken during the run. Its organizing unit is a decision trace: a decision linked to supporting evidence, explicitly modeled policies, the execution plan, tool calls and their results. Token use and elapsed time belong in that record too.
This makes a completed run useful beyond its final answer. A later agent can encounter both the domain entities and the recorded route through the tools. Lyon’s intended setting includes hundreds or thousands of agents with overlapping tool access. Persisting their traces in one organizational context graph gives them a shared record from which to learn, rather than confining each run’s experience to one agent.
What information must remain connected for another agent to understand a decision? The diagram separates conversational memory from execution memory, then joins them in the shared graph. That connection preserves both what was discussed and what was done with it.
The conversational record supplies mentions and relationships.
The context graph connects domain knowledge with evidence about agent actions, making both available across agents.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Give reusable skills explicit execution steps
Memory records what happened. A skill packages guidance for what an agent should do next. The skill format described here starts with metadata explaining its purpose, then uses progressive disclosure: more detailed Markdown and reference material enter context when the agent needs that part of the procedure. Packaging makes the guidance reusable across agents without loading every reference upfront.
Prose still leaves interpretation work. A procedure can refer ambiguously to an entity, describe a step loosely or make its output difficult to debug. Neo4j’s research addresses these problems by extending the skill representation with metadata that treats the procedure as a typed execution graph. Steps become nodes, and a strict schema governs their descriptions and how they are actionable.
The graph representation makes the procedure’s parts explicit while retaining the skill as a shareable package. It also creates a more structured object to inspect: the system can reason about individual steps instead of treating the entire procedure as one block of prose. Lyon reports an increase in successfully completed SkillsBench tasks when the research protocol was applied to human-curated skills. The recording gives no numerical result or evaluation details, so this supports a reported benefit in that comparison rather than a quantified expectation for other workloads.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Check the evidence across the skill—and keep checking it
Skill distillation starts from a context graph scoped to a workspace or project and turns observed experience into procedural guidance. The goal is to retain well-understood steps and their supporting data. That support also gives the system something to monitor: when the memory changes, the procedure derived from it may need to change.
Three checks ask different questions about a candidate skill:
- Grounding: Does the skill draw on data actually present in memory?
- Coverage: Are the steps and descriptions supported throughout the skill, rather than only in one part?
- Coherence: Does the selected subgraph describe a sufficiently unified topic, or has it combined material that should become multiple skills?
These checks constrain what gets packaged. Evidence for some statements does not establish support for every step, and a well-supported collection can still mix several topics. Governance continues after generation: a dynamically loaded, governed skills registry can provide an up-to-date skill and expose whether it has become stale or drifted. The useful unit is therefore a procedure with maintained connections to its evidence.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Choose a scope before synthesizing the procedure
The Neo4j Agent Memory Service, abbreviated NAMS, implements this pattern as a Neo4j Labs service. A workspace-scoped graph holds the three memory types. Background workers distill skills, while background curation surfaces skills whose supporting memory has changed.
The first consequential choice is scope. A skill can start from the subgraph around an entity or from a specific conversation. That selection determines which experience the system will attempt to turn into a procedure. Lyon describes most pipeline steps as deterministic; the LLM synthesizes some claims and generates text for the skill description. Deterministic processing surrounds the language-generation work, rather than making the entire process deterministic.
Each component retains support in the underlying graph. Curation can then watch for three consequential changes: supporting data changes, disappears or becomes contradictory. Coherence uses graph algorithms such as community detection. If the selected material spans several topic communities, that can indicate that the proposed skill should be decomposed. It is a signal for deciding how to divide the material, rather than a claim that every cross-community procedure is invalid.
How does a change in memory reach an already generated skill? The flow below makes the evidence connection visible. Distillation produces the procedure, but curation follows the supporting data afterward; packaging does not end the skill’s relationship with memory.
Start from an entity subgraph or a conversation.
The skill graph retains support for its components, allowing later memory changes to trigger staleness signals.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
From a healthcare conversation to a patient-intake skill
The demonstration starts with memory already ingested. NAMS exposes REST APIs and MCP tools for ingesting, retrieving and working with the three memory types. Its dashboard provides a graph view that can be traversed, a view of flagged entities and the domain ontology. Here, the healthcare ontology organizes data about encounters and providers, and the workspace contains healthcare-agent conversations.
The next action narrows that collection to a skill-sized input. The interface offers an entire workspace, an entity, specific conversations or a class in the ontology as possible scopes. Lyon selects the most recent conversation. That choice queues a distillation job, which fetches the relevant data, runs a seven-stage pipeline, constructs the skill graph and packages it.
While that job runs, the demonstration opens an earlier result derived from a patient-intake conversation. The observable change is from a record of one interaction to a procedure covering patient intake through charting. The extracted steps remain grounded in the actual tool calls and entities that supplied their components. The earlier conversation therefore contributes both the procedural sequence and inspectable support for it. This is a previously generated result; the demonstration does not establish completion of the newly queued job or show the resulting skill executing a fresh intake.
The output includes SKILL.md; a downloaded package would also contain references following the progressive-disclosure pattern. The graph representation and the familiar skill package serve different parts of the same workflow: structured steps retain their evidence, while the package carries guidance that another agent can load as needed. The talk ends by pointing toward slides, service documentation and open-source tooling for implementing these patterns.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
Related talks
- Beating RL With Reflection: GEPA and Optimize Anything
A complementary approach to learning from execution traces: reflection changes prompts, agent programs and repository skills, while Lyon focuses on structured memory, evidence connections and skill governance.
Read the complete timestamped transcript
- 0:12
Let's go ahead and get started. I see some folks wanting to, to come on in, so feel free to, to come in. Uh, this session is gonna be all about actionable knowledge and context graphs. So my name's Will, I'm a product manager at Neo4j. Um, there's been a lot of discussion in this conference so far about loops, right? The, uh, React loop, the Ralph loops. Sometimes I
- 0:42
feel like we're in maybe like a, an amnesic loop, right? Where our, our agents go through this, uh, reasoning phase, acting, they successfully complete some task, but then kind of forget, uh, what they've learned. And memory systems today are largely focused around, uh, embedding text, then at retrieval time, sort of finding the most relevant data to, to stuff into the context, right? Um, so, so something like this. We
- 1:12
find, you know, similar chunks of data in, in our, our corpus. We inject that into the context window and trust that the agent is able to do something useful with that, right? Now, the, the challenge that we have is, you know, recall isn't really, like, actionable knowledge. So one challenge we run into is, like for example, we don't have a canonical representation of a thing, right? So if we have, you know, three different ways that we refer
- 1:42
to Dr. Nguyen, Robert N, the cardiologist, right? Like, depending on the, the context of this discussion, we need to have some canonical representation of the thing, right? So really, agents need knowledge that is both connected, typed, and traceable, right? Not just retrievable. So retrieval is, is, uh, just a part of the problem when we're talking about, uh, agent memory. So at Neo4j, this
- 2:12
is how we think of agent memory. We think of it as a connected graph composed of short-term, long-term, and reasoning memory. Uh, we'll take a look at, uh, this in a bit more detail, uh, but bear with me on, on this idea of context graphs for agent memory, right? So what are a few of the, the key components here? Well, one is going from unstructured data to a knowledge graph, right? Going from the, this process of
- 2:44
agent messages, right, both u- user and assistant messages, going through some entity extraction process, where we're identifying what are the entities, uh, and how are they connected. Um, doing this in a graph where we have strong types, right? We have a, a relationship that describes how these entities are related.
- 3:08
The entity resolu- resolution phase is, uh, one of the most important parts of building up this, this knowledge graph, right? Understanding and making sure that you have a canonical representation of the thing. Uh, so when we're talking about Dr. Nguyen, we know this is a provider, his name, the role that he has, uh, and so on. And a shared ontology is an important piece of this, right? So having some description of your data model of the domain that you're working with, this is one
- 3:38
of the key pieces, uh, of making sure that you're successful in going from that unstructured text data to, uh, to a knowledge graph. So for folks that have, that have worked with memory, uh, i- in agents in the past, like, this, this should look a little familiar. Uh, one thing though that I think is really important when we're talking about agent memory and actionable knowledge is this idea of the reasoning graph, right? So, um,
- 4:09
if we saw in our, our, uh, sort of three components of agent memory here, reasoning memory is a first-class citizen, right? We've talked about short-term and long-term. These are the messages, the entities extracted. But what about the actual actions that the agent takes? What about the, the reasoning? Um, we wanna make sure that we're capturing that because that is an important piece, uh, as well. So we wanna make sure that we're storing thinking, the facts, not just, uh, not just the data, right?
- 4:38
So how an agent decided, and this is typically represented as decision traces, right? So every decision, uh, that our agent makes is linked to, uh, the reasoning grounded in evidence, right? We have policies, um, that, that we're modeling explicitly, and we understand for that execution plan, the tools the agent called, the results of those tool calls. Uh, we also wanna capture things like how many tokens did this burn,
- 5:08
the, the timing that this took, and so on. This is all part of, uh, capturing the reasoning memory. And th- this is important, like, not just so that a single agent can get better the next time you do it to do the exact same thing, but rather think about systems where you have hundreds or, or thousands of agents that have some shared grouping of, of tools that they have access to, right? So we're able to persist these agent runs, these, these decision traces in the graph,
- 5:38
and then share that with other agents in our organization, right? So that we have one shared context graph, uh, that enables our agents to essentially learn from, uh, each other in this shared memory piece. Okay. So that, that, that's the memory piece. Like, th- this, this is, um- Uh, a system that we've built at, at Neo4j, a pattern that, that lots of folks are, are, are following for working, uh, with agent memory. But what about this idea of, of actionable knowledge, right?
- 6:08
How do we enable our agents to, to take action? Uh, memory, one way to think of memory is, is this is a good representation of what happened, right? We know, um, the people we were talking about, how they're connected. We have the- these decision traces. Um, the next step is typically creating skills, right? How many people are, are using skills w- with their agents today? Mo- most folks? Cool. H- how many people have, have written skills yourself?
- 6:39
Cool. Mo- most folks. How many folks have had an agent write the skill for you, right? Cool. Yeah, so that, that's what we're, we're talking about here, is going from this memory graph, this context graph, to how can we use that to create, like, grounded, uh, executable skills? So if you're not familiar with, uh, with a skill, this is, um, an open standard Anthropic, um, put this out, uh,
- 7:09
agentskills.io, I, I think is where the, um, open standard is hosted. And the basic idea here is that we have some, uh, some, some metadata, right, some description of, like, what this skill is about, uh, and then this progressive disclosure, right? So I, I have lots of more detailed information. These are often in like markdown files, references that we can, um, progressively choose if, if we're going down this path of one piece of the skill, we can, uh, retrieve and load that data, uh,
- 7:39
into context, right? So this is, this then gives us some, like, shareable unit, uh, that we can then, uh, take the skill, package it up, uh, and reuse that across different agents. But skills have a, a similar challenge to some of the issues that we saw, uh, previously when, when we were talking about, uh, memory, right, is that we're, we're typically working with prose. Um, and so we're still often
- 8:09
limited in, in this, uh, challenge of understanding, is this the canonical thing that we're talking about, making sure that we have, like, debuggable steps and, and output. Um, and so to address some of these challenges that, uh, we're seeing with skills, uh, some folks on the Neo4j research team, um, have done some interesting research and published some work on AIP, uh, graph representation for learning and governing agent skills. Um, this is a screenshot
- 8:39
from the, the paper. Zach is, uh, is here somewhere. If he's not here, he's, he's at, um, the Neo4j booth today. So definitely, uh, chat with Zach if you're interested in this. But you can think of this, uh, AIP as essentially an extension of the agent skills, um, protocol with some additional metadata that now treats your skills as typed execution graphs, right? So we have, uh, steps modeled as nodes. We have a very,
- 9:09
um, very strict schema that governs how we describe that, how these, uh, steps are actionable, right? So think a, think of this as a way of, of representing, um, a skill as a graph broken up into steps that are executable with a schema-governed description. Um, that's the, the AIP protocol, and this is from the, uh, the benchmark that was used in the, the
- 9:39
paper. Found that, yeah, like, this, this actually matters. When, when they applied this to, um, SkillsBench, saw like a significant increase in, uh, tasks successfully completed when looking at applying, uh, the AIP protocol to human, uh, curated, uh, skills.
- 9:59
So that looks interesting. How can we leverage some of those, uh, some of those ideas, some of that research for skill distillation? Uh, so essentially what, what we want to do is take this, this memory graph, right, this context graph that's maybe scoped to a, a workspace or scoped to a, a project in an organization, and we want to distill that into a skill, but not just like a markdown file. We want this to be grounded
- 10:29
in actual data that we've observed. We want this to be like deterministic, right? We want, uh, well-understood procedural steps in our graph, and we want to be able to govern this over time, right? If, if the underlying data that makes up our skill changes in memory, we want to be able to understand that and, and know about that. Um, grounding is, is, is an important piece here, right? Making sure that the, the data that makes up our skill is, is grounded in our memory system, but
- 10:59
that, that's, that's not enough for it to be useful, right? There are other heuristics that we need to look at here. So, uh, things like coverage. Are we making sure that the, uh, steps and descriptions of our, um, throughout our skill are grounded across the skill? Coherence. Coherence is interesting. This is a way that we can detect, uh, maybe if the, um, information, the piece, the subgraph going into the skill
- 11:28
distillation is split across multiple topics, and we can suggest, well, you may want to create multiple skills here, um, and so on. So those are, are some of the pieces that lead up to generating the skill. Uh, skill governance is an important piece that I mentioned, making sure that we can understand, uh, as that skill changes, uh, as the data changes, are, are we able to update the skill? And with dynamic loading, you can think of, uh, having sort of like a
- 11:58
governed skills registry that allows us to retrieve, uh, for any agent, the, the most recent up-to-date skill. Understand, uh, if that skill may have been, uh, stale or, or, or drifted, right?
- 12:15
Cool. So we've implemented this in, uh, the Neo4j Agent Memory Service, or we, we call it NAMS. This is, um, an agent memory as a service, as part of our Neo4j Labs efforts. Um, we also have open source tooling ar- around that, that, uh, implements these patterns. And essentially, the way this works is we, we have a, a workspace-scoped context graph, right? That has our three types of agent memory, short-term, long-term, reasoning memory. Um, we
- 12:45
have, uh, background workers that are capable of going through this distillation process. We'll, we'll take a look at, at what this looks like in a minute. Um, and then having some background curation process that is still making sure to take care of that governance to surface when, uh, our skills become stale based on the data that we have in the memory system. Um, this is the, the pipeline that, that we go through t- to generate these. Um, the... I won't go through this in, in
- 13:14
detail. The, the one thing that we're trying to point out here is that most of these steps are deterministic. Um, we're really only leveraging an LLM here, uh, for synthesizing some of the claims, uh, generating some of the, the text that we use for some of the skill description. Uh, and then the other thing I wanna call out here is the, the first piece is really identifying the scope. Like, what, where, what is your starting scope for distilling one of these skills? Is it around a certain
- 13:44
entity, the subgraph around that? Is it a specific conversation, uh, and that sort of thing. And so we can decompose, uh, the, the way we represent the skill and make sure that each of these components is grounded, again, in underlying data, and that's the piece that we're looking at for if that skill, uh, essentially becomes stale, if any of that underlying data changes, is removed, or, uh, becomes, uh, contradictory.
- 14:16
Um, cool, and then we can also, as I mentioned before, we can also compose skills, right? So this is where that coherence piece comes in. We use, um, graph algorithms like community detection, right? So if we're mixing, uh, topics from multiple communities, that can be an indication that we need to decompose our skill, and we'll catch that in, um, the skill governance piece. Cool. So I've got a few minutes left here. Let's see what this looks like. Um, so this is NAMS, the Neo4j Memory Service. Um, this is,
- 14:46
this is, uh, free. Currently, anyone can, can sign in and, and try this out. Um, this is what the, the dashboard looks like. The, the basic idea here is that we have, um, REST API and MCP tools that we can expose to our agents that map to, um, ingesting, retrieving, working with long-term, short-term reasoning memory. Um, we're going through this entity extraction and resolution
- 15:16
process, so I can, uh, look at some graph representation of my, uh, agent memory here. I, I, I can traverse that and, and so on. Um, we can look at entities that have been flagged. Um, one important piece here is this idea of an ontology. So here, we're using a healthcare ontology, so we- we're going to be working with data about, you know, encounters, providers, uh, that sort of thing.
- 15:46
Uh, and then I have ingested a bunch of conversations here. So we can see the conversations that have been, um, ingested related to, uh, a healthcare agent, right? And so now we're ready to distill a skill. Um, and we said the, the first piece is to decide, like, what is the, the scope of that s- of that skill. We can do this for, you know, an entire workspace. That's often not what the case we want. We can do this
- 16:15
around, um, you know, a certain entity, specific conversations, a, a class in the ontology. Let's do this around, um, our most recent conversation, and we're gonna see, this is gonna kick off, um, in the, the distillation queue and go through and fetch the data, go through that seven-stage pipeline and, um, construct the skill as a graph and then package that up for us. Um, here's one... While this is running,
- 16:46
let's just take a look at this guy. So here's one, uh, that we ran previously. This is, uh, was run on a patient intake, uh, conversation, and as you can see here, that we've essentially extracted out the steps that make up the skill to run from a patient intake, um, all the way through, uh, charting for the patient. But we can see each one of these steps is grounded in the actual tool calls and the
- 17:16
entities that constructed the underlying components of the skill, and we package this up with SKILL, um, .md file. Uh, if we downloaded this, we would... This would be packaged up with, uh, other references and, and, and so on, uh, following that progressive disclosure, uh, standard that we use with agent skills. Cool. So that was a, a quick look at
- 17:46
kind of how we think about agent memory as part of this context graph, uh, with Neo4j, and I'll leave up some, uh, resources. You can ca- grab the slides here. There's a link, uh, to the slides and a QR code. Um, the Neo4j Memory Service that I mentioned, um, is listed here as well as lots of documentation, uh, and resources for some of our open source tooling. So that's it. I'm out of time, but we
- 18:16
have a Neo4j booth, uh, so I will be there as well as lots of other folks from the Neo4j team. So we'll see you there. Thanks, folks.