AI Engineer World's Fair 2026

Giving AI Agents Memory That Learns — Jake Broekhuizen, LangChain

Read the talk

Giving AI Agents Memory That Learns

Jake Broekhuizen explains how a financial assistant’s tone failures can become instructions for its next run—and why selecting lessons, refreshing context and reviewing rule changes matter as much as storing them.

From a talk by Jake Broekhuizen

At a glance

Ideas worth remembering

  • A trace becomes memory when a selected lesson is written into persistent context that future runs actually read.

  • Behavioral corrections such as advisory tone belong in procedural memory: rules and skills that guide how the agent engages with users.

  • Keep most execution evidence as history or evaluation data; promote only the subset useful for future behavior.

  • Memory updates need a working read path. Caching and retained runtime state can leave an agent using stale guidance.

  • Surface consequential procedural changes for human review before loading them into execution, and use evaluations to protect important behavior.

A budgeting assistant starts telling users what to do

A financial assistant can offer sensible budgeting advice and still fail its product requirements. Jake Broekhuizen, who leads LangChain’s labs team, opens with a company building an agent for spending, budgeting and non-investment advice. Its tone needs to remain advisory and helpful. Yet some interactions become pointed and directive: the assistant starts telling users to cancel subscriptions and move money into savings.

Selected presentation frame from Giving AI Agents Memory That Learns — Jake Broekhuizen, LangChain at 120 secondsOpen full source frame
Jake Broekhuizen presents on the World's Fair stage.

The distinction concerns the relationship the assistant establishes with the user. Offering a spending cap as an option leaves room for the user’s judgment; directing the user to change their finances crosses the company’s tone guidelines. Those guidelines live in the context, rules and policies the agent references. Correcting the response therefore requires changing what guides the agent’s behavior.

Initially, that correction is manual. Someone reads the traces, identifies where the tone slipped, edits the relevant context and checks for regressions. Each failure can teach a useful lesson, but the work of carrying that lesson into future runs is hard to scale. The question for the rest of the talk is how the system around the agent can perform that learning loop.

0:170:47
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:17 · section reference included

A stored trace does not change the next run

Observability makes an agent’s execution inspectable. A trace can show the tools it called, the artifacts it referenced and the recorded steps leading to its decision. That information helps explain a bad result. But if the next run receives the same misguided instruction, storing the trace gives the agent no reason to behave differently.

Memory, in this account, is persistent context that guides future runs. Logs, transcripts and traces are evidence of an experience. They become memory when a lesson from that experience is converted into context the agent can actually reference later. The consequential step is the conversion: a record of a tone failure must become guidance about how to respond.

Selected presentation frame from Giving AI Agents Memory That Learns — Jake Broekhuizen, LangChain at 241 secondsOpen full source frame
A slide contrasts “Traces as logs” with “Traces as memory.”
2:433:03
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

2:43 · section reference included

Facts, experiences and rules change behavior in different ways

A taxonomy borrowed from cognitive science separates three kinds of agent memory:

  • Semantic memory: what the agent knows. Facts and preferences provide information that guides a response.
  • Episodic memory: what the agent has experienced. Past interactions, examples and learned patterns preserve useful experience.
  • Procedural memory: how the agent should behave. Instructions, skills and rules guide its actions.

For the financial assistant, the useful correction belongs in procedural memory. A new financial fact would not address the tone problem, and a banned-word list would miss the broader behavior. The agent needs a rule governing how it engages with users: keep the advice measured and avoid directing their decisions. Broekhuizen identifies procedural updates as the source of most visible gains in this example.

Selected presentation frame from Giving AI Agents Memory That Learns — Jake Broekhuizen, LangChain at 362 secondsOpen full source frame
A slide highlights Skills and Instructions as procedural memory.
4:355:07
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

5:07 · section reference included

Separate the current workspace from what survives the run

Memory type answers what the context contains. A second distinction answers when it is available. Working memory is the context used in the current run: intermediate scratchpads, tool results and retrieved files. Long-term memory holds instructions, skills and other context that remains available across future turns and runs.

Selected presentation frame from Giving AI Agents Memory That Learns — Jake Broekhuizen, LangChain at 392 secondsOpen full source frame
A slide presents Working Memory and Long-Term Memory in separate panels.

The agent harness determines how long-term context enters the current task. It may inject context into the prompt, retrieve it through tools, read it from a store, load files or change runtime state. These approaches differ in implementation, but they all need to connect persistent information to the agent’s current working context.

A trace can therefore be read as a record of how working memory changed during execution, rather than merely a record of the final answer. That history contains more than the agent should carry forward. Some details explain this particular attempt; others reveal a lesson useful for later attempts. Selecting between those two uses is the core design problem.

6:156:45
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

6:15 · section reference included

Read context, run the agent, filter evidence, write a lesson

The connection between working and long-term memory forms a read/write cycle. At the beginning of a run, the agent reads relevant skills and instructions into its working context. Progressive disclosure—loading the skills relevant to a particular task—is one way to perform that read. The agent then executes and produces a trace of retrievals, tool calls, decisions and any sub-agent activity.

The filter step decides which parts of that evidence deserve to influence future behavior. Copying everything into long-term memory would burden later tasks with material they do not need and make reasoning harder. The selected signal is written back as persistent context. For the financial assistant, that signal becomes a rule to remain measured rather than pointed or directive.

Selected presentation frame from Giving AI Agents Memory That Learns — Jake Broekhuizen, LangChain at 574 secondsOpen full source frame
A circular diagram shows long-term memory, a run, a trace, and a filter in a read/write cycle.

Where does one run affect the next, and where does ordinary history leave the loop? The diagram follows the financial assistant’s tone correction through the read/write cycle. The return path carries a selected behavioral lesson; the history branch keeps the remaining evidence available without making it part of every future task.

This also explains the role of background work sometimes called sleep-time compute or dreaming. Memory processing can happen while the agent is working or while it is idle. The simpler implementation model is to capture what happened, analyze what matters and update the context future runs will use. Those three responsibilities remain useful even as frameworks differ in how they perform them.

How it fits togetherA tone failure becomes guidance for the next run

Instructions and skills, including the rule to keep financial advice measured.

Only selected signal returns to long-term context. The rest can remain referenceable history.

8:018:31
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

8:01 · section reference included

Context Hub carries the correction into the financial assistant

LangSmith supplies one concrete implementation of this loop. Its observability captures the execution trajectory, including tool calls, sub-agent calls and retrieved context. A background analysis process extracts signal from traces according to the behaviors the team cares about. LangSmith’s Context Hub provides the remote store where updated context can be kept for the agent to retrieve on its next run.

The financial assistant begins by loading instructions, skills, relevant Markdown files, policies and rules from Context Hub. Its run produces a trace that can expose failed tool calls, user corrections or, in this case, a tone violation. The analysis engine looks for patterns and identifies which skill Markdown files need changing. It then updates those files so the correction is available to subsequent runs.

Selected presentation frame from Giving AI Agents Memory That Learns — Jake Broekhuizen, LangChain at 725 secondsOpen full source frame
A slide titled “One Implementation” diagrams an agent, traces, and LangSmith Engine.

The observable change comes after the revised context is loaded: Broekhuizen reports that the assistant’s tone changes in later user interactions. The causal sequence is specific—a directive response supplies evidence, analysis turns that evidence into a procedural correction, and the next run reads the correction. This is a reported behavioral improvement; the recording does not quantify its size or explain the pattern detector well enough to establish a general success rate.

10:4411:14
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

10:44 · section reference included

Keep the loop selective, fresh and reviewable

The first operational lesson is restraint. Agents produce substantial execution “exhaust,” and most of it should remain referenceable history. Some traces can become datasets for offline evaluations. Only a small subset should change memory. These are different destinations for the same evidence: preserving an attempt for inspection or testing does not require promoting it into instructions.

The second lesson is about freshness. A memory store can contain the right update while a long-running background agent continues using old context. Cached data or retained runtime state may prevent the change from reaching later work. Broekhuizen describes struggling to understand why behavior was not changing before recognizing this access problem. The practical decision is what can be cached while still ensuring that future runs receive memory updates.

The third lesson is to distinguish what may update itself from what deserves human review. Procedural memory contains instructions, policies and tone guidelines, so changing it can affect a large share of the agent’s behavior. Broekhuizen estimates about 75% for his agent; that is his estimate rather than a measured general proportion. Proposed changes to these rules should be surfaced for review before they are committed and loaded into active execution, with evaluations protecting important behavior.

Selected presentation frame from Giving AI Agents Memory That Learns — Jake Broekhuizen, LangChain at 906 secondsOpen full source frame
A slide lists “Not everything should be a memory update,” “Make sure future runs actually read the update,” and “Protect important behavior with evals.”

These lessons qualify the opening ambition that the next run should be better. Storing more information does not achieve it by itself. Experience must produce a useful correction, that correction must reach the next run, and consequential rule changes need review and regression checks. Memory gives the next run a chance to improve by changing the context from which it acts.

12:5913:29
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

12:59 · section reference included

Resources

Read the complete timestamped transcript
  1. 0:12

    Testing.

  2. 0:15

    Awesome.

  3. 0:17

    Good morning, everybody. Uh, I know it's the last day of the conference, so super excited to have you here. Hopefully, uh, hopefully I can keep you engaged for the next, uh, fifteen to eighteen minutes or so. My name's Jake, and I lead our labs team here at LangChain. And today, I wanna talk to you about a topic that I have, uh, been speaking a lot about with, with our team and, and with a lot of teams that I speak with who are building agents, and that's how to give your agents memory. Whether you're building a coding agent, a support agent, a,

  4. 0:47

    uh, a deep research agent, most teams arrive at the general same instinct that it should get better from one run to the next. So let me give you an example to hold onto for the rest of this talk today. Recently, my, my-- the, the labs team at LangChain have been working with a company in the financial services sector who's building a, uh, spending, budgeting, and non-investment advice agent. Given the highly regulated nature of the financial services sector, obviously, you would imagine that it would need to follow some

  5. 1:17

    strict tone guidelines. And so the way that it speaks and, and the-- and its vocabulary, and the way that it interacts with different users is obviously the, the tell on whether or not it's following and abiding by the correct tone. It will derive the way that it generates this tone from the context and from the memories and from the references that it has to its rules and policies. And you can imagine, in some contexts, and actually in some contexts that we experience, its

  6. 1:47

    tone would slip, and so it would move from being advisory and helpful to actually pointed and directive, saying things like, "Users in your similar situation would generally set a spending cap." And it would actually say, "Oh, well, actually, it looks like you need to cancel some subscriptions and move some money into savings." The company that we're working with thought that this is a violation of tone, and there were similar instances of this too. And so to correct that, the review process is manual. You have to have someone understand the traces and understand where its tone slipped, make those

  7. 2:17

    changes to the context that it needs to reference when it's giving its tone, and then make sure that there are no regressions as well. Obviously, that, that doesn't scale, or rather it's challenging to scale. And so the system should be-- or, or the system around the agent should be a place where it can continuously learn and benefit from its interactions and its experiences, and that's the version of memory that I wanna talk about today.

  8. 2:43

    Together, I wanna start with the tension that I think a lot of teams have run into. I think all of us here today, the reason why we're at this conference is we're curious, we're building agents. I think we've gotten very good at observing them. But observing them and then turning those observations and those agent experiences into lessons for future runs are obviously very different things.

  9. 3:03

    And so that's where memory... And so that's where memory comes into play. Agents are producing more signal and more traces than ever before, and this is gonna be something that's obviously-- that will obviously increase as more agents are built. That's fantastic for us because every run produces a trace, and that trace is stored somewhere. That's incredibly useful. It means that we can understand the tools that were called, the reasoning that the agent made as it was on its way to making its decision, the artifacts that it referenced.

  10. 3:34

    Observability is obviously a critical foundation for building very reliable a-- very reliable agents. But I think the challenge really is after the trace is stored or after the log or after the transcript of what the agent did is stored. If the future context that that agent references never changes, then a misguided skill or instruction will mean that the agent-- the mistake that the agent makes today will still persist in future runs. The goal or the system that we wanna work towards

  11. 4:05

    is where a trace stops just being a log and becomes a source of signal, memory. Agents that are observable tell you what happened. Agents that are adaptive are able to change, are able to change and guide themselves on the following runs. So I've mentioned memory. I think it'll really help to sort of define what I mean when I say memory. It's the durable context that an agent references that will guide future runs. A trace, a transcript, or a log is

  12. 4:35

    evidence of what happened. It only becomes memory when that lesson is converted into durable context that the agent can reference on future runs. Now, you'll see six things on this slide behind me: facts, preferences, patterns, previous examples, skills, and instructions. I think that it's helpful to group them into three distinct buckets, a taxonomy that we borrow from cognitive science and actually that applies a lot when we're talking about agents that are language-based.

  13. 5:07

    First, there's what the agent knows. This is often referred to as semantic memory. This includes things like the facts and the preferences that guide the way that it responds.

  14. 5:19

    Whoops. Then there's the, then there's what the agent has experienced. And this is the episodic memory. This is the learned patterns that it's seen, past interactions and examples.

  15. 5:31

    And finally, there's how the agent should behave. This is also known as procedural memory, and that's the instructions, the skills that you're all probably very familiar with, and the rules that guide how it should, how it should behave. Procedural memory is where actually most of the visible gains come from. So in that financial assistant that I, uh, the example that I gave at the beginning, tone changes that we wanted to, to promote to the agent and, and make it follow came in the form of updating its procedural memory. We weren't giving it a new fact,

  16. 6:02

    or we weren't telling it not to use a banned word. We were giving it a rule to follow- As it engages with its users. That's the procedural memory that guides our agent.

  17. 6:15

    Another distinction that I, that I wanna make, and I think this is very important when thinking about the general agent, uh, agent memory paradigm, is the separation between working memory or short-term memory, and long-term memory. Working memory is what the agent is referencing in its current run, so these might be things like intermediate scratch pads, it might be tool results, it might be the files retrieved as part of its operation. Everything that it needs to do to perform its current job. Long-term memory is what

  18. 6:45

    exists typically in some data structure or some store somewhere else that has instructions and skills and things that it needs to reference as part of, uh, future turns and, and continuous behavior. Different agent harnesses will leverage and inject long-term memory in different ways, whether you're building your own SDK or, or using a variety of different popular agent frameworks. Some inject it directly into the prompt, some will use tools to retrieve that, some will in- some will pull it directly from

  19. 7:15

    a store, some load it through files, or some change the runtime state. The imple- the implementation can vary, but the distinction's really important. Some of the memory or some of the context that an agent references is temporary, while some of it is durable and long-term and is useful to guide future runs. And I think that distinction's really important when we think about traces and logs and artifacts that the agent produces as it runs. A trace isn't just a record of the final

  20. 7:45

    output. It's a record of how our working memory changed over the course of the run. Some part of that should stay as history. Some part of that should become durable long-term context.

  21. 8:01

    And so that's what gives... Or, or that's what lends itself to this really helpful mental model that I have, or that I and, and the team at LangChain, and generally I think, uh, is, is, is a very useful, uh, way to sort of think about the relationship between working or short-term memory and long-term memory. It's a read/write cycle. And so you might have heard of sleep time compute or dreaming. Those terms refer to the way that memory is updated or pulled from long-term into short-term whilst the agent is

  22. 8:31

    doing work, or actually in some cases, whilst the agent is also not doing work. That's the sleep time compute. If we start somewhere, at the beginning of an agent run, it needs to read relevant context from long-term memory into its short-term memory. So this is anything re- rel- this is anything important for the current run that it has. It might be skills, it might be instructions. Uh, you've probably seen progressive disclosure with agents that actually pull in the relevant skills for a certain task. That's that phase, moving from long-term memory into short-term memory. Then when the

  23. 9:00

    agent runs, it produces a trace. This evidence is really important to us because now we can understand how it retrieved context, the, the, uh, the tool calls that it made or the decisions or the sub-agents, or anything that's part of its working job. As I mentioned, most evidence should stay as history. If we turned everything that was from a tr-- if we turned all of that evidence into long-term memory, then it would make the agent struggle and it would make it harder to reason over future tasks. And so what I think is really interesting is the filter step.

  24. 9:31

    How do you si-- how do you decide what is useful signal and what should become durable long-term context, and how do you decide what should stay as history? Once you've made that dec-- and we'll get into that in a moment. Once you've made that decision, that is, that signal is then written back into durable long-term context. So in the case of our financial assistant, that was be measured and not pointed or not directive so that it can update the way that it behaves on future turns. And that's the core loop. Memory informs the run,

  25. 10:01

    the run makes the evidence, we filter that evidence to make signal, and then we use that signal to update memory. If you can build that flywheel, then we trend towards a place where we can build very reliable and very powerble-- powerful agents that just feel like they get better with time. And so once we have that m-mental model, I think, you know, an even more simple way to break this down or to, to abstract this away is, is capturing what happened, analyzing what might matter, and then updating the

  26. 10:31

    context so that future runs can use this. The implementations around how different, how, how different people building different frameworks are still evolving, but this three-part shape is a really useful starting point.

  27. 10:44

    One implementation of this, and I won't, I, I won't claim it's the only implementation, but just to, just to kind of like put some concrete structure to what I'm talking about, is the way that we do this with LangSmith. The capture step to be able to understand and see the trajectory and the things that the agent did is where our observability comes in. You're able to see the tool calls, you're, you're able to see the way that the agent's called sub-agents, you're, you're able to see the context that was retrieved. The analyze step in the middle is something new that I'm really excited about. But

  28. 11:14

    basically, it is that intelligent process that's doing the background analysis and doing the background work to extract signal from traces based on the things that matter to you, and then promoting that signal to the memory store or the point of reference that the agent will use for future runs, and that is LangSmith's context hub. So you can imagine now if there is a place where this analysis can be done, signal can be extra-extracted, a remote store where you can now go and update what's in that so that the agent

  29. 11:44

    can pull that down the next time it runs, that is what LangSmith context hub does. And the abstract loop, it, it maps cleanly. You capture the experience, you extract the signal, and then you store that durable context for future runs.

  30. 11:58

    To kind of put that in some more detail and, and also to frame it in the, uh, i-in the, the context of the financial assistant that I was speaking about at the beginning and, and sort of the way that we, we built this loop- When that agent runs, its context is loaded from the context hub, and so that might include its instructions, its skills, relevant markdown files that matter to it, policies, and rules. When it functions or when we invoke that agent, then a, a trace is produced. That shows us the, the, uh, failed tool calls

  31. 12:28

    or user corrections, or in this case, it flagged where our tone strayed from what we wanted it to be. Engine then looks for patterns and then identifies where different static skill markdown files that are based in context hub need to be changed or promoted, and it'll make that change.

  32. 12:47

    And then the next time our agent runs, we're able to actually go and, and understand and see that its tone changes when it interacts with, engages with users.

  33. 12:59

    So some things that I and my team have currently, uh, or, or kind of always think about when, when designing memory, and, and this should also as well ideally guide the, the way that you design and build some of these abstractions, is that agents produce so much-- we call it exhaust or we call it kind of like feedback. Not everything should become a memory update, and actually the devil is in the details around what should become one, as, as we've sort of discussed over the course of this presentation. It's understanding and deciding what matters to you. You know, most

  34. 13:29

    trace data actually should stay as history, as referenceable history. Some, some might become datasets to do offline evaluations and things like that. A small, a small segment of that or, or subset should end up becoming what changes your agent's memory. Now, this second one is sort of like a gotcha that I, I ran into a lot, and I think that you sort of tread the fine line of trying to build a system that's optimized, but then also one that has a memory store or updates available to it in the hot path. And what I found was,

  35. 14:00

    uh, for-- particularly for long-running background agents, right? Imagine if you have an agent that's running over long time horizons, and you make some memory updates in the beginning of its operation, but because of the runtime state and because of the way that it's accessing memory, you don't actually access that or that's not available to it for future runs. And so, you know, understanding what you can cache, what you can't cache to be able to make sure that memory updates are available for future runs. Very important, uh, and sort of-- I've-- I was banging my head on a wall for a while, while, whilst trying to sort of understand

  36. 14:30

    why my behavior wasn't changing, and I needed to recognize and think about that second principle. And then the third one as well. I think that, uh, OpenClore and, uh, you know, Hermes Agent and, and a lot of agents that are, uh, becoming widely popular that sort of do this self-update process, right? I think that it's another really important consideration of like what do you make self-updatable as opposed to what do you make something that needs human review? And so if we're talking about procedural memory with instructions and policies and tone guidelines,

  37. 15:00

    those are probably gonna drive about seventy-five percent of the behavior of my agent. And so that actually probably should have some level of human in the loop and human interaction. You should have some way to surface it, the changes that you make to the parts of the agent's memory that are procedural before actually committing those and, and putting those in the, in the, in the agent's sort of like hot reload path. So protecting important behavior with evals is, I guess, another really important, uh, third, third kind of like level to this that I find really important.

  38. 15:31

    And so I wanna come back to this initial refrain. The next run should be better. Memory allows us to do this. And the, the word memory is thrown around a lot. I think in the context of building agents and working with harnesses and working with like providing context to your agent, memory really does give us the ability to do this, and it's more than just a place to store information. It's how experience becomes context, and then that context is how our next runs get the chance to improve. If you're curious

  39. 16:01

    about an actual implementation of this, by the way, that QR code has a video of me walking through how I've built an agent that does exactly what I've just described. Uh, hopefully that'll kind of like make it tangible and, and will sort of connect some of those dots. Uh, and if you wanna have any further discussions about this, please come and see me. We're at Booth UG nineteen. Uh, I'd love to-