AI Engineer Code 2025

Stop Fine-Tuning to Fix Retrieval Problems — Anant Srivastava

Read the talk

Stop Fine-Tuning to Fix Retrieval Problems

Selected presentation frame from Stop Fine-Tuning to Fix Retrieval Problems — Anant Srivastava at 69 secondsOpen full source frame
The slide links files, databases and APIs to the model under the heading “The interesting engineering is not in the model.”

Anant Srivastava explains how to assign behavior to prompts, changing knowledge to memory, and stable task reflexes to weights—and how to move information among them as an agent learns from its work.

From a talk by Anant Srivastava

At a glance

Ideas worth remembering

  • Prompt edits, indexed documents and training samples each decide where behavior or knowledge lives. Review their combined effects as architecture.

  • Use authored prompts for small, stable behavior; use memory for changing, large, citable or access-controlled knowledge.

  • When an assistant lacks the right document chunks, fix retrieval. Fine-tuning on runbooks can encode stale facts while preserving the original retrieval failure.

  • Fine-tune stable task reflexes when human judgments have converged and capability or cost justifies the change. Keep humans at contested edges and monitor drift.

  • Learning a recurring format changes future retrieval needs: format examples may become unnecessary even while factual knowledge remains in memory.

Every prompt edit decides where knowledge lives

Enterprise knowledge lives in files, databases and APIs. A model consumes tokens and produces tokens; the system around it decides which knowledge reaches inference and how. Anant Srivastava opens with that architectural question: should something arrive through authored instructions, through retrieval and memory, or through a change to model weights?

Teams often answer by escalation. An answer is wrong, so someone edits the prompt. If that fails, someone investigates retrieval. If answers remain wrong, fine-tuning becomes the next proposed fix. This treats three tools as successive levels of intervention, even though each changes a different part of the system.

A quieter mechanism produces the same architectural drift during ordinary product work:

  • Prompt edits decide which behavior lives in instructions.
  • Indexed documents decide which knowledge lives in retrievable memory.
  • Training samples decide what the model absorbs into its weights.

These decisions happen even when nobody calls them architecture. Their effects accumulate across teams.

0:130:18
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:13 · section reference included

The catalog changes, but the old product names survive

The support assistant example follows several individually plausible changes. After launch, a product manager adjusts its tone through the prompt. The support team adds refund-policy documents to retrieval. When answers still disappoint, an engineer inserts the product catalog into the prompt. Later, the machine-learning team fine-tunes the model on support tickets covering up to six months of history.

Then a new catalog launches. Updating the catalog in the prompt looks like the obvious maintenance step, yet the assistant continues returning product names that no longer exist. Srivastava attributes this failure to the historical tickets: fine-tuning absorbed product information alongside the support behavior those tickets demonstrated. Replacing the prompt’s catalog changes one source of information while leaving the learned associations in the weights.

Where did the obsolete names survive the update? The diagram follows the two paths by which catalog knowledge entered the assistant. The prompt path receives a replacement; the training path retains historical product information. That split explains why updating the visible catalog can leave the underlying failure intact.

Six months of normal product work created a system that nobody owned as a whole. Everyone owned a piece. Srivastava includes himself among engineers who have made these decisions without enough thought; the useful correction is to ask where each piece belongs before adding it.

How it fits togetherTwo catalog paths, only one updated

Launched after the assistant had accumulated several changes.

Replacing the prompt catalog leaves product information learned from historical support tickets in the weights.

2:383:03
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

2:38 · section reference included

Prompt: small, stable instructions about behavior

The prompt’s job is to specify tone, persona and behavior through instructions that are small, stable and editable. For the SaaS support agent, that means maintaining a professional tone, offering human escalation after three failures, and giving concrete next steps. Those requirements remain useful across users and queries.

Selected presentation frame from Stop Fine-Tuning to Fix Retrieval Problems — Anant Srivastava at 380 secondsOpen full source frame
The slide “A customer support agent’s persona” displays a prompt and a separate “Where it lives” note.

A product catalog has different maintenance and inference costs. Carrying it in the prompt supplies factual context even when much of it is irrelevant to the question, and the additional tokens cost money. Srivastava also warns about the “lost in the middle” problem with long context. His diagnostic is whether the material is small, stable and concerned with how to behave.

Here, prompt includes authored instructions wherever they originate. A system prompt and instructions loaded from Markdown files serve the same architectural role. The important distinction is what they tell the model to do, rather than which file or message holds them.

5:156:01
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

5:15 · section reference included

Memory: retrieve current knowledge with the right scope

Memory grounds an agent in knowledge that is current, large or citable. Current knowledge changes faster than it is practical to fine-tune. Large collections exceed what belongs in a standing prompt. Citable knowledge must remain connected to a source that the production system can identify.

Selected presentation frame from Stop Fine-Tuning to Fix Retrieval Problems — Anant Srivastava at 449 secondsOpen full source frame
The slide lists Current, Large and Citable under “Right When,” beside warnings about using memory to fix behavior or reasoning.

Memory covers two arrangements in this talk. An agent can acquire information about the person interacting with it. An external application can also maintain enterprise knowledge—such as refund policies—and retrieve it through retrieval-augmented generation, or RAG. The agent need not manage the underlying collection for that collection to serve as memory.

Retrieval supplies material to reason over; the model still needs the ability to use it. Srivastava’s warning is concrete: if a model cannot reason across two, three or five documents, supplying fifty does not solve that capability problem. More context can add confusion when the missing ingredient is reasoning.

Access control adds another reason to keep knowledge in memory. If one user must not see another person’s information, retrieval must control which information reaches that user’s context. The diagnostic therefore includes scope: does this knowledge belong to a particular user, repository or other permitted collection?

An organizational code assistant makes these requirements practical. Repositories change frequently, so fine-tuning their contents creates a maintenance problem. Inserting the entire codebase into a standing prompt also carries irrelevant material into inference. Relevant code can enter the context window when a task needs it; the collection itself belongs in retrievable memory.

Two retrieval mechanisms keep that collection useful:

  • Code-aware chunking. Chunk code with its structure in mind, potentially using an abstract syntax tree (AST), rather than treating it only as undifferentiated text.
  • Metadata and filtering. Attach repository identity and relevant user-permission information to chunks. Keeping that metadata alongside the chunk makes it available when selecting the material for a request.

The aim is to return the code relevant to the task and its scope. Returning many loosely selected chunks produces what Srivastava calls “RAG mush”: the database supplies volume, and the model inherits the confusion.

6:587:28
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

6:58 · section reference included

Weights: avoid freezing knowledge that still changes

Rate of change is the starting diagnostic for weights. Fine-tuning can encode a stable pattern, but stability alone does not justify training. There must be a strong reason to make the change. Encoding a pattern too early risks freezing a decision boundary while the task is still evolving.

The subtle mistake is an internal document assistant that answers poorly, prompting the team to fine-tune on its documentation, runbooks or procedures to teach the domain. In Srivastava’s example, the missing ingredient is the right document chunks. Training on those documents gives the model an older version of the facts while leaving the retrieval defect unresolved.

The repair follows the failure’s cause: improve retrieval so the assistant receives the relevant current procedure. Runbooks belong in external memory, where the system can update and retrieve them. A wrong answer alone does not diagnose a need for fine-tuning.

Selected presentation frame from Stop Fine-Tuning to Fix Retrieval Problems — Anant Srivastava at 795 secondsOpen full source frame
The slide “Weights: specialized reasoning, reflexive format” contrasts “Fine-tune what’s crystallized” with “Freeze what’s still moving.”
11:2411:30
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

11:24 · section reference included

Fine-tune the stable center, retain humans at the contested edge

Content moderation and claims processing offer a different path. Start with a capable model making recommendations and humans correcting them. Early corrections may be inconsistent. Over time, repeated decisions can reveal a stable center of the task and a contested edge where human judgment remains necessary.

As human override rates settle, that stable center becomes a candidate for fine-tuning. Convergence is something to observe, rather than assume every task will reach. The resulting model still needs monitoring for drift: if the previously settled pattern changes, its learned response can become inappropriate. Human involvement remains at the contested edge.

Medical claims coding sharpens the distinction between facts and reflexes. Doctors’ notes are mapped to ICD-10 codes for uses such as insurance. Srivastava cites seventy thousand codes, yet recommends against fine-tuning simply to memorize the collection. Codes can change, and covering the collection through training would require substantial training data.

The training target is the repeatable task behavior: understanding the input format and making the appropriate selection across formats. The code collection remains knowledge to consult. This is the practical force of “reflexes, not facts”: fine-tuning should teach a settled operation the model must perform repeatedly.

After stability and human convergence, there are two reasons to consider the expense of fine-tuning:

  • Capability. The model needs to acquire a task behavior it cannot perform adequately.
  • Cost. A repeated, well-understood pattern may let a smaller fine-tuned model handle volume more cheaply.

Srivastava expects cost to be the more common motivation when frontier models already have sufficient capability. He gives no universal savings figure; the case depends on the pattern and workload.

The decision lens is compact: prompt for behavior, memory for facts and knowledge, weights for learned reasoning or task reflexes. It is a heuristic with exceptions. Its value is that each proposed change must answer a question about the job it will perform, rather than merely occupy the next rung after an unsuccessful fix.

Selected presentation frame from Stop Fine-Tuning to Fix Retrieval Problems — Anant Srivastava at 1037 secondsOpen full source frame
The slide “The diagnostic” shows a comparison table with columns labeled “Prompt,” “Memory” and “Weights.”
13:4714:17
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

13:47 · section reference included

The harness moves information as the agent learns

Assigning information to a place is only the beginning. A useful architecture lets information circulate. Work performed in the context window generates signals, some of which become lasting memory. A later session retrieves relevant memory into context. The harness—the surrounding system that stores information and manages these flows—makes experience available beyond a single interaction.

Selected presentation frame from Stop Fine-Tuning to Fix Retrieval Problems — Anant Srivastava at 1072 secondsOpen full source frame
The slide “The circulation” diagrams Prompt, Memory and Weights as connected circles.

A further flow turns stable experience into training material. Over time, retrieval patterns and recurring formats reveal things the model could learn reflexively. Those selected patterns move from memory into fine-tuning. This requires the earlier diagnostic: a recurring fact does not become a good training target merely because it appears often.

What changes after a format becomes a learned reflex? The cycle below separates the flow of information into context from the effect of fine-tuning on future retrieval. The return arrow from weights to memory means retrieval needs change; it does not mean the system copies the model’s weights into a database.

Return to medical coding. Before fine-tuning, the system may retrieve examples to show the model the notes-and-code format on each request. Once the model learns that format reflexively, those particular examples need not be supplied every time. The system still needs the relevant factual knowledge, but the material required to teach the format at inference has changed.

That is how an agent can improve through doing its job: experience becomes memory, settled patterns become reflexes, and those reflexes alter what future work needs to retrieve. Srivastava closes by calling the model the easy part. The substantial architectural work is the harness that puts information in the right place and manages its movement over time.

How it fits togetherExperience changes both the model and what it retrieves

Instructions and retrieved information support the current session.

Context produces signals for memory; stable formats can enter weights; learned reflexes reduce the need to retrieve format examples.

17:3217:37
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

17:32 · section reference included

Resources

Read the complete timestamped transcript
  1. 0:13

    How do you decide what to put where?

  2. 0:18

    For most enterprise AI systems, the interesting enduring problem is not the model itself. The model consumes tokens and produces tokens. Your knowledge lives in files, databases, and APIs. The key architecture decision is how that knowledge reaches inference, um, from the prompt or the context through the retrieval or memory layer,

  3. 0:49

    or by model adaptations such as fine-tuning. Today, I'm going to argue that this is a decision that most enterprise teams make by accident, and the path to recogn-- to, to making it on purpose is to recognize that this is not a ladder that you climb. These are three different tools for three different jobs.

  4. 1:20

    So how is this an accident? Well, there are two mechanisms. One is the obvious one, the other one not so much. So the obvious one is that teams escalate. You don't get the right answer from your model. The easy thing is go put something in there for the prompt, like, go fix the prompt. Still wrong? Maybe there is something with the memory system or the retrieval system. Let's take a look there. Even then, if you're

  5. 1:49

    wrong, hey, you know what? Let's fine-tune the model. Doesn't happen very often, but yes, sometimes teams reach out to fine-tune the model.

  6. 2:01

    There is a layer underneath it. The teams that are not escalating are still making these decisions every day. Every edit that you make to the prompt decides what behavior lives in the prompt. Every document that you index decides what knowledge is held in the memory, and every

  7. 2:27

    s- training sample that you bake in will... Every, every training sample that you identify will bake something into the weights. Um,

  8. 2:38

    all these are architecture decisions, none call like that. Okay? Let me make it a little concrete for you all. Um, so, and I've seen this play out actually. Uh, an internal support assistant just released. In a few weeks, the product manager makes some changes to the prompt to fix the tone of the support assistant.

  9. 3:03

    After a few weeks, the support team adds refund policy documents to your retrieval and memory systems. I'll talk about retrieval and memory. I'm teaching it. I'm, I am treating it all as memory, but I'll come to that. Um, and then since your model is not-- your, your agent is not giving you the right answers, an engineer shoves the product catalog into the prompt,

  10. 3:29

    and after a few weeks, the machine learning team decides to train the model or fine-tune the model on support tickets from six months ago, from up to six months ago.

  11. 3:43

    All decisions look fine, look okay, uh, except one, by the way. But even I have made such decisions in my, in my journey. I've made such decisions of doing such things without really giving it a lot of thought. Um,

  12. 3:59

    so when the new product, product catalog is launched,

  13. 4:04

    looks easy, right? Oh, the engineer has shoved the product catalog into the prompt. Let's put the new product catalog there. And you would think that everything is working fine, but your model or your agent is still giving you product names that don't even exist. How did that happen? Well, the product catalog leaked into the weights when you fine-tuned it on your support tickets, and that is the accident. Your model...

  14. 4:34

    Your, your architecture was accumulated. It was not designed. That is the accident. It is what six months of normal product work led to. Nobody owned it, meaning everybody owned a piece of it, okay?

  15. 4:55

    Um,

  16. 4:57

    nobody asked the diagnostic question. Nobody asked where this should th-- where this thing should live. That diagnostic-- Those diagnostic questions are the lens we are going to talk about, and the rest of, uh, the talk is, is about the lens, okay?

  17. 5:15

    So the first part is prompt. What is the job of the prompt? The job of the prompt is behavior, is tone, is the persona of the agent. Things that are small, stable, and editable.

  18. 5:31

    The wrong job for the prompt is to store facts. You don't put product catalog into the prompt because you are unnecessarily providing context to your model and paying the cost for it, and when you have too many tokens, you run into lost in the middle problem. The diagnostic for prompt is, is the small-- knowledge small, stable and about how to behave rather than what to know.

  19. 6:01

    A concrete example for a good prompt use case is a customer support agent for a product SaaS, right?

  20. 6:15

    Now, this agent, it has a professional tone. It is-- It offers human escalation after three failures, and it provides you concrete next steps. None of this changes per query per user. This is-- This, this exact thing belongs to the prompt. And one other thing is prompt is something... I'm not only talking about system prompt, I'm talking about authored

  21. 6:45

    instructions for your, for your model. So they could also come from MD files, they could also come from other things, right? But this is what I mean by prompt here.

  22. 6:58

    The second path is memory. The job of memory is to ground your model or your agent in, in knowledge that is current, large, and citable. Current meaning it changes faster than you can fine-tune. It changes every now and then. Large is because it cannot be held anywhere. It cannot be held in your prompt. And citable because production AI systems need to

  23. 7:28

    be able to point to the source of the knowledge. This is in-- This is what belongs to memory. Now, when I use the term memory, people think of agent memory, n-no, knowledge that is acquired about people, about people who's... about a person who's interacting with the agent, right? What I'm also referring to here is external memory, which is-- which you retrieve using RAG and all that, like the refund policies, right? It's not being managed

  24. 7:57

    by your agent, it is managed by an external application. That is included here. So I'm bo-talking both about the agent memory and the memory of the enterprise, okay? The wrong job of-- for memory is shoving behavior in or thinking about how to reason with the memory. Now, of course, you can have a graphical kind of RAG which will help the agent do reasoning, but that's it. That's

  25. 8:27

    where you stop. For reasoning, you've got to have the right model because if your model cannot reason over two or three or five documents, it won't be re-able to reason over fifty that you provide. The more context doesn't help if it's not able to reason. So the diagnostic or the question you need to ask is: Is the knowledge too large to fit, changing faster than I can retrain? And another important aspect is whether it has some access control, whether it is

  26. 8:57

    scope per user. You have got to have that in memory. If there is some kind of access control that this user cannot see the other person's information or the data or whatever, you've got to have that knowledge in memory, because then you can control who sees what.

  27. 9:18

    A good example for storing something in external memory or memory is a code assistant, right? Which has access to your-- all your repos as an organization.

  28. 9:33

    So, uh, would you fine-tune your model to learn the code? No. It's, it's a simple one, right? It changes very often. You wouldn't do that. You'd fine-tune your model to learn reflexes, not facts, typically. And would you stuff it in the prompt? No, you don't stuff such things into the prompt. You may s-stuff or you may add certain things that are important in the context window, right? That's something different. But you wouldn't stuff the entire code base into the prompt, okay? Or, or even a part of

  29. 10:03

    it if it has no use there. Uh, now, the other important thing is, um, when you build a RAG system, right, like for, for, for your code, for your code, you need to make sure that you are chunking properly, you are using code-aware chunking with AST perhaps. Uh, and then the other important thing is use filtering. You should denormalize-- You should perhaps denormalize things such as who's the-- what is the repo this

  30. 10:33

    code is for. You get a code chunk, you have to create metadata about which repo it belongs to. What is the user-- what are the users who are allowed to commit, and all that because that is what is going to help you get the exact data that you need versus creating a RAG mush where you are returning a lot of chunks back from the database without... and, and then confusing the model.

  31. 10:59

    So again, the diagnostic is something that's too large to fit, changing the faster than you can retrain the model, and it's scoped at a certain level. It could be, it could be a user, it could be a repo or whatever. All that knowledge belongs to memory. It could be external memory, which is your, like, your applications database, or it could be your agent memory as well.

  32. 11:24

    Now let's come to the interest-- more interesting part, which is weights.

  33. 11:30

    So what belongs to the weight?

  34. 11:36

    Something that stopped changing. That is what belongs to the weight. And the only axis that drives this slide is rate of change.

  35. 11:47

    You can lock in what has stopped changing and perhaps retrain. I'm not saying just lock in what's stopped changing and just go retrain your model. No, that's not what you do all the time. You would-- You have to have a very strong reason to retrain your model. So what you do is you lock in what's changing and then go from there. But if you get it wrong, then you could freeze a boundary that's still moving, okay? And it's not as subtle as like, "Hey, can I re-- would I re--" Nobody's gonna retrain their model on like...

  36. 12:17

    Or fine-tune, I meant, not retrain. Sorry. Fine-tune their model on pricing list. Nobody does that. It is more subtle than that, right? And a current version of that is, um, there is a-- Let's say there is an internal document assistant and the s-support team or the-- or your team starts-- decides to fine-tune the model because it's not giving the right answers. So to teach the domain, it says, "You know what? Let's fine-tune this

  37. 12:47

    model on, on our documentation. Let's fine-tune the model on our..." Or maybe not all your-- all the documentation, but runbooks and procedures. And what happens with that? There is an issue. If you do that, when a... Your model will be giving you information that's old. This was actually a retrieval problem. If the model was not giving the right answer, what the model needed was the right chunks, the right pieces of documentation.

  38. 13:17

    So basically, instead of fix-fixing the retrieval problem, the team actually went ahead and fine-tuned the model. That's not what you should do. Runbooks, procedures, these are facts. These belong to s- either your me-- either your agent memory, not there, but of course, your external memory. That is where these things should reside. Not-- If the, if the model is not- Good afternoon, everyone. Hello. Oh, sorry. That was ... I was like, "Who's that?" So if your, if your model is

  39. 13:47

    not giving the right answers- -fix the retrieval problem. Don't just go ahead and fine-tune your model, okay? Uh- And f- if you flip it, same axis. If you reverse it, um, a good use case where I could fine-tune is, let's say it's a hard ambiguous problem, such as content moderation or claims processing, right? When you begin, you can use a frontier model that makes recommendations and humans correct it,

  40. 14:17

    although inconsistently. Over a period of time, patterns will emerge. Patterns will emerge, and you will have a sort of a center which is sort of fixed and stable, okay? And then you will have the contested edge.

  41. 14:35

    The, the rates of human overrides will flatline over af-after a certain point, and you'll, you will most likely emerge in a pattern where you have a contested edge, where you still need humans, but a center which is pretty stable, okay? Now, that center is something you can fine-tune your model on, okay? But when you do that, you have to monitor for drift. Is there a drift in my center? And if that happens, then you're going to run into problem. And, um,

  42. 15:05

    actually a stable version of that is, uh, I don't know. Have you heard of medical claims coding? I think we must have. Like ICD-10 codes. Doctors' notes are mapped to like codes for different purposes, like insurance purposes and all that. So Mount Sinai- We'll be starting our talk really soon. Wow. Mount Sinai and IMO Health actually- Our partner team also needs- And the-there are seventy thousand of such codes, by the way. So would you fine-tune your model to learn those codes? They don't change so often. You won't.

  43. 15:36

    That's fact. That's knowledge. That's-- You should not be looking at fine-tuning your model. One of the obvious reasons is m-- is that they change sometimes, and the other is you will need a large test da-- training data to even fine-tune your model on it, okay? What you fine-tune your model is on the reflexes, whether, whether it is understanding the format of the input and whether it is able to... And when it-- you have to make selection across

  44. 16:06

    the formats, like which one is the right one, that's where you fine-tune your model to do that, okay?

  45. 16:15

    So the diagnostic for fine-tuning is, has it st-- has the information ch-stopped changing? Have the humans converged? And more important, you should either ha-have one of these two reasons. Is the, is... Are you fine-tuning because your model is not capable, or are you fine-tuning to save cost? And a lot of times it's not capability. Frontier models are quite capable. It usually comes down to cost, where you can have a sm-- If you have a l-nice

  46. 16:45

    pattern in your work, you can actually perhaps use a small model, fine-tune it, and run the volume cheap, right?

  47. 16:56

    So

  48. 16:59

    if you don't remember anything from the talk, this is a simple one, right? This table helps me decide. You can bo-- You can bounce your current AI system against this and, and decide. This is the lens basically that helps you decide what lives where. Of course, there can be exceptions. I'm not saying this is always the rule. There can be exceptions to this, but prompt is more about behavior, memory is about what to know and facts, and weights are about how to reason.

  49. 17:29

    Okay?

  50. 17:32

    And this is the architecture.

  51. 17:37

    Your prompt or your context window, your memory, and your fine-tuning needs to circulate. You need to build a right harness around it so that these things can circulate.

  52. 17:50

    Your prompt or your context window has information. It generates signals. Some of them becomes durable memory, okay? Now, when you are restarting a new session, things get pulled out of from the memory into the prompt or context window, and that's when you're going from the memory to the prompt. The more-- so this is-- a lot of teams have built this architecture, but what people are looking to are focusing on building is the right side, the memory-to-fine-tuning flow, okay?

  53. 18:21

    Now, what happens is when you build an agentic system over a period of time, you will have retrieval patterns. You will have formats that you think the model should reflexively know. So those things move from your memory to fine-tuning. Once you're fine-tuned, so that's m-memory to, uh, weights. Once you're fine-tuned, what is worth retrieving changes. If you have

  54. 18:51

    changed-- if you have fine-tuned your model on the formats of like notes and the ICD codes, you don't have to retrieve certain examples from your memory or external memory each time and provide it to the model because now the model knows the format r-reflexively. You don't have to provide examples to it that, "Hey, this is the format," or, "That is the format." So that's weights to memory. And so basically, this becomes a loop over a period of time, and the agent gets better

  55. 19:21

    at its job by the virtue of doing its job, okay?

  56. 19:28

    That's the architecture. So what I would like to conclude with is the model is, is the easy part.

  57. 19:38

    What you build around the model, the harness around the model that helps you store the right information at the right place and circulate among them is the key architecture that you got to build. Thank you. Thanks for coming out.