AI Engineer World's Fair 2026
Why Bigger Context Windows Won't Save Your Agent — Elizabeth Fuentes Leone, AWS
Read the talk
Why Bigger Context Windows Won’t Save Your Agent
Elizabeth Fuentes Leone shows how to manage growing agent histories in Strands Agents: keep recent conversation, store large outputs elsewhere, share pointers between agents, and give tool work explicit limits and asynchronous handles.
From a talk by Elizabeth Fuentes Leone
At a glance
Ideas worth remembering
A larger context window does not remove the need to select information: accumulated tool output can overwhelm context and leave relevant material neglected.
Conversation managers control immediate history; long-term and graph memory let the application retrieve past information and relationships separately.
Memory pointers keep large payloads available without repeatedly placing them in context, and can be shared through invocation state across agents.
Clear tool responses guide the next decision; invocation limits bound repetition; async handles separate starting slow work from collecting its result.
Compression is useful only if the retained information still supports the task. Preserve a way to retrieve underlying evidence when a summary is insufficient.
A log-watching agent accumulates more than it needs
An agent supervises an application by retrieving its logs. Each retrieval adds another batch to the conversation history. If subsequent invocations carry that history forward, the model receives both the latest logs and the accumulated results of earlier checks. The monitoring task may stay the same while its context keeps growing. Elizabeth Fuentes Leone, an AWS developer advocate, uses this example to explain why large tool outputs can overwhelm an agent.
A larger context window gives the application more room to accumulate data, but capacity alone does not ensure that the model uses all of it well. Fuentes Leone describes an attention “U-curve”: information near the beginning and end remains easier to use, while information in the middle can be neglected. This is a qualitative failure pattern in the talk, rather than a measured threshold that applies to every model or context length. The architectural question is therefore which information should reach the model on this invocation.
Context engineering means “giving the model the information it needs when it needs it.” That includes reducing unnecessary data and considering what the data contains: tool output can bring prompt injection into the context as well as consume tokens. Curating the input serves both the answer and the amount of work required to produce it.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Four ways to change what enters the context
The talk organizes context management around four complementary strategies:
- Externalize: move large data into persistent storage and retain a memory pointer that identifies it.
- Select: retrieve only the information relevant to the current task.
- Compress: summarize or compact information so it occupies fewer tokens.
- Isolate: separate context across agents, giving each agent the information its work requires.
These strategies act on different parts of the problem. Externalizing changes where data lives; selecting changes what comes back; compression changes how much detail travels; isolation changes who receives it. A log-monitoring system can combine them rather than choose a single technique for every kind of memory.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Start by managing conversation history
Strands Agents is the open-source, model-agnostic framework used for the examples. Its agent construction supplies the agentic loop behind a system prompt and a set of tools, without requiring the developer to assemble an explicit graph of nodes. The presented setup uses Amazon Bedrock; other model integrations can be supplied through model configuration.
A conversation manager is the first, relatively simple place to control growth:
- Sliding window: retain only the most recent messages. Earlier conversation falls out of the model’s immediate context.
- Summarization: condense older messages while preserving recent conversation.
- Combined management: use summaries of older material alongside retained recent messages.
The summarizing example preserves the last four messages and configures a summary ratio. The spoken explanation gives conflicting percentages—50% and then 60%—so it does not establish a consistent numeric ratio. The useful mechanism is clear: recent exchanges stay intact, while older exchanges are represented more compactly as the session continues. This gives the agent continuity without replaying every old message in full.
That convenience creates a decision about what may be lost. Keeping recent messages protects the immediate exchange, but the older summary must still carry whatever facts future answers will need. The ending returns to this limitation: a summary can be short and still omit the detail that matters.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Conversation history is only one kind of memory
Managing a session’s history does not by itself solve remembering across sessions. A new session may need something learned earlier, even though that earlier conversation is no longer present. The talk separates three forms of memory by the information they retain:
- Short-term memory: recent conversation history within the current session, managed with the conversation techniques above.
- Long-term memory: information saved for later retrieval, with a vector database offered as a storage approach. Prompts can determine which kinds of information from a conversation should be saved.
- Relationship memory: entities and their connections, represented in a graph. Fuentes Leone points to Neo4j as a place to learn more about this approach.
The three can coexist. Recent history supplies the immediate exchange; retrieval brings back relevant information from the past; graph memory retains relationships that the application wants to represent separately. Remembering something therefore need not mean carrying its entire original conversation into every future request.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Replace repeated log batches with a retrievable pointer
Return to the monitoring agent. Initially, every log retrieval enlarged the conversation. With a memory pointer, the tool retrieves the logs and stores them outside the model’s context. The agent retains an ID identifying that data. Another tool can inspect the stored logs and return a compact status, so the agent can answer a routine monitoring question without receiving the whole batch.
The observable change is in what the model receives: a pointer and a useful result replace the repeated raw log payload. If a later question requires the underlying logs, the agent passes the ID to a retrieval tool. Externalization keeps the original data available; selection determines when it returns to context. The design still depends on tools that can store, interpret and retrieve the data—the ID alone contains no diagnosis.
Where do the logs go, and what reaches the model instead? The flow below separates the large stored payload from the compact pointer and assessment. Its return path matters: the application can recover the logs when needed rather than relying entirely on an earlier summary.
In the Strands example, tools are defined with a tool decorator, and tool context exposes agent state for managing this information. The memory pointer ID remains in the context while the associated data is handled through state and storage. This makes the division between model input and tool-accessible data part of the application’s architecture.
The same division becomes useful in a multi-agent Strands swarm. Copying one agent’s full context into another gives the recipient material it may never need. Instead, the agents share a memory pointer through invocation state. They can refer to the same large dataset while keeping their individual contexts focused. Shared access to data does not require identical conversation histories.
A tool retrieves a large log batch.
A tool can assess stored logs and return a compact result. The retained ID allows later retrieval when the underlying data is needed.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Give repeated tool calls an explicit stopping point
An agent can also accumulate work by invoking the same tool again and again. An unclear tool response leaves the agent without a clear basis for deciding that the work is finished. Fuentes Leone pairs two remedies: make the tool’s response clear, and impose a maximum invocation count.
The example allows each of the illustrated tools to be invoked three times. A clear response helps the model choose the next step; the count limit bounds repetition even when that choice goes wrong. The limit prevents unlimited calls to those tools, but reaching it does not establish that the task succeeded. It supplies a stopping point for execution.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Separate starting a slow job from collecting its result
Some tool problems are about elapsed time rather than context size. An MCP tool or external API may take longer than the calling application can wait. AWS API Gateway is the example of an intermediary with a maximum response time: a slow downstream operation can outlast that limit and interrupt the request.
An asynchronous handle separates starting work from retrieving its answer. One tool starts the long job; another checks its status. The application retains the handle and can collect the result on a later invocation. The presented implementation uses FastAPI. Fuentes Leone reports improved responsiveness in her tests, without specifying a measured latency reduction; whether this pattern fits depends on whether the application can return before the result is ready.
What changes in the timing of the agent’s work? The diagram shows the long operation continuing between two invocations. The first call starts the job and preserves its handle; a later status check reconnects the agent to the result. The slow work remains, but a single request no longer has to wait for its entire duration.
The first invocation initiates external work.
Starting the job and retrieving its result are separate tool operations.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Preserve useful evidence without passing everything
The closing anti-patterns put a limit on enthusiasm for every technique above. Context stuffing includes everything because it is available, even when the answer needs only a small part. Indiscriminate summarization makes the opposite mistake: it reduces the history but may discard information needed to answer. Passing whole contexts between agents repeats the stuffing problem across the system.
The final cheat sheet maps concrete problems to concrete controls:
- Large log or database outputs: externalize the data, select relevant information, and keep a memory pointer.
- Several agents needing the same data: share the pointer through invocation state rather than copy the full payload into every context.
- Repeated tool calls: provide clear responses and limit invocation counts.
- Slow external work: use an asynchronous handle to retrieve the result later.
For the log-watching agent, this produces a practical division of responsibilities. Storage keeps the evidence, tools inspect and retrieve it, state carries references, and the model receives what it needs for the current answer. The important decision is what crosses into context on each invocation—and how the application recovers detail when a compact result is no longer enough.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Read the complete timestamped transcript
- 0:12
Hi. Well, first I have to say I am Elizabeth. I am no Morgan, but the guy for the agenda, they never change agenda, so we are almost the same but with different color of hair. Okay? So it's, it's going to be similar. So today I'm going to talk about the infinite context window is a myth. Yes. It's because right now agent system
- 0:42
breaks when tools returns a large amount of data. Context window overflow and the agent start to hallucinate and start the reasoning is degraded. So right now, context engineering help, help us to, uh, give us a better architecture for this agentic application. So we have some techniques that allows us to improve the context in-- the context window for this
- 1:12
a-agentic application. So imagine that you have a agent that is responsible to, uh, supervise some application. So this agent have to retrieve log, logs for that application. So every time this agent is doing the retrieve for the logs, he's putting data for the context window. And every time this context window is going to grow. It's going to grow every time because when we have, uh, one
- 1:42
session, and each session is going to retrieve the last session because it's going to remember everything. That's the way that a-agents remember, and they use the context window for that. In the past, about probably one year ago, before this agentic era, we saw that just add more token, that will fix it. Yeah, let's use more bigger context window. That is going to fix it. And today we know that is not true
- 2:11
because we have the, uh, attention U-curve. What is this? We have a lot of data, but if that data is super, super big, the agent is going to forget the middle in the, in the, in the curve. It's only going to remember the first part and the last part that we sent to that context window. And I'm talking here about context engineering. But--
- 2:41
I can move. But context engineering is giving the model the information it needs when it needs it. Sometimes the agent doesn't need all the information inside that context window. Sometimes the context window is poisoning the agent because you can have a prompt injection inside your context window. And the, uh, context engineering help us to curate the context window for optimal outcomes and token consumption. So
- 3:11
today I'm going to show you some techniques that can help you to improve the context window. And these techniques, uh, I manage for a strategy. We have, we have the externalize because move this, that, this big data to a persistent storage. You can create a memory pointer to store that data. I'm going to share the deck at the end, so you need to-- you don't need to have to... Please pay me attention,
- 3:42
and because I'm going to share you all the information that I'm going to talk here. And the second one is select. You can retrieve only what is relevant. We have compress, that reduce token because you can summarize and you can compact that information. And the last one is isolate because you can separate context across agents. And let's start. So the starting
- 4:12
point, conversation manager. This strategy, yeah, we have some strategies that are a little complicated, but we are going to start for the easy one that we can use inside the framework. So here in this strategy we are going to use a Strands agent. It's a framework like we are... AWS we maintain. It's completely open source. Yes, it's free and it's AWS. Yeah. It's open source, free, and it's completely model-agnostic.
- 4:42
So here with Strands you can use three different strategy inside the framework. So we have sliding window that allows you to keep only the most recent messages. We have summarization inside the framework that allows you to summarize older messages and keep recent one. And the last one is a combination of two above. We have... That you can combine summarize of older messages with recent one. So first let's see
- 5:12
what is a Strands. A Strands is this. So you only need one line of code to create the agentic loop. So this line here create all the agentic loop behind the scene. You don't need to create the nodes, the graph, nothing, because a Strands can do that for you. And you have the system prompt and the tools. And yes, this is simple because it's using Amazon Bedrock behind the scene. But if you want to use
- 5:42
Ollama, uh, Gemini, Anthropic, you only need to add one line because you can do that. You import the, the library for OpenAI, for example, and then you add the model line, and that's it. So let's return to our strategy. If you want to use summarization conversation manager- That's the only line that you add-- that you can add for this agentic loop. So you have the summary ratio, that is, in this case, is fifty
- 6:11
percent, and you want to preserve, preserve the last four messages. So every time that you invoke this agent in this session, this session is going to manage your context window in this way. So it's going to keep the four last messages, and it's going to summarize the other sixty percent. And yes, this is amazing, but what happen when you have a lot of memory and you don't want to lose, you don't want to lose all the memory inside that context
- 6:42
window? So you can use... Well, always we have short-term memory because short-term memory is when you have the recent conversation history. So you can use this, uh, applica- this, uh, this function that I just mentioned in the short-term memory. But what happen when you want to return and, uh, you for-- you are-- you have a new session for this agent, and you don't have short-term memory? You can have long-term memory, and you can create a vector database for
- 7:12
that long-term memory. You can have different type of memory inside that vector database with different prompts that can understand your conversation and save whatever you want. And the third one is relation memory. You can use Entity Graph. You can go to the, uh, to the guide for Neo4j. They, they can explain you more about that. You can use traversal-based it or health in connect. So you can understand. You can use a combination of three of those. You can have long-- short-term
- 7:41
memory for the mem-- the session that you have. Then you can retrieve the memory that is in the past, the long-term, long-term memory. And if you have some relation thing that you want to keep in a different way, you can add for that a graph memory as well. Let's continue with the other one. So what happen with this application where a lot of logs that I told you? You don't want to put that logs inside your context window every time that you invoke that agentic
- 8:11
application, that agent. You can save that inside a memory pointer. So where... how this work? You have the, the tools. The tools, uh, returns you oof and a bunch of, uh, uh, logs. You can use another tools that is going to capable to took that logs and save that logs in a storage, and you are going to retrieve an ID for that storage. And for that, using a Strands, you can have that using the agent
- 8:41
state. In the agent state, you are going to have all information for the tools. Then you can use another tool that is going to capable to understand that logs, and it's going to send, uh, um, it's going to send to the agent, "Hey, the logs has okay." You don't need the logs to understand that. But what happen is you need the logs. You can go and say, "Hey, this is my ID. Come on, retrieve the logs because I needed to know." So this is something that you can build in your architecture,
- 9:11
and it's something like this. You can have two different tools. That's the way that we create tools inside the Strands safe only with the tool decorator. So we have fetch application logs, and if you can see there, you have the tool context agent state, state. So we retrieve that information, and we save that in a memory pointer. We create a memory pointer ID that is going to preserve inside our context window. But what happen is
- 9:41
your application super complicated, and the tools is not enough. You can use, ah, I forget that. You can use a multi-agent application. So you can create a Strands swarm with multi agents. But you share the context window between agents is like add garbage to the other agents because not all the agents need the information between them.
- 10:11
So you can share the information between the invocation state pointer. So you can create an ID between that swarm, agentic swarm, and you can share the memory pointer between the la-- the invocation state between the agents, and that's the way that we create the swarm in a Strands Agent. And oh my God, nine minutes. I, I want to be fast. And the next one is...
- 10:41
Wait.
- 10:44
The next one is what happen when the agent, uh, wanna stop never. And sometimes we have this agent that is stay in the loop, like, forever, invoking the same tool, the same tool again, and it's never ending. So you can stop that with, um, a minimum amount on invocation inside the loop. So what is the problem? I just read the problem. To sell more soul may be available. So what's... that's happen when you no-- when you,
- 11:14
you, you don't have a clear response when you invoke the tool. So you can avoid that with the tool way, with a clear response and with amounts of invocation for that tool, and you do that like this. So you have a limited tool count. So you set max tool count. For this example is my tools search fly. You can-- You only can invoke that tool three times. And the other tool? Yeah, three time as well. So the agent
- 11:44
never is going to do, to do this, uh, eternal loop. It's only going to invoke that tools three times. And the last one, yes, I almost there. The last one is, uh, is not, uh, something that help you to the context window. It's something that you can use to help your context window because sometimes we are invoking our MCP tools or IMCP, or we have external API That
- 12:14
never answer you at the time that you need it. Sometimes, for example, we have API Gateway, AWS API Gateway. It have a, a maximum time that is going to receive the, the answer. So if you don't have that time, it's going to, uh, like stop your application because the delay. So you can avoid that using a, async handle. So when you use synchronous, your agent is going to wait forever for the answer for your MCP.
- 12:45
But if you have an async handle, you can invoke your MCP, your external API. So then you can save that, and you can retrieve the answer in the next invocation when you need it. Yes, this is going to depend of the case that you are building. And you can do that like this. Create MCP. This is a function that use FastAPI. And with FastAPI, you only need to create the tools. So I'm going to have two
- 13:14
different tools. One is the start long job, and another one is check job status. And I did a lot of tests using this, uh, async handle, and the difference is a lot. For example, if you have also the different, um, for, uh, the... Well, the agent never stop to answer me. So because I wait, then it can retrieve the information. And to close this session, six minutes,
- 13:45
we have the antipatterns. So what are you have to avoid? You have to avoid context stuffing. So you need, you don't need to include everything and the kitchen. You know? You don't have to include everything in the context window because the agent doesn't need all the information there. You can only share what the agent need to create the answer that your user or you're asking for. You can have, uh, native summarization. Sometimes
- 14:15
when you do native summarization, the agent, uh, can, uh, you can, uh, forget what you need. If you are doing whatever summarization, what you might need for, probably in that summarization, you are not going to have everything that agent is needed to answer you. And context population. Again, don't pass everything to the agent. The agent doesn't need all that information. And what is a good thing here? So tool returns large data.
- 14:45
Logs, uh, database, a lot of information. Externalize, select, and use a memory pointer. Many agents need the same large data. You can share invocation between steps. And agent loops or tool hops. You can have clear steps. You know, give a clear answer for the tool. And limited tools counts. And async handle. Thank you.
- 15:16
And of course, if you want to learn more about Strands Agent, we have a huge booth there. We have a lot of swag. We have a little, uh, Legos for Amazon that you can have. And see you there. We have the specialist. And I am Elizabeth, no Morgan, remember. Thank you so much.