AI Engineer World's Fair 2026
Why Your Company Needs a Context Graph (and How to Build It) — Gil Feig, Merge
Read the talk
Why Your Company Needs a Context Graph (and How to Build It)
Gil Feig explains how live APIs, synced data, company skills and reusable summaries work together to answer business questions—and why freshness and provenance must travel with the answer.
From a talk by Gil Feig
At a glance
Ideas worth remembering
MCP connections inherit source API limitations. Broad questions across paginated datasets may require a synced copy rather than many request-time lookups.
Company skills should guide retrieval: defining customer happiness determines whether the agent should inspect tickets, surveys or another source.
Live calls, synced records and derived summaries trade off operating cost, freshness and repeated retrieval work. Choose them according to the question and its data requirements.
Preserve provenance with the context: source, fetch time, access identity and permissions, transformation history, and invalidation rules make answers inspectable.
A company brain needs more than connections
An agent needs access to company systems, but access alone does not tell it which information matters or how the company interprets it. Gil Feig, co-founder and CTO of Merge, frames a context graph as a “company brain”: the information and processes an agent needs to answer questions across the business. His focus is the architecture and its components, rather than a particular technology stack.
Here, “graph” does not require a specific representation of nodes and edges. The context layer brings together third-party data from systems such as NetSuite and Jira, structured information and static documents supplied by the company, memories formed during agent use, and skills that describe how work should happen. These components contribute different things: records describe the business, while memory and skills help determine what to retrieve and how to use it.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
One upset customer is a lookup; all upset customers require a different access pattern
Start with a narrow question: “Why was customer A upset last week?” The agent can look up that customer’s Zendesk tickets for the period. In Feig’s example, it finds three tickets describing a broken API sync and the details of the failure. The customer and time window already identify the records to retrieve, so a live lookup can supply the explanation.
Now change the question to “Which customers were upset last week?” The customer identity has become the answer rather than an input. The agent must inspect tickets across customers to discover who belongs in the result. Extend the window to a year, and the same approach may require thousands of paginated requests, pulling tickets a hundred at a time until the task times out. The question sounds like a small edit; the required data access changes substantially.
MCP exposes tools, but those tools still inherit the underlying API’s access patterns. An API that retrieves pages of tickets does not acquire a semantic search operation simply because an agent can call it. Feig imagines an endpoint that could directly answer an English question about unhappy customers. His syncing recommendation addresses systems that lack that capability; it is not a claim that every semantic query must always use a local copy.
The proposed solution is to sync relevant third-party data into a locally controlled context store, using a vector database for semantic retrieval. That moves the broad search away from a chain of request-time API calls. It also creates engineering work: the team must decide how to store the records and retrieve useful context from them. For a simple Stripe lookup about a known customer, or a single action in Stripe, live MCP access can remain the appropriate path.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Route the question, then apply the company’s definition of happiness
Once both synced retrieval and live lookups exist, the next problem is combining their results without dumping a massive amount of context into the requesting agent. A router sends the incoming prompt to the appropriate handler. A summarizer gathers the retrieved information from the synced and live paths, condenses it, and returns it to the agent that asked the question.
Before retrieval, the handler needs to know what “upset” means inside this company. Ticket count is one possible signal, but tickets could contain positive feedback. Another company might assess customer happiness through Qualtrics surveys and require the agent to consult those first. A skill encodes that process: how to determine whether a customer is happy, and consequently which systems to query. Memory and skills belong in the context layer because selecting context is itself company-specific work.
Return to customer A. The original lookup explained a broken sync using three tickets. With the broader dataset available locally, the workflow can gather customers, count their tickets, summarize their contents and produce a status for each customer. The observable change is from an explanation for a named account to a set of customer statuses. The company’s skill determines how those signals should be interpreted; the synced store makes the cross-customer retrieval practical. “There’s no magic here,” Feig says. The result follows from arranging the data access and business logic.
Where does the company’s definition enter the data flow? The diagram places skills and memory before source selection, then brings the two retrieval paths together at the summarizer. The sequence makes a useful dependency visible: deciding what counts as customer unhappiness affects what data the system should fetch.
Ask which customers were upset last week.
The handler consults relevant skills and memory, selects synced or live context, and returns a condensed result to the requesting agent.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Four tiers balance retrieval cost, freshness and repeated work
Feig uses a broad definition of context: anything passed into the agent. The four tiers serve different roles and carry different operating costs.
- Prompt, skills and memory. These guide the task and the choice of data sources, including company-specific processes.
- Live API calls. These are easier to add and, in the architecture described here, cheaper to operate than maintaining synced context. Prefer them for simple lookups and retrievals.
- Cached or synced data. This supports broad semantic retrieval and deeper questions, but requires a database and ongoing synchronization. Keeping the copy current adds complexity and expense.
- Derived context. This stores information produced from the source data, such as a summary of a customer’s tickets, for reuse in later questions.
Derived context avoids repeating work. When a customer’s full set of tickets comes in, the system can summarize it and store that summary as structured context. A later, common question can retrieve the summary instead of fetching and summarizing everything again. The stored summary is another representation of the source data, so its usefulness also depends on the freshness and transformation history that the provenance layer will track.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Use freshness requirements to choose the retrieval path
The selection procedure begins with a skill or memory match. If a matching skill can complete the task without external data, the system can generate a response immediately. If there is no match, or the skill requires external information, the workflow moves to data selection.
At that point, check whether the cached or synced data is fresh enough under the applicable service-level requirements. A usable cache supplies the context. Otherwise, the workflow calls the live API. Retrieval may repeat until the system has enough information; it can then store a useful discovery in memory and generate the response.
This is a freshness policy, not a guarantee that live APIs can replace every stale dataset. The earlier year-long ticket question still has the same pagination problem. A live call works well for a bounded lookup; it does not remove the need for a maintained synced copy when the answer requires searching broadly. The retrieval decision therefore depends on both acceptable data age and the shape of the question.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Keep the origin and history of every piece of context
“Do not drop where the data’s coming from.” Source links help a reader inspect an answer, but provenance has a deeper role when an agent makes business decisions. A record may be versioned, transformed or outdated. Without its history, the system can treat information synced a year ago as if it described the business today. For synced records, store provenance in metadata associated with the data.
The information to retain answers several distinct questions:
- Origin and time. Where was the data fetched, and when?
- Identity and access. Under which ID and scopes was it obtained? Who accessed it, and which permissions were used?
- Transformation. Was the record changed or summarized, and what did it originally look like?
- Invalidation. What makes it unusable—for example, exceeding a specified age?
A month-old cutoff is Feig’s example of an invalidation rule, rather than a universal freshness requirement. The same metadata supports both the answer interface and investigations after something goes wrong. A traceable answer should leave enough information to work backward from the result to the data, permissions and transformations that produced it.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Build a bounded account update, then retain what makes it useful
The final example asks for a one-paragraph update on Acme, including renewal risk and the last support interaction. Assuming no relevant skill, the workflow first checks the cache for an account summary. If that summary is missing, it performs a live CRM lookup. It then retrieves the support interaction from cached ticketing data if that data meets the freshness requirement, or through a live lookup otherwise.
The live ticket lookup is deliberately bounded to five tickets. Pulling the full history would reintroduce timeout risk. The system combines the retrieved information into a derived summary, may store that summary, and returns it. The resulting paragraph serves a narrow account-status request; five recent tickets should not be mistaken for a complete historical analysis of renewal risk.
That example makes the closing advice concrete. Tools supply operations; the context layer also needs the company’s data and processes, including a skill that explains what to do when a customer complains and where to route the issue. More data does not automatically improve an answer. Pull what the question needs, keep it fresh, organize synced data for the retrieval it must support, and preserve the trail that explains a bad decision.
Merge’s closing product pitch maps onto parts of this architecture: Gateway is an LLM router, Unified syncs third-party data from customers’ platforms, and Agent Handler supplies prebuilt live MCP connections with governance, identity and provenance. These are the commercial components Feig presents for internal and customer-facing use cases. The architectural decisions remain the same: choose the appropriate access pattern, apply company knowledge before retrieval, and keep the evidence behind the answer.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
Related talks
- Beating RL With Reflection: GEPA and Optimize Anything
Develops the skills component from another direction: learning reusable repository knowledge from agent trajectories so later attempts spend less time rediscovering how work should happen.
Read the complete timestamped transcript
- 0:12
Right. Thank you all for, for coming. Um, my name is Gil Feig. I am the co-founder and CTO of Merge. Merge helps companies build production AI very easily. We are the connective infrastructure, uh, to make everything come together. Uh, today I'm gonna be talking about why you need a context graph and how to build it. What I'm not gonna do is talk about specific technology you should be using. Uh, you can talk to your agent about that. It will help you. There's a lot of good technologies that can back this. Uh, the focus of this
- 0:42
is, uh, why you need a context graph, and then sort of in a, in a higher level, how you should structure it, what the components should be, uh, and how things come together to build the best possible, uh, s- brain for your company. If you haven't gotten a hat yet, by the way, my team is walking around passing out hats. They're, they say, "Well connected." Um, so let's dive in here. So I'm not gonna spend much talking about Merge, but just a quick overview of us. Um, we have over four hundred enterprise customers, OpenAI, Ramp,
- 1:12
Netflix, Perplexity, a lot of companies across the board using us for both internal and customer-facing AI use cases. Uh, we have three different products, Unified, Agent Handler, and Gateway, that cover a lot of the components of it. Um, we were founded six years ago by me and my co-founder, Shensi. We've raised three rounds. Uh, we're a hundred and thirty employees in the three most beautiful cities in the world.
- 1:35
All right, so quick agenda here. Step one or number one, we will talk about what a context graph even is and why it's so important that every company has one. We'll go into the different tiers of context, what, what types of data, what types of, of, uh, sort of logic from your company should be making it into your graph. We'll talk about context selection strategy. How are we picking the most relevant content for our agents or context for our agents at any moment in time? Traceability and provenance. How do we know where that data came from, who requested it,
- 2:05
when it, when it's from, is it fresh? All that. And then lastly, some final key takeaways. At the end, I'll stick around if you do have any questions. Uh, so feel free to, to come up to me afterwards.
- 2:18
All right, so what is a context graph in this talk? We're not gonna go into the nodes and the edges that connect each node in a true graph infrastructure. We're gonna talk about the company brain. Your company context graph consists of a lot of different components. It has third-party data. What are the other systems you use like NetSuite, uh, like, like Jira? Um, it has structured context. Uh, so that could be, uh, data coming in from other third parties. That could be data that,
- 2:48
that you all have put into your context layer. Uh, things like, things like static documents that you want to always be passed to your agents. We have memories. These are things that are formed. We all know memory, things that are formed as you're using the agent, telling it to remember things, things that your agent thinks it should remember on its own. And then we also have skills. Uh, and skills are a very, very important part of your context graph, which we'll get to.
- 3:15
All right, I see some picture. Okay, cool. Uh, so moving on. Let's get-- let's go through an example here. We're gonna use this example to guide why it's not, it's not enough to just connect a single MCP server or a couple MCP servers and call it a day. So let's ask a simple question. One of, one of your employees or, or customers asks your agent, "Why was customer A upset last week?" Pretty simple question to answer. The agent goes, it sees it has access to your Zendesk ticketing, and it can check and say, "Hey, what did-- what
- 3:45
tickets were filed by this customer last week?" Great. We saw that they saw that the API sync was broken, um, and specific details about what went wrong. We also see the customer filed three tickets total, all of that. Great. That answers that question. But what if we want to extend that? And what if we want to say which customers were upset last week? That becomes a lot harder because you now have to have your agent load in every single ticket from Zendesk that happened in the last week. What if we're saying which customers were upset in the last year?
- 4:15
We should be able to answer that. Our agent should be able to answer that, but it can't if it has to live look up and pull tickets a hundred at a time, thousands and thousands of those until it times out and it just can't answer the question. So MCP is not fully solving a lot of these problems because the underlying access patterns of the APIs of these other platforms don't support a live lookup methodology. If you could ask your ticketing system something semantically, if Jira exposed an endpoint that said, you know, that you could query in English
- 4:45
which customers were upset last week and it returned that data, that would be great. But no platform's incentivized to do that because it costs them money to vectorize it, and it makes it so you don't have to visit their platform as often.
- 4:59
So you need a syncing layer on top of this. You can't just use MCP to live lookup. For things that you wanna go deeper into, that you want to semantically question, you need to actually sync a copy of the data locally. Okay? This doesn't apply to everything. So you might have, you know, let's say static lookups to Stripe. We're always gonna be asking Stripe a question about a specific customer. We're always gonna be taking a single action in Stripe. That can be MCP. That can be live. But if you need any form of semantic lookups, any form of analysis that is beyond a
- 5:29
simple lookup, simple question, but requires looking across a, a vast data set, you have to sync a copy. So that becomes the next part of your syncing layer-- of your, of your, uh, context layer. So you have your MCPs, your live lookups. You have your synced or your cached lookups as well. And that, you're pulling in from a third party. You're putting in a vector database locally. Again, AI can build that vector DB for you relatively easily. Uh, you gotta, you gotta, you know, work on your strategy of how you store data in the context layer and how you
- 5:58
retrieve from it, uh, for that synced cache layer. Um, but that's sort of the next component here. But the next problem is, all right, now we have all this relevant data. We're, we're pulling from our vector database to get, you know, semantic data, and we're doing live lookups for other data we might want on top. Now we need to bring all that data together 'cause we can't just dump a massive amount of retrieved context back to our agent. So that's where we introduce the router and the summarizer. The router is sending things to the right place. At the beginning, the prompt comes in, it decides where, you know,
- 6:28
who should be handling this, where it should be sent. Once that data's all aggregated in there, our summarizer then takes that data from, from the, the synced context and from the live context, summarizes it, and sends it back to the agent that sent the prompt over. But you all notice we have one thing here that we haven't talked about, and that's memory and skills. That could live separately. Why is it important for your context layer? The reason is that getting context is actually a set of company principles and
- 6:58
processes in and of itself. For example, which customer was upset last week can mean something very different to one company versus another. It could be number of tickets that are opened, um, but actually, what if those tickets were positive? What if actually we don't use a ticketing system and that we know that for customer happiness, we need to actually always go look at Qualtrics first and see survey data 'cause that's just how our company works. We're different. That's why skills are important, because you need skills like how do we determine if a customer is happy? So now your agent should first check the skills
- 7:28
before it then goes and starts checking your synced layer, your, your, uh, live layer for what data to get, and then summarizing and returning it. So the flow here is prompt comes in, gets routed to the right agent. The agent looks to see if it has any memory or skills associated with the question or the, the task at hand. Based on that, it then knows where to look in the synced cache DB and, uh, from the live MCP lookups, summarize, send back.
- 7:57
There's no magic here. There really is none. It's all physics. As long as you know how to calculate, you know, within, within your own, uh, sort of use case, you're great. This flow here, where we looked at the number of open tickets, again, would not have been possible with live lookups, but with a local lookup, it was really simple, and we were able to basically gather all customers, number of tickets they have, summarize their tickets, and then give a status on each one. So it's a mix of some live lookups, some cache lookups. Uh, that's all coming together here to build really, really rich data for
- 8:27
any question that gets asked.
- 8:31
All right, um, and so there are different tiers of context, and some might debate whether prompt skills and memory are context. Context is anything that's being passed into the agent. We're gonna put this all here because they all belong in your context layer. Um, so you have your prompt skills and memory. We talked about how those are gonna really drive the, the where do we get data from. You have your live API calls. These are easy, easier to build, easier to add on. They're cheaper to run than synced context, so you should prefer these when you know you're gonna have simple lookups, simple
- 9:01
retrievals. Cache data is absolutely essential if you want that, that sort of semantic data, the, the complex lookups, deep, deep questions about, about any data. Um, but keep in mind, it's a little bit more complicated to set up and build. Uh, it's more expensive to run because you have to have a vector database, and you're constantly having to sync that data. So it is tougher to keep up to date. And then lastly, something we didn't quite discuss yet is derived context, which is basically any additional context you create
- 9:31
when you're pulling in data from the third party. So this could be, you know, as we're live pulling in all tickets, we're also, for every time a customer's full set of tickets comes in, we're summarizing those, and we're actually storing in our, uh, stored context, our structured stored context, we're storing a summary of the data that we pulled in so that in the future, if we have a common question, we don't have to then also go, go re-look up and re-summarize everything.
- 9:58
Awesome. So all four of those come together. You saw the, the chart for how they do that. Um, and yeah, those are the tiers. So quick flowchart of how this works when, when a query comes in or a prompt comes in. So first thing here is do we have a skill or memory match? If so, yes, we wanna execute that skill. Does the skill need external data? No? Then we're just generating a response, and we're sending that right back. Unfortunately, it's not always this easy. So what happens if we don't have a skill or memory match? No?
- 10:28
Well, then let's go right to, to data selection here. Uh, and same thing for if that skill needs external data. So we say yes. The first thing we're doing is checking our cache or our synced data store, and we're saying, "Is the data here fresh? Is it within our SLAs? Can we use this data 'cause it's viewed as fresh enough?" Yes? Then let's use that cache context. No? Then unfortunately, we have to go down, and we have to call the live API. Now, once we do have enough context, we're gonna loop through these a few times till we build up the context we need.
- 10:58
Once we do have enough context, we might wanna store a memory of, of whatever we just, we just discovered or, or figured out. Um, store that to memory and then generate the response and send it back. And then same thing here, if we didn't have enough context, uh, I just showed that, but we have to call the live API, and then we will, we will, uh, generate that response back to the client.
- 11:21
I'll leave it here for a second. I, I see some people looking and taking pictures.
- 11:30
All right. So last piece here, pretty much at the end, traceability and provenance. Doesn't seem important up front, uh, up front. Do not drop where the data's coming from. Always store it. Uh, you, you have that from MCP. You also have that in if y- for anything you're syncing in vector DBs, there's something called metadata. You wanna store it in there. Uh, it's associated with the data. It is critical. It's similar to, you know, the way that this surfaces in ChatGPT or any surface that you're using is when it tells you a fact or it says something, at the end it says three
- 12:00
sources, and it links to where it came from. That is even more important when you're making business critical decisions off of that data, when that data is, is versioned, when that data can become outdated, when you're letting it make a decision on data that it synced a year ago. You don't want that. Um, so you need to know where it came from. Where did we fetch it? When did we fetch it? Which is not here. Under what idea-- Under what ID and what scopes? Who accessed the data? What permissions were used to access it? Was it transformed? What did it originally look like? And then lastly, what should invalidate it? That could be age, for
- 12:30
example. The data's more than, more than a month old, we don't wanna consider this ever. Really important for showing in your platform, and if not, for showing in your internal platform or whatever you're building, even for, for being able to investigate when something does go wrong, 'cause it will.
- 12:47
So putting it all together, here's another example of one that we might use. Uh, so give me a one-paragraph status update on our customer Acme. Include renewal risk and last support interaction. First, we're gonna check the cache for the account summary. We're gonna assume there's no skills here. Uh, so we'll check that. If it's nil or missing, we're gonna call the CRM using a live API lookup. Then we'll query the ticketing system for the last support interaction. Again, that's cached if within the SLA. Otherwise, it might have to do a live lookup. We're only-- we only wanna
- 13:17
pull five tickets. We can't pull the full history because, again, when we have live lookups, we run into timeout risk. And then lastly, we'll generate a derived summary. We might store that derived summary, and we'll pass that back.
- 13:33
So key takeaways, tools are not context. A context graph is gonna help you reason across the whole company. It's not just tools. It's data. It's synced data. It's processes that you have as a business. Who is documenting, you know, when a customer complains, what do we do? How does it know to always route to the right place? That is a skill. These-- It's your company's skills, it's your processes, it's your data, and it's your tools and actions you need to take. More data, not always better. You want fresh, high-quality data. You wanna pull only what you need.
- 14:03
When you do have to sync data, you wanna structure it properly so it's accessed in the right way. And then lastly, every answer has to be traceable. You do not wanna be left in a position where your AI made a bad decision, and you have no idea why it did that.
- 14:19
So last piece, plugging Merge here. Enterprise AI runs on us. We power these features for a lot of companies. We have Gateway, which is an LLM router. We have Unified, which helps you sync data from third parties out of your customers' platforms. And then we have Agent Handler, which is live MCP connections pre-built across hundreds of tools, governed, identity all built in, uh, including provenance. Uh, and yeah, it's plug-and-play for both internal and customer-facing use cases. So if you wanna build an easy context layer,
- 14:49
hit us up. Thank you everyone for listening.