AI Engineer World's Fair 2026
Beyond RAG: A Relational Context Engine That Cuts Token Burn — Unblocked
Read the talk
Beyond RAG: A Relational Context Engine That Cuts Token Burn
Peter Werry explains how Unblocked connects code with organizational history, then builds a document query engine that turns natural-language questions into validated, time-scoped database queries.
From a talk by Peter Werry
At a glance
Ideas worth remembering
Task-specific context connects code with architecture documents, PRs and conversations, reducing the discovery work an agent must perform.
Identity, time, state and aggregation questions require structured queries alongside semantic retrieval.
Schema discovery and identity matching can use procedural methods, reserving the model for query synthesis or ambiguous matches.
Validate model-generated operations and inject tenant constraints outside the model’s control.
Retries can produce an allowed query that misses the intended meaning. The authentication example needed full-text search as well as error feedback.
The local context engine simulator tests task-specific bundles through A/B comparisons on the same task.
The new employee problem follows agents into the codebase
A new employee can read the code and still miss why it works that way. Architecture documents explain some decisions; pull requests and conversations explain others. Incident history carries the “battle scars” that make an apparently straightforward change risky. Peter Werry, founding engineer at Unblocked, starts with this familiar discovery problem: before coding agents, people assembled that context themselves.
A context engine pulls that scattered institutional knowledge together. In Werry’s definition, the output is grounded context: information connected across sources, selected for the task, personalized to the user and aware of permissions. The useful unit is therefore more than a relevant document. It is a set of related evidence that helps explain the work in front of the agent.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Give Claude the architecture and the discussion behind the code
The first comparison asks Claude to plan an optimization of an internal component. Without a context engine, Claude starts by searching the codebase and assembling its own understanding. It eventually produces a recommendation, although Werry adds the familiar qualification: it at least “confidently projects” that it has figured things out.
With Unblocked attached, Claude calls the context engine and receives a task-specific bundle. The bundle connects source code with Notion architecture documents, pull requests and Slack conversations about optimizing the component. Those links also give Claude places to investigate further. Instead of discovering every relationship through successive searches, the agent starts with related material already assembled.
Werry reports $1.29 and about a minute and thirty seconds for the session with context, compared with up to $2.60 and three minutes without it—approximately a 50% cost reduction. These are results from the displayed sessions, rather than a general benchmark. They illustrate the proposed economy: useful context can reduce the exploration an agent performs before it can produce a plan.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A regression needs history as well as code
The next example starts with a change in product behavior. Richie, a colleague working on Unblocked’s code review product, notices that it is surfacing fewer issues. Unblocked narrows the investigation to a likely cause: a model switch changed the review behavior. That remains a diagnosis presented in the demonstration; the recording does not establish a controlled causal test.
Richie then asks Unblocked to fix it. The system creates a pull request whose description connects the earlier PR associated with the regression, a Slack discussion and relevant data. The important relationship runs across systems: a code change records what changed, a conversation supplies the surrounding interpretation, and product data helps explain the symptom. Together, those sources inform the proposed fix.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
“My PRs last week” requires a query
The build portion begins with a distinction inside the same organizational data. PRs, reviews, comments, tickets, documents and conversations contain searchable content, but they also have structure. RAG can find code, the PR that changed it and the thread debating it. An agent loop with additional file-search tools can carry that investigation further.
Other questions require selecting and computing over records:
- Identity and time: Which PRs did I merge last week?
- Counts and rankings: Who reviewed payments the most?
- State: Which PRs are open?
- Time series: Show a weekly ratio chart since January 1st.
The time-series request was how Richie’s investigation began. Retrieving semantically similar text does not itself compute a weekly series over the underlying records. That requires a query engine. Werry describes retrieval as having two halves: semantic retrieval for relevant content, and structured retrieval for filters, temporal conditions and aggregations.
Generating the query is the easy part of this design. The surrounding system must tell the model which fields exist, resolve a person’s name to the identifier used by the source system, validate the proposed operations and recover from errors. A fluent query can still be dangerous or cross tenant boundaries. Natural language makes the interface convenient; it does not make the resulting database operation trustworthy.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Discover the schema, resolve the person, then synthesize
The open-source Document Query Engine starts as an unwired demo: a question produces nothing. Werry adds the surrounding capabilities in stages. The first is automatic schema discovery. For the Mongo document store, the engine samples documents and derives a schema at runtime, giving the model a description of the data it can query.
Inferring enums—fields with a limited set of possible values—is one tricky part of that discovery. An LLM can help, but Werry’s preference is to get as far as possible with procedural methods. Work that looks uncertain at first can still admit a deterministic solution. The resulting demo exposes the schema of a GitHub pull-request collection along with indexes that can help query performance. Because discovery uses samples, it describes the sampled documents rather than guaranteeing that every possible document shape has been seen.
Identity resolution follows the same preference for simple machinery. A human name in a question may differ from a GitHub username or an identity in another system. Procedural matching, including fuzzy matching, connects those identities. If several candidates remain, the model can act as a tiebreaker over a supplied list rather than inventing an identifier.
Query synthesis then receives the discovered schema, the current date and time, and potentially information about the requesting user. Each input removes an ambiguity: the schema constrains field choices, the clock grounds relative periods, and user information lets “my PRs” refer to the right person. With these pieces connected, the demo generates a query and pulls documents.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Keep database permissions outside the model’s proposal
Validation is the most important stage in Werry’s account. The proposed query is untrusted even when synthesis has the right schema. Two kinds of control matter:
- Operator restrictions: Reject dangerous operations, such as a function operator that executes raw JavaScript on the database server.
- Protected fields and tenant constraints: Keep hidden metadata out of the model’s choices. If a tenant ID must constrain every query, inject that constraint behind the scenes rather than asking the model to manage it.
The next question makes validation observable: find PRs with authentication in the title. The model produces a regular-expression approach, and the demo’s validator rejects it because regular expressions are disallowed. This is an application policy demonstrated by this engine. Rejection preserves the restriction, but it also leaves the user’s useful question unanswered.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A safe retry can still answer the wrong question
The retry mechanism feeds the error back to the LLM and asks for a revised query. Technically, it is a loop with a maximum retry count. Validation errors—and, in the broader design, database execution errors—become information the model can use to correct its next attempt.
The authentication-title question now gets through after a couple of attempts, but the correction introduces a different problem. The query uses an exact title match, and no PR has that exact title. The model has found an allowed operation while losing the intended search behavior. Passing validation establishes that a query is permitted; it does not establish that the query expresses the user’s intent.
Adding full-text search supplies the missing capability, and the next attempt succeeds in the demonstration. The engine needs a permitted way to search text, so correction can preserve the question rather than retreat to exact equality. How does the same request change across these stages? The diagram follows the rejected operation, the safe but empty result, and the final search path.
The relationship to notice is that feedback and capability do different jobs. Error feedback helps the model stop using a forbidden operation. Full-text search gives it an operation capable of answering the original question. The completed engine combines structured retrieval with a text-search capability rather than expecting retries alone to solve the problem.
Find PRs whose titles contain authentication-related text.
Validation changes which operations are permitted; full-text search changes which questions the engine can express.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Turn retrieved records into an answer—and test context locally
The final agent demonstration combines the capabilities in a broader question: “What authentication-related things have we been working on last week?” The engine supports filters, aggregations and time-scoped queries; the chat interface shows the queries it creates, then reasons over the retrieved data. Structured retrieval supplies the records that the answer needs, while the agent turns those records into an explanation.
The ending returns to the practical question raised by the first cost comparison: does pre-gathered context help your agent on your tasks? Werry introduces an open-source context engine simulator for environments where connecting Unblocked immediately is impractical. It runs locally, creates a context bundle for each task, and uses that bundle in an A/B test against the same task. That offers a way to investigate the value of context in a particular workflow before wiring in the full product.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
Related talks
- Beating RL With Reflection: GEPA and Optimize Anything
Extends two mechanisms from this recording: using execution errors to improve subsequent attempts, and carrying repository knowledge forward so agents spend less time rediscovering it.
Read the complete timestamped transcript
- 0:12
Are we up? Okay. Can you hear me just like this just fine? Everyone good? Okay. Um, so we're, we're gonna be doing a bit of a speed run here 'cause, uh, uh, we've only got twenty minutes. But, um, I'm c- I'm Peter from Unblocked, and at Unblocked, we build a context engine. So I'm gonna show you a couple of things. The talk's gonna be split into two parts. Um, I'm gonna show you what a context engine is and how it can add value to your organization, and then we're
- 0:42
gonna actually go and build a component of a context engine, and I'm gonna show you a couple of open source projects that you can, uh, go and play with after this.
- 0:54
So there's the agenda. What's a context engine? How does it deliver value? And then we'll build. Um, so let, let's just hop back in time maybe a few years to the times before AI agents. And in those days, um, you were the context layer.
- 1:15
Imagine that you're a new employee to an organization, and you're just trying to get to grips with what's, what's happening. What you usually end up having to do is look at the... look at a whole bunch of code, rip around documentation systems, just trying to find the information that you need to get your work done. Um, and not to mention, like, all of the, the sort of shared history that's built up over time through an organization dealing with incidents and the battle scars that come with it. Um, so what a
- 1:45
context engine does is it takes all of that institutional knowledge, and it pulls it in. And what you get, the scattered context that goes into the context engine, and out the other end, you get grounded context, which means that the context is all interrelated to each other. It's intent-specific. It's personalized. And it's permissions-aware.
- 2:12
And so I'm gonna show you exactly what I mean by that now. Let's imagine
- 2:21
that I'm gonna go into here into Claude. I'm gonna show you a couple of interesting sessions that I had. So, um, I asked Claude here, this is without a context engine attached, to create a plan to optimize a component in our layer called the Source Mark engine. And, uh, I'm sure you're all familiar with what Claude does when it starts a new task. It's like a new employee. It doesn't understand anything about your, your code base or your organization. So the first thing it does
- 2:51
is it greps around your code base and, and tries to figure out what's going on. So you can kind of see that here. And, you know, it, it figures it out eventually, or at least, you know, it, it confidently projects that it's figured it out. Um, and so it provides kind of like a, a recommendation for me there. But now I'm gonna show you what it looks like when you attach a context engine to the back end. So now what it does is it calls the Unblocked context engine, and Unblocked now delivers a very
- 3:20
comprehensive but task-specific context bundle for this task. And what you can see is that it, it's referencing not only the source code, but also, you know, Notion documents that talk about the architecture, um, pull requests, Slack conversations that were had about optimizing this particular component. Um, and the, the, the net out is that you get this really tight context bundle with all these links. So now, uh, if, if, um, you know,
- 3:50
Claude wants to go back and fetch more information, it has all these great jump-off points. But the really interesting thing is, is this. So here I spent, um, you know, $1.29 on this, and it took about a minute and thirty seconds. Um, that was with the context engine. But without, I'm spending up to, you know, $2.60 and three minutes. So what you see essentially is a fifty percent reduction in cost, which is pretty incredible.
- 4:21
Um, now I wanna show you another interesting demo, um, which is, uh, an interesting session that my colleague... Whoops, sorry. Bring this out here. It's an interesting, um, session that my colleague had with Unblocked the other week. So what he... We also have a code review product, and Richie works on that. And he discovered that the number of issues that the code review product was surfacing was, was starting to drop. So he just asked Unblocked, and
- 4:51
Unblocked was able to pull this information, um, and, and, you know, narrow it down very quickly to, uh, a likely cause, which was that we switched from Opus four-eight... or Opus four-six to Opus four-eight, and the behavior characteristics behind the scenes, um, you know, caused all those issues to drop. And, uh, so what he did is he asked Unblocked to just fix it. And now Unblocked, um, was able to just
- 5:21
create a PR behind the scenes, and it has all that context, and it knows what to do. But the really cool part is you can see that it created this description, and it's, it's just remarkable. Like, it, it, um, understood the PR that caused the regression in the first place, and then it found a Slack conversation where the convers-- where, where this was being discussed and pulled on some of the data here to help inform, um, the fix, which, which is just nuts.
- 5:53
So I'm gonna go back to... Whoop, I think I lost my keynote presentation here. Let's see. There we go.
- 6:03
So you have the same problem with agents, right? Cost... Oh, sorry, this is the wrong it's the wrong keynote.
- 6:17
There we go. So now I think what we're gonna do is we're just gonna get right into the building part 'cause we're, we're quite short on time today. Um, and so I- I'm just gonna summarize this. You've got context scattered all over the place. You've got context in PRs, reviews, comments, issues and tickets, docs, conversations. Um, and, and all of that is, uh, semantically queryable, but it's also structured.
- 6:45
And RAG nails the content part. Um, it, it can be, it can handle things like content questions, find the code, the PR that changed it, the Slack thread that's debating it. And if you wrap that stuff in an agent loop, it can go really, really far. You give an agent loop tools to do file search after this, and it, it's very powerful. But it doesn't handle some of the other type of, uh, structural questions, like PRs that I merged last week.
- 7:15
Who reviewed payments the most, open PRs, et cetera. Um, and these are queries, they're not semantic searches. So I wanna just show you again, um, going back to Richie's example here, that he started the, the conversation with Unblocked by asking it a query-based question 'cause there's a temporal element. Can you show me a chart of ratio weekly since January 1st? Okay. So in order to do this,
- 7:46
it's not gonna be able to just pull this information from a vector store 'cause that's not where that information's encoded. So it's, it's gonna have to do that, um, through some sort of query engine.
- 7:59
And so, you know, retrieval has two halves essentially. We've got semantic retrieval, and that can handle those types of semantic-based questions. But today we're gonna work on this, the structured part. We're gonna turn structured queries or natural language queries into structured queries.
- 8:20
But generating the queries is actually the easy part. The hard part is the scaffolding that sits around it. So you need to do things like understand the schema, so you can pass it to the LLM, so it generates the, a tightly bound query. Uh, you need to also be able to map identities. If I ask a question like, "Open PRs by Rasheen," um, it's not gonna know who Rasheen is because his identity in GitHub might be completely different. So we have to be able to map that somehow. And
- 8:50
then when you get the query out of the LLM, it's untrusted. It could have, uh, operations that are dangerous. It could cross tenant boundaries. And so you have to have some way of validating that the query is, is, uh, correct and safe. And then finally, uh, even though you do that validation, there could be errors. Um, and when you actually do the query against the database, there could be errors there as well. And so we can actually get the LLM to try to correct itself if it gets into that
- 9:19
situation. So this is what we're aiming for here. We wanna go to, uh, to this. We wanna be able to say something like, "Open PRs by Rasheen," and then we wanna get this kind of query on the back of that. So just take a sec here. Uh, this is the open source repo that we're gonna be building against, and I'm going to jump right into what that looks like actually.
- 9:48
So this is the repo here. It's called the Document Query Engine, and I wanna show you what it looks like when we just boot it up the first time. So I can say, you know, "PRs by Peter at Ground Zero." Oh, actually, we're actually not running it, so I'll go back.
- 10:07
Let's run it first.
- 10:15
Now I can ask my question.
- 10:19
So when you start, it's gonna have... It's not gonna be wired up at all. It's gonna produce nothing. So we're gonna just gradually build this thing up here, okay?
- 10:29
First thing we're gonna talk about is automatic schema discovery. This is completely dynamic. Mongo is schemaless, not... A- and not all databases are schemaless, but the same kind of principle can apply. You can dynamically discover the schema at runtime. Um, and so what, what the technique involves is sampling a whole bunch of documents in the document store and then deriving a schema from that. And, uh, you know, I just wanna show you, like, one tricky part of that, which is
- 10:59
inferring enums. Um, now the, the, the outcome from this is that you can use an LLM as a Thor hammer, um, and it can do a pretty good job. But actually, you get, uh, really, really far just using traditional procedural methods. So you can... Y- you know, don't always fall back to an LLM to do this kind of what looks like non-deterministic work. This is actually quite deterministic, this kind of thing. So
- 11:29
now that we've discovered schema, I'm just gonna show you what that looks like in our demo. Whoops. First, we have to level up to the next...
- 11:41
And then run. So now we've discovered the schema from our, um, GitHub collection, and this is just pull requests, okay? So it has not only the schema here, but also, um, indexes that will help improve the performance of queries.
- 12:05
The next step is identity resolution. It's that thing I talked about, about mapping, you know, users to, uh, real identities. And, um, again, this is- this can be solved just using simple procedural tricks. So, um, that- that's actually more or less how we do it, uh, in underneath the hood at Unblocked. When we take usernames or identities from different systems and map them across, um, you can do this kind of fuzzy matching, and then in the end, if you have to, you
- 12:35
can pass a list of names to an LLM and have it act as a tiebreaker.
- 12:45
And finally, uh, the last step before we can actually show the cool stuff is this. It's the synthesis part. And roughly speaking, this is actually the easiest part. We just take the schema that we dynamically resolved from earlier, we take the current date and time, uh, maybe some information about the user that's asking the question, so it can easily resolve questions like my PRs and not just Peter's PRs. Um, and then we pass it to the LLM and ask it
- 13:15
to synthesize a query for us. And the outcome, uh, will look something like this.
- 13:24
Let's get to this one.
- 13:29
So now I can go back to my engine and try to run that query again.
- 13:37
And it's going to pull the documents for me now and generate this query.
- 13:45
Now, the next step I- I was talking about was, was validation. This is actually the most important thing. Um, your, your, uh, query could contain really dangerous operators, like the function operator, for example, which would execute raw JavaScript, um, on the DB server, which is definitely not ideal. Um, but there could be other things that you wanna do with this. Like, for example, you might have hidden fields that you don't want, uh, to be surfaced either to the LLM or for the LLM to, to derive and execute.
- 14:15
If you have metadata fields, for example, that contain, um, tenant IDs and you wanna isolate each query to a specific tenant, you're better off making sure that the LLM doesn't interact with those fields at all, and then just constraining it behind the scenes. You can inject those fields underneath.
- 14:37
And that's kind of what that looks like. So again, we're gonna just quickly show the payoff for that and what that looks like.
- 14:47
Let's go up to our validation step. Oops. Run it again. So if I now ask a question like this, PRs with authentication in the title, that is probably gonna derive some kind of like regular expression thing. And there we go. Okay. So, uh, our validation engine said, "Sorry, regular expressions aren't allowed. You're getting bounced." And now we're not allowed to do that.
- 15:17
So validation is cool, um, but what do we do with that? You know, we've got an error, um, but what we can do instead is just take that error and just loop it back in to the LLM and have the LLM kinda self-correct.
- 15:34
And that's kind of what it... That's all it is. That's all it looks like from a technical standpoint. It's just a loop with some sort of maximum, uh, retry.
- 15:45
And so if I try this again...
- 15:52
Bounce up to my retries and run that.
- 15:57
Now, if I try that query again... Oops, that's the wrong one. Let's try this one.
- 16:12
Now what you can see is that it actually took a couple attempts . Um, but it got there in the end. However, um, it constrained it down to use match, and unfortunately, this is an exact title match. We don't have any PRs with that exact title, so we need one more thing,
- 16:32
which is this, full text search. So now if we add full text search to our engine, um, hopefully the outcome that we get is something that looks a little bit like this. So let's try it one more time.
- 17:05
Awesome. So there you go. We got... We went from start to finish there. This is all open source, so you can take a look at how it's implemented and play around with it yourself.
- 17:17
And this is what we built, something that can effectively handle all of the, uh, other half of RAG. So anything that has to do with, uh, structural based data.
- 17:29
You can have filters, aggregations, it's time scoped, and it self-corrects. But what's really cool is that it also powers a chat agent. And so if I go over here and I say something like, um, "What authentication-related things have we been working on last week?"
- 18:00
Now you get this whole thing that will show you exactly what it's doing, all the queries that it creates.
- 18:08
And now it has the data, and it's just kind of reasoning about it behind the scenes, and there we go.
- 18:14
Very cool.
- 18:19
So there you are, structured, queryable context for agents and humans. Now, there's just one last thing. Um, on the description you probably saw that we were promising a context engine simulator. If you don't want to wire up Unblocked, uh, right away because you've got some constraints in your, in your environment, um, we have an open source project that will allow you to do this entirely locally yourself. And so the idea is, uh, it will create
- 18:49
a context bundle for you on a task by task basis, and then take that context bundle and run an AB test against the same task and show you the benefits of having all of that context at your fingertips for your agents. And that's the, uh, PR code for that, if you want that.
- 19:16
All right. And one last thing . If you want some coconuts, come by the booth
- 19:25
. Thank you.