AI Engineer Code 2025
Your Coding Agent Is 6 Months Out of Date — Jakub Hojsan, Exa
Read the talk
Your Coding Agent Is 6 Months Out of Date
Jakub Hojsan explains how Exa supplies current evidence to coding agents, why a search tool needs rules for when to run, and how query-dependent highlights keep retrieved context small.
From a talk by Jakub Hojsan
At a glance
Ideas worth remembering
A dependency-related diff can require upstream evidence to explain its purpose. Parameter removal may be a migration requirement rather than a cleanup.
Search integration has two parts: provide the tool and define triggers for using it, such as a dependency or version bump.
Query-dependent highlights select different passages from the same source, keeping model context tied to the question rather than to a fixed page summary.
Queries, sources and supplied highlights give developers a retrieval trace they can inspect when an answer goes wrong.
The ending separates relevance from output structure: embeddings, filtering and reranking select documents, while caller-confirmed schemas organize the returned information.
A dependency migration can look like a cleanup
A pull request removes a parameter, and a coding agent reads the diff as a tidy cleanup. The actual reason is a dependency upgrade: the calling code must change because the upstream API changed. That difference matters in review. Recognizing a consistent edit does not explain whether it is required, or whether it completes the migration.
Jakub Hojsan, a Forward Deployed Engineer at Exa, introduces this problem through the gap between a model’s knowledge cutoff and its release. He describes a typical lag of around six months; that is the talk’s framing, rather than a fixed interval for every model. Repository changes made after the cutoff can therefore be unfamiliar even to a newly released model. A review that relies on remembered API behavior may miss the reason for a perfectly reasonable diff.
The human debugging path supplies the missing step. A developer sees a compilation error, searches for the error or asks whether the parameter was removed, and uses the answer to prepare a fix. For the agent, searching the upstream repository and changelog can turn the same deletion from an apparent cleanup into an explained migration: the dependency changed, the old parameter no longer fits, and the caller needs to follow the new interface.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Give the agent a reason to search
Exa’s search interface returns context-rich highlights for a model to consume. In the migration example, the useful input is the passage explaining the breaking change, rather than the entire GitHub page. Hojsan contrasts a page containing 100,000 characters with a relevant excerpt of about 500 characters. Those quantities describe the presented example, not a guaranteed compression ratio for every search.
But retrieval only helps if the agent invokes it. Hojsan describes an agent rejecting a request to use a newer model because its remembered knowledge says the model does not exist—even though a search tool is available. The failure occurs before retrieval: the agent treats its old knowledge as sufficient and never checks.
A code-review rule can make the trigger concrete: when a diff bumps a dependency or version, inspect the upstream source and ground the review in what it says. This gives the harness two responsibilities. It must provide the search capability, and it must instruct the agent when to use that capability. The version bump becomes a visible reason to verify current behavior, without requiring the model to recognize an unknown API change from memory.
Where does the missing evidence enter the review? The flow below follows the dependency example from the search trigger to the migration explanation. The important connection is between the version change and the upstream passage: a small excerpt can change the interpretation of the diff because it supplies its cause.
A change in the diff triggers the review rule.
The harness rule triggers retrieval; the upstream explanation gives the parameter removal its migration context.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Make retrieval inspectable and independent of the model
Once search runs, the integration determines what a developer can inspect and how much context the model receives. Hojsan presents four reasons to use Exa as a separate search service:
- Transparency. Exa exposes the queries, sources and highlights passed into the model. Sending that trace to telemetry lets a developer inspect the retrieval step when an answer goes wrong. Hojsan contrasts this with native search experiences that return sources without exposing all the retrieved content.
- Small context. Highlights select the passage needed for the question, reducing how much page text enters the model.
- Cost at scale. Hojsan claims more favorable search pricing and quality than native providers for sufficiently large workloads. The talk supplies no comparative pricing or evaluation protocol, so this remains a vendor claim rather than a demonstrated cost advantage.
- Model independence. One search API and parameter set can serve models from different providers. Changing the model does not require changing to that provider’s search interface.
The trace is especially useful because retrieval and generation are separate places for an answer to fail. Knowing which query ran, which source returned and which lines reached the model gives the developer something concrete to investigate. Model independence addresses a different concern: keeping the retrieval interface stable while the application changes its choice of model.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The same page returns different lines for different questions
How does Exa make those highlights? Hojsan describes selecting lines at runtime through computation, without using an LLM to synthesize the page. He claims this adds zero extra latency to the search call; the recording does not establish the timing conditions behind that claim. The mechanism he does explain is query-dependent selection: the question determines which parts of the source become model context.
The demonstration keeps the website constant and changes the request. A query for Hojsan’s biography returns information about photography and riding motorcycles. A query for his phone number returns a different set of information from the same website. The observable change is in the selected content, rather than in the source being searched. A biography passage can be relevant to one question and useless to the next.
What changes when the query changes? The comparison below makes the shared source visible. Each request selects different lines at runtime, so the model receives context shaped for its task rather than a fixed summary of the page.
The coding application follows the same relationship. A large repository may contain a “bajillion lines of code,” but the review needs the part that explains the API change. Hojsan also gives a news example: a request for the latest NVIDIA news that mentions Jensen should retrieve content satisfying that question. The useful unit is the relevant passage, not simply the page that happens to contain it.
The source used for both requests.
Changing the question changes the returned context while the source stays the same.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Exa Agent takes on search orchestration
The talk then moves from a search API inside a coding harness to Exa Agent, an offering for applications that need search without building their own orchestration. Hojsan names Cursor, Cognition, Warp and CodeRabbit as users of Exa’s web search. Exa Agent extends the offering beyond web pages to partner data, including Similarweb for web analytics, Particle for podcast intelligence and Crunchbase for private markets.
The conference example changes the task from finding a passage to finding people. A query looks for people who have mentioned the AI Engineer Conference at its location, and Hojsan recognizes familiar people among the results, describing many as speakers. Public mentions can supply candidates for such a list, but they do not by themselves establish attendance or make the list exhaustive.
A related example searches for Exa employees and contact information for outbound work. These examples explain the packaging decision: Exa Agent handles search orchestration for a broader research task, while the API lets an application control that work itself. For coding agents, Hojsan points to Exa’s MCP integration as the route to making search available in a compatible harness. The earlier requirement still applies: the agent needs instructions telling it when to search.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Ranking finds relevant documents; schemas shape the result
The first audience question asks how Exa organizes and reranks billions of documents. The answer describes a multistage search: convert the query into an embedding, combine semantic search with keyword filtering, apply reranking and drop irrelevant results. Hojsan does not specify the internal ordering of every filtering stage or the reranker design. The distinction is useful nonetheless: finding candidate documents and deciding which candidates deserve to survive are separate parts of retrieval.
Exa also curates its index. Hojsan describes it as containing tens of billions of documents while being smaller than Google’s, with an emphasis on document quality. That choice makes coverage and selection part of the product’s tradeoff: search operates over a deliberately curated collection, rather than assuming that the largest possible collection is always the most useful one.
The final question asks whether the caller can supply a schema—for example, to receive a phone-number field rather than free-form prose. Hojsan describes entering the desired fields in natural language, generating a schema and clicking Confirm. The result then follows that schema. Enabling additional properties allows fields beyond those explicitly requested.
Hojsan gives a limit of up to ten fields on Deep and tentatively recalls up to a hundred on Agents; the latter limit is uncertain in his answer. The practical capability is clearer than that tentative count: callers can describe the information they need and confirm an output structure before receiving results. Retrieval determines which evidence reaches the application; the schema determines how that information is arranged for the next step.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Read the complete timestamped transcript
- 0:12
Hello, everyone. Can you all hear me? Good? Sweet. My name's Jakub Hojsan. I'm a Forward Deployed Engineer here at Exa, and, uh, today's presentation is on how we built search for coding and review agents and how we power most of the Bay Area, um, in this regard. Um, coding agents don't really need these ten blue links that you see when you search Google. Uh, what we do is we have a semantic search engine that essentially you pass in a query, and we provide you context-rich highlights to pass into your LLM.
- 0:45
The core problem that we have with large language models in code review and coding agent applications is that there's a knowledge cutoff to all of these models. As you can see from this graph, we're typically lagging around six months from the knowledge cutoff date to the release date. In these instances, you have this gap that exists between when the model was released and when you might have an important change log or a PR push to a repository that you do have to review. Um, in this instance, I do have an example where from the,
- 1:14
uh, Cuda vector store PR, um, that was made a few months back, this is within our model blind spot after GPT 5.5's cutoff date. So in these months leading after the cutoff date, you have all of these changes being made to real repositories that, one, you can't create net new repos with, but you actually can't review at all. Um, so that's kind of where Exa comes into play in this instance. Uh, and we'll actually just review this diff real quick, so you can
- 1:44
kinda see what I'm getting at. But before we do that, like, how would a human even review this? Very typically, a hu- a human would probably go into Stack Overflow and say, "Was this, like, random variable removed?" In this case, it's called inertia check. Um, and, like, "Why isn't my code compiling?" And, uh, if your code's not compiling, you'll probably just paste the error code into Stack Overflow, and you'd actually get the response that you would want, and then you would make a PR to the repo in order to resolve it. Now, in this case, in-- with
- 2:14
using a large language model, many people without web search will go see, okay, here's the error. The diff looks consistent, and it looks like a cleanup. But that's really not the case. The reason why we're removing this parameter, uh, in this instance is because we actually need to refactor to a dependency bump. So using web search, you're actually able to take a look at the repository itself, any breaking changes from this repo's change log, and then explain the migration in depth.
- 2:48
So what's interesting here is that we don't return the entire GitHub repository page. What we'll return is a very small snippet from the page explaining the change. So if the query is exactly this, we have an interpretation step which will take that entire page and distill it down into exactly what the large language model needs to answer the question. So you're not feeding the model a hundred thousand characters anymore. You're feeding it quite literally only five hundred characters to answer the question.
- 3:20
Now, adding a search tool to your agent is not enough. There's obviously kind of a two-step process here. Um, if you just bolt on a web search onto your agent, many of you've seen maybe when you're using Claude Code, you'll say, "Hey, I'd like to use Sonnet four-six or four-seven now." And the model will say it doesn't exist, right? So even though it has access to a web search tool, it's not actually, like, doing anything, um, 'cause natively, Claude Code does have a web search tool. Um, and in these instances, the model simply doesn't know that you're requesting a version
- 3:50
that does exist. So you actually have to provide the model with a set of rules. So when you're doing a code review agent, the rule might look like when you have a diff that bumps a dependency or version, you might wanna look at the upstream source to verify that this is actually true and then ground the actual review in what you find. Um, so when I work with these code review and coding agent customers, there's several instances where this might happen. They might actually add the tool to their agent, and then adding it to the harness is simply
- 4:20
not enough. So it's really a two-step process. It's instructing your agent when to use web search, uh, since m- oftentimes it's not built in, and then also actually just executing the search, uh, in the sentence. So there's four things that Exa kinda gives a model. Uh, we have transparency. When you're using a native web search tool within OpenAI or Anthropic, you actually have this black box that exists. They, they call a web search API. They synthesize all the information. It might take up to ten seconds.
- 4:50
And then you get some of the sources and then none of the content. So that's really point one. When you use Exa, you get the entire span. You get the entire trace of what we looked for, the exact queries, the exact sources, what highlights were passed into the model. Um, this is all given to you and can be passed into any of your telemetry to go investigate when something does go wrong. Um, we also provide token-efficient highlights, which I'll cover on the next slide, which distills an entire page down into exactly what the model needs. And then it also comes down to cost.
- 5:21
Many, many of the times the cost from these model providers are actually so large that you don't notice your web search bill. Um, but given, like, sufficiently large workloads, we're priced actually much more effectively than native web search providers and much better quality in most instances. Um, and then another thing that most of our customers are very happy with is that not being tied to one model provider. Uh, when you use a third-party web search tool, you have this level of standardization that exists. Uh, you can use GLM when it comes out. You can use the new Anthropic
- 5:51
models when it comes out. You can use the new OpenAI models when it comes out. And then you have, like, one fixed API that you're calling instead of everything that the other model providers are using. And you have a fixed set of parameters that are extremely flexible, um, to what you would need across all model applications. So here's a quick GIF of kind of how this works since I'm more of a visual learner myself. Um, but in this instance, if you're asking about the Exa Search API, we're quite literally just pulling certain lines at runtime, not using an LLM.
- 6:21
This is completely computational. So we actually cut zero-- we add a- zero extra latency to our search call by providing you the contents of pretty much any page on the internet. Um, and I do have an example of how this works.
- 6:37
So let's say you're looking up...
- 6:42
Hopefully, you can see that all right on the query there. If you're looking up myself and then a biography,
- 6:49
you'll quite quickly get a result that I love photography and then what I do in my free time. So I like driving motorcycles, as you can see. It's right here. But then from that same website, I could actually ask for my phone number, and then at runtime, we'll actually just get a completely different set of information. So you can see this is generally applicable to pretty much any use case. In a coding use case, you might have an entire GitHub repo that has, like, a bajillion lines of code, but you don't want the bajillion lines of code to be passed into your model, and you don't wanna use a
- 7:19
model to synthesize that information either. Um, so this works across any type of query. If you want the latest news about NVIDIA, and it has to mention Jensen, then you could also get that. Um, but yeah, this is like one of the many applications that we have, um, in these instances.
- 7:39
And coding is really just the beginning. Um, this is one of the reasons why I joined Exa in the first place. We have-- We had such a large swath of people that wanted to use Exa as a search API to get coding docs, um, and I was super excited for the application. Um, but as we went forward and we took on many popular providers and they loved using us and we provided them a great experience, Cursor, Cognition, Warp, CodeRabbit, all of these companies now use Exa to power their web search. Um, but now we have a new
- 8:09
offering. Uh, it's called Exa Agent. Now, there's many instances where you don't wanna orchestrate your own search, but you need a good search experience, uh, and that's kind of where it comes in. We've partnered with several data providers to not just surface from the web, but then also surface highlights from many different places, such as Similarweb for web analytics, Particle for podcast intelligence, and Crunchbase for private markets. So we're powering the biggest financial firms and hedge funds, uh, with this information. Um, and if you wanna learn more about this, you can go to Exa Connect.
- 8:39
And then I do have a small little message here as well.
- 8:45
I ran a quick query. If I wanted to find all of the people that are attending the AI Engineer Conference, you can quite literally find anyone that has ever mentioned this conference at this location. And it's kinda crazy because I see people that I know and have emailed me here. So, uh, these are most of the speakers that are attending the conference, um, and this applies broadly across every single use case. If you wanted to find all the employees that work at Exa, you can also do that. And you could also get all their emails and LinkedIns for outbound.
- 9:15
But yeah, um, I think this is quite exciting, and we wanted to package our search in the best way possible, um, 'cause we can orchestrate our search quite well.
- 9:27
But yeah, we're u-- If you wanna use Exa today, you can simply call us via our MCP, and then we're on basically every single provider that supports it. Um, and if you'd like any help in setting it up, that's my job, so you can feel free to send me an email. Um, but yeah, there's also a QR code at the top. Um, if you want free credits, you'd also talk to me after the talk. Um, but yeah, if there's any questions, happy to take them. I think I could take one question or two questions.
- 9:58
Sorry, I'll come down real quick.
- 10:20
I'll get it.
- 10:23
Hello. Hello. Uh, he asked how we manage the re-ranking and organization of, like, billions of documents on the internet, basically. Um, and it's a multi-stage process. Um, so essentially we pass in your query, we turn it into a query embedding, and then once we've turned it into a query embedding, there's several steps in between for keyword filtering, uh, alongside semantic search. But, um, there's, uh, many different methods of doing this. Uh, but we do employ re-ranker, re-ranking steps during our search, uh, in order to give you the best information. Um, we drop
- 10:53
results that are not relevant. Our index is highly curated, so we don't have as many documents as Google, but I bet you that we have very high-quality documents in our, like, tens of billions of documents index. Uh, but yeah, if you're interested about any specifics, feel free to talk to me after. But yeah, cool.
- 11:12
Uh, [REDACTED], with like some of your results have a schema, is that something that I can supply? Like if I ask you, "What's your phone number?"
- 11:20
Yeah.
- 11:21
Can I supply a schema where I say, like, natural language-
- 11:24
Mm-hmm
- 11:24
... or like a-
- 11:25
Yeah. Yeah. So we actually have this feature that's quite useful. Um, so you could actually generate a schema, so y- up to ten fields on Deep, and then I, I believe it's up to a hundred on Agents. Uh, but you basically type whatever you want in natural language, and you could generate it.
- 11:46
And then in this instance, it'll generate, and then you'll just click Confirm, and there you go. So and then we adhere to it, and then if you want additional properties, you just set this to True, and then you can get whatever you'd like. You'd also-- My favorite demo is, uh, hobbies. Uh, but yeah. Um, great. I'll take some questions, uh, off the stage, but, uh, thank you all for listening today.