Grep or Embeddings? Agentic Search Over Company Documents — George He, LlamaIndex
Read the talk
Grep or Embeddings? Agentic Search Over Company Documents
George He explains how company-document agents can use hybrid retrieval to find a direction, then list, filter, grep and read remote files to build answers grounded in their contents. The hard work includes preserving document layout, keeping permissions current and avoiding repeated ingestion.
From a talk by George He
At a glance
Ideas worth remembering
Direct file search works well when the corpus is manageable, textual and navigable. Large, mixed-format company collections make pre-indexed retrieval useful for avoiding repeated exploration.
A document-search harness can combine hybrid retrieval with listing, metadata filtering, grep and reading, letting the agent expand or narrow its investigation after the first search.
Prepare structured text and page screenshots together. Preserving table structure and layout matters because downstream models must reason from the representation they receive.
Production freshness includes permissions and metadata as well as content. Ingestion, parsing and indexing are distinct stages with different synchronization work.
Remote grep and bounded reads give an agent familiar file operations without downloading the whole corpus. Connecting a synthesized answer to specific source files makes it easier to check.
Why grep works so well on code
A coding agent can find useful context by following imports, searching identifiers and reading nearby files. Why give it a vector database as well? George He, head of engineering at LlamaIndex, opens with that practical question: better instruction following and tool use have made file traversal increasingly capable, but the choice between direct search and pre-indexed retrieval still depends on the data.
In He’s account of Claude Code’s design, local search won partly because an index would introduce another thing to maintain. Code changes frequently; a vectorized copy must stay synchronized with those changes. Searching and reading the files already on disk avoids that synchronization problem.
The repository also supplies its own navigation aids. Code is already text, usually fits within a manageable local corpus, and has a semi-structured hierarchy. Imports and tests leave “breadcrumbs” that tell an agent where to look next. Those properties make direct traversal useful: a match can lead to a definition, then to its callers or tests, without first building a separate retrieval system.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Company documents change the cost of exploration
A company’s document collection rarely offers the same starting conditions. PDFs, images, PowerPoints and manufacturing schematics mix formats and representations. They do not necessarily contain useful textual references to one another, and their organization has not received the syntactic discipline that code requires. At company scale, access rights also matter: finding a relevant document is insufficient if the person asking cannot read it.
MCP servers, skills and instruction files such as CLAUDE.md can explain how to navigate a collection. They still require someone to maintain useful structure. The decision therefore combines two questions: how much data must the agent explore, and how much curation will the organization do to make that exploration affordable?
He offers roughly 100–1,000 files as a range where local file search may be reasonable. An agent can inspect filenames and iterate through files without spending too much context on discovery. With thousands or millions of files, repeatedly reading the collection becomes wasteful. Pre-indexed search supplies a “rough compass”: enough direction to choose a promising subset before paying to inspect its details. These ranges are practical guidance rather than a fixed cutoff; document complexity and curation affect the choice.
Larger context windows do not remove the cost. He points to million-token context as an available capability, then asks whether it is sensible to pay for document ingestion, rendering and extraction inside model context every time a question arrives. Subagents can separate work and return compact answers, but their own context and execution add tokens and latency. Pre-parsing and pre-indexing move repeated preparation out of the query path.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Give the agent a compass and tools to investigate
The proposed harness gives the agent both indexed retrieval and file operations. Semantic search narrows the initial field; familiar operations such as grep, reading, filename filtering and metadata filtering let the agent investigate what it finds. The agent can choose its next operation instead of following a manually fixed choice between RAG and file traversal.
The five tools answer different questions:
- Hybrid search: Where should I begin looking? Retrieve a promising set of documents or chunks.
- File listing: What else exists in this part of the collection? Inspect the surrounding corpus.
- Metadata filtering: Which files match useful attributes? Narrow the collection using information attached to documents.
- Grep: Which files or passages contain a specific word or pattern? Check precise textual evidence.
- Reading: What does the document actually say? Inspect relevant content, including rendered pages when text alone is insufficient.
For the compass, LlamaIndex recommends combining keyword retrieval, such as BM25, with semantic similarity. Keywords help when the agent knows a particular term; semantic retrieval helps when relevant content expresses the idea differently. Fusion combines and weights the results. A subsequent reranker can examine, for example, the top 100 candidates with an LLM or another deeper scoring process. He reports better retrieval from this combination, without specifying a benchmark or effect size that would establish a universal improvement.
The first ranked result set does not have to become the final evidence set. Once the agent learns relevant metadata or finds a promising file, it can list related files that the initial search missed, expand its investigation and narrow again. This matters because the agent begins without knowing every attribute or document in the corpus. File operations make retrieval an investigation rather than a single top-k lookup.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Reading a document requires preserving its structure
A scanned chart or schematic does not become understandable merely because its words have been extracted. Its meaning also depends on spatial relationships. Document preparation must therefore produce useful representations before the agent asks a question: text that preserves structure, and rendered page images that retain the human-readable arrangement.
LlamaParse and the open-source LiteParse enter here as tools for structural content extraction. Tables, charts and segmented information can be represented in Markdown or another form the agent can read. During rendering, the pipeline should also persist page screenshots. Keeping both representations gives the agent a way to read extracted content and inspect its original arrangement without rendering the document anew for each query.
What information must survive parsing? The diagram separates the two representations of the same document. Structured text supports search and focused reading; screenshots preserve layout for visual inspection. Their shared origin lets an agent connect a retrieved passage to the page it came from.
Parsing quality should be judged by the relationships downstream reading needs. For an OCR pipeline such as Textract or Tesseract, He asks whether it captures the table representation and how faithfully it preserves layout. If the representation loses location or structure, the later model has to interpret damaged evidence. Scaling and distributing parsing helps process the corpus; it does not substitute for checking that the resulting representation remains useful.
A PDF, chart, scan or schematic contains text and spatial relationships.
Preparation preserves searchable structure and rendered layout; the agent can inspect both when deciding what a passage means.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Keep content, permissions and indexes synchronized
Production introduces two moving targets: the document and the rules governing access to it. Multi-tenancy requires preventing one customer from seeing another’s data. Connectors to Google Drive or SharePoint must also bring in permission metadata and keep it current. A permission update should not force the system to re-parse all the underlying content. Freshness includes metadata as well as document text.
He separates synchronization into three stages:
- Ingest: Bring documents into a synchronized system.
- Parse: Convert them into a standardized representation—Markdown or text for this demonstration, with other representations possible for multimodal uses.
- Index: Transform and export the prepared data into the indexed state that the search and traversal tools use.
Storage then trades memory residency against loading data from disk into faster memory. Drawing on LlamaIndex’s experience with millions to tens of millions of documents across tenants, He recommends persisting data on disk with a memory layer for efficient retrieval. The demonstration uses turbopuffer. The recommendation addresses the size of the stored corpus; keeping everything in memory is a different cost and retrieval tradeoff.
Metadata goes into the vector-storage layer during ingestion so queries can filter large collections, including by attributes used for permissions. This makes filtering part of retrieval rather than something the agent discovers only after downloading content. The talk identifies permission synchronization as necessary production work, but does not provide a complete access-control implementation; storing permission tags alone does not explain how every read operation enforces them.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
From 135 Alphabet filings to a cash-flow table
The demonstration starts with 135 pre-indexed files derived from Alphabet’s financial reports, mixing HTM files and PDFs. Source-location metadata accompanies the documents. Additional tags such as year and company would give the agent further ways to filter the collection. The initial question is deliberately broad: find cash-flow statements.
With keyword and semantic retrieval set to an even balance, the search returns chunks paired with screenshots of their source pages. The observable change is from a collection of files to specific candidate passages with visual context. A chunk points the agent toward relevant content; its page helps the agent inspect the surrounding document, and the tool can supply more context around that chunk.
The next operation becomes precise. Grep can look for a cash-flow phrase or regex pattern inside a selected file, then a read operation can fetch text using an offset and a maximum length. In a separate file example, He reads 1,000 characters beginning at offset 500. The important mechanism is bounded access: the agent receives the passage it requests rather than downloading the entire remote collection.
How does a broad search become a source-backed table? The flow below follows the cash-flow example from indexed candidates to focused inspection and synthesis. The return path matters: inspecting one result can change what the agent searches for next. Remote file operations let that investigation continue without requiring a local copy of every filing.
The chat request asks for cash flow as a table covering 2021–2025. While the rerun is still thinking, He opens an existing response and describes a cash-flow summary assembled from multiple files. The retrieval process can search, review screenshots and context, read particular files, then search again with a revised scope. The displayed result illustrates multi-document grounding; the recording does not establish the completed rerun or supply the table’s values for checking its numerical accuracy.
The ending makes trust a property of the whole investigation. Search finds likely evidence, but the harness lets the agent inspect it, revise its search and connect the answer to particular files or chunks. For the cash-flow task, that connection gives a reader somewhere concrete to check the summary. The index supplies direction; the subsequent reads supply the substance of the answer.
Pre-indexed HTM and PDF documents with source metadata.
Hybrid search narrows the corpus, focused reads inspect the evidence, and revised searches can add missing material before synthesis.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Read the complete timestamped transcript
- 0:12
Uh, we're gonna get going. Nice to meet you, everyone. Uh, my name is George. Uh, today, uh, we'll be going over a talk on search, document search, and, uh, how we get better results here at LlamaIndex. Um, my name is George. I'm the head of engineering here at, at LlamaIndex. We are a Series A startup. We focus quite a bit on document parsing, orchestration, and, uh, knowledge management. Um, today, the topic of the talk will center around how we've helped enterprise customers scale up, search,
- 0:42
grep, and generally manage context using agentic orchestration and just general infrastructure tooling.
- 0:51
Uh, the plan today is gonna be going through a discussion or a debate around, uh, how to best orchestrate and perform document search, uh, how to get the best results, and we'll iterate through their, uh, the technical challenges that we've seen all the way to a live demo at the end if we do have time.
- 1:12
All right. So to get started, um, the question that I like to pose to everyone is when and why specifically do agents need vector context in some discussions or just context, period? Um, and the major takeaway here is over the last few years, models have gotten much better at orchestration and instruction following. They're able to traverse local file systems. They're able to work with skills, MCPs, in order to retrieve information in ways that are better than before. But
- 1:41
you still have your local search and your local context problems, right? Um, in order to manage large corpuses, in order to manage complex documents, easiest way is still to manage it locally. Um, and so there are kind of two evolving approaches to orchestration, uh, between grep or what I like to call kind of local file search and file execution, or embeddings or pre-indexing, pre-retrieval. Um, we'll kind of talk through what the different approaches offer and when you might prefer one over the other
- 2:11
or when it's appropriate to do a hybrid of both.
- 2:16
Uh, to get started in understanding where this problem kind of popped up in industry, um, when Claude Code launched, uh, you know, there was a lot of discussion around the internal mechanics of how Claude Code works. It's actually a really interesting problem to think about how you're going to traverse and manage your local file system if you're going to go through coding, right, as a solution or a problem. Um, and Anthropic's taken the approach of using your local file system, grep, read, using blob management basically
- 2:46
to manage your local file system, and it's perfect for your local execution needs because code is easy to manage locally, it's concise, and it's already in text format. But you'll also notice that, um, this tweet by Boris, uh, Claude Code creator, kinda highlights that, uh, Claude Code experimented with a vector database, didn't end up choosing it, and the reasons aren't just for performance reasons, but because of practical scaling and maintenance reasons. It's hard to keep a pre-indexed vectorized
- 3:15
index in sync, especially as code keeps on updating.
- 3:20
And so specifically, the solution that Claude Code took, which was no pre-indexing, no RAG, no vector search, uh, works really well, but only because of a few characteristics of code. Um, they did not care about pre-indexing information because the code in your repo is usually small. It's usually semi-structured. Usually, your code leaves breadcrumbs for what else is being imported, what else is being tested, um, and it leaves a natural hierarchy for how an agent might move about your file system. Um, and the largest part
- 3:50
there is code is already on disk and local, uh, especially if your agent has context, so why complicate it with a pre- pre-retrieval system?
- 3:59
And the reality is when you switch over to larger company datasets, uh, where you have images, you have PDFs, you have PowerPoints mixed in, uh, it's harder because there's not really a folder to grep. Uh, you can't really grep the folders across a company dataset because it's not organized like code, right? You have to think about how much effort has been put into curating the code that you write, right? It mu- it must be syntactically sound, it must have proper references. And then you see your company database or, you know, whatever is in your
- 4:29
data lake, um, and you realize it's really hard to keep things in sync, right? The scale is at a completely different level. Uh, you care quite a bit more about security, uh, who can access what information. Uh, the format is not standardized. You might have really complex PDFs, schematics, manufacturing diagrams mixed in, and the scale just makes it so that you can't run this locally anymore. So how do, how do you keep a system in sync, right? The s- industry has gone through many different solutions. Um, you can use a MCP server. You can try to manage this using
- 4:59
Anthropic's CLAUDE.md, so you-- in addition to having just this blob of a data source available somewhere, you keep track and let your agent know, "Hey, you can traverse the data this way," right? So MCP skills, Claude, Claude, uh, agent files, instruction files, these are all options for keeping track of some type of hierarchy. But the major takeaway here is actually it really comes down to the scale of the data that you have and how much curation effort you're going to put in in order to maintain that retrieval system or that data
- 5:29
management system that makes sense, right? People are talking about context engineering, about harness ma-- you know, making a harness, but at the end of the day, it's all about keeping the right primitives in place. Um, but in order to do that, you have to understand what's the scale that you're dealing with, right? So if you have hundred to h-- you know, a hundred or, you know, a hundred to a thousand files, uh, it may be appropriate to just keep down to local file search 'cause your agents can manage this in context. It won't burn too many tokens just iterating through files or managing the literal file
- 5:58
names as it traverses. Um, and- But as you scale out, right, to thousands or millions of files, that's not possible. You don't wanna burn those tokens. Uh, you don't wanna read the files every single time. You wanna have some needle in the haystack sense of roughly where you're looking for. At least if it's not the needle in the haystack, it's a rough compass on the direction you should be going, right? Um, and so we'll talk a bit about what-- kind of what the right approach is, uh, between hybrid-based search, semantic-based search, as well as, uh, your old-school file traversal here.
- 6:29
Uh, the other reason for why you might care about this is obviously as each model increases in its memory capabilities, right? We got one million token context length now, uh, before it used to be two hundred K, before that a hundred K, and even smaller before then. Um, the important thing to think about here is actually as the model capabilities increase, what you're willing to pay as you're going through this process, right? Do you want to shove those tokens through the model's ingestion pipeline? Do you want to, uh, send binary data over to your model for it to try to render
- 6:59
in maybe its own engine and then extract data out of, uh, for your search? Um, it just becomes extremely expensive, right, in order to use your model context for ingestion and search. Um, and, you know, there have been solutions around how you can use subagents, right? Have one larger agent issue subcommands so that a smaller agent can deal with the memory context, the memory bloat, and just give a response back. Um, but the issue there is because each subagent needs full context, because subagents, uh, even though they might use
- 7:29
cheaper or simpler models, uh, if you've noticed in Claude Code, sometimes your, you know, main session may issue commands to weaker model, uh, sub sessions in order to execute the jobs. Um, the environments are actually rather token-heavy and latency-heavy, right? So in order to speed up your process and make sure that you're not burning tokens, unnecessarily reprocessing documents each time, uh, you may actually think about building out a pre-indexing system. Um, and so we'll kind of talk about that here,
- 7:59
um, in kind of building out a harness and letting an agent choose what's going on, right? So, uh, specifically, we don't really want our system or any system that can scale out to be manually deciding, uh, should I use a hybrid RAG approach or should I just individually search through files? And the cleanest way that we found that in this era of engineering is, uh, managing that with a harness. So give your agent the tools needed
- 8:29
to either search through specific files using your old-school grep, your old-school cat, file name filtering, or metadata filtering in addition to semantic filtering. Um, and that way, you can get kind of the best of both worlds going through traversal, uh, so you can get the rough direction of which files you want to search and use semantic search to down scope your files, and then using your agent's built-in file traversal capabilities to dig into the meat of the details so that you can actually ground your answers instead of hallucinating them.
- 9:01
So the five tools that we'll talk through today, um, enabling kind of document search or, uh, complex document search, uh, as you might see in like insurance, finance, healthcare industries, right? Where you just have a ton of uncategorized or semi-labeled data, is what are the tools that we're going to use, right? So the question first is, where am I even going to look? What's the compass direction we're gonna go? Um, and then you have three tools that are going to be there, and, you know, you have a lot more of these in your local file system to figure out what
- 9:30
actually is in this subset or this subcorpus of documents that we've identified with our compass direction, right? So listing files, traversing files by metadata, uh, grepping specific files to see if there are multiple keywords available that make a specific document more desirable. These are all things that you could run in a pre-indexing pipeline or a live pipeline. Um, and the last one is kind of, you know, what does the document actually say? And specifically, for multimodal documents, you need a way to pre-render the documents, and this
- 10:00
is one kind of unique approach, uh, to multi-page or long schematics, which is in your pre-indexing process, one of the most important parts is to chunk or break your document down into digestible portions, right? But that applies to not just the textual side, but also the rendering side. Uh, 'cause one of the hardest parts of freeform document rendering is you actually have to rasterize or render the document into a human-compatible view. Uh, these use cases are usually where you have an actual person that would have done this, uh, instead of AI,
- 10:30
right? So these schematics, these PDFs, these PowerPoints, they're not really built for pure machine consumption. Um, so giving the agent a pre-built screenshot index to search through is also very useful as you're performing RAG.
- 10:46
All right. So we'll take a look at some of the details of the tool itself. Um, so in terms of retrieval itself, right, which is the compass direction, there are tons of approaches you can take here. There's graph RAG, there's traditional RAG, there's kind of your old-school keyword-based BM25 search. Um, we, we usually recommend a hybrid for enterprise use cases, and that's just mostly because, uh, keyword-based search has its place for developers. Keyword-based search also has its place for AI agents that are looking for specific pre-indexed words in a corpus,
- 11:16
right? Um, at the same time, semantic meaning and similarity is really important. Uh, combining the two through a hybrid-based search usually gives too much-- gives you much better results. Um, you can see in various benchmarking data sets that, you know, just raw keyword or vector will get you to a baseline. But when you put in a fusion, you know, you combine and weight, uh, keyword as well as vector search, you get better results, uh, 'cause you aren't blinded by, you know, specifics of your search, uh, method. And re-ranking the final results, you know, you might take the top one hundred,
- 11:46
re-rank them using an LLM or some other deeper process, uh, ultimately gets you much better retrieval results. Um, but that's just one part of the pipeline, right? That's the document search. Uh, we care about results.
- 11:58
Um, and so tools two through four we'll talk through and kind of show later, which is, um, once you have the compass direction of what are the-- what is the general group of files that might matter, right? What's the metadata that might associate most closely with your query? Uh, 'cause your agent at the start might not know all the metadata or the extracted content of your system. Um, it can then traverse through the file hierarchy, uh, search for similar files inside of your corpus that weren't even returned in that initial search in order to expand out and then further narrow down its results. Um, so
- 12:28
giving your agent the primitives of listing out the files, grep, and read that you have locally is really important in a shared and distributed file system, especially if it's one that you don't want to download onto your agent's context every single time, which can get expensive.
- 12:45
And, um, diving a bit deeper into the content itself, right? Um, this is where, uh, complex documents like, you know, stuff with charts, stuff that's scanned in, stuff that has, uh, you know, schematics, for example, kind of breaks apart for your traditional document indexing and chunking workflow because it is not a human-read line of text, right? It's a separate representation. Um, and so thinking about how you do your pre-indexing flow and parsing flow, uh, so that you don't have to reread the
- 13:15
documents every single time using your agent is also important. You don't wanna burn those, uh, you know, binary-encoded tokens, uh, every single time you wanna read a new document. So we usually recommend pre-indexing, uh, or pre-parsing your documents using a structured parsing flow. Um, we have a product called LlamaParse. We've also open-sourced a tool called LiteParse, which handles structural content extraction. So this is like if you have tables, charts, or segmented, uh, information, you want to encode that as either Markdown or just something that your agent really can understand the spatial reasoning of.
- 13:45
Uh, because when you have multimodal data, uh, spatial reasoning is important. Last one there is, of course, uh, in your pre-parsing flow, you usually wanna generate those screenshots as you actually render the page, right? So whatever you're doing for pre-parsing, you probably also want to store and persist screenshots as they're being generated from document rendering. Uh, it's not always clear if something is text or is, you know, got metadata associated with it. You probably wanna combine the two together.
- 14:13
Um, and so parsing is an important part, but it's not the only part of a pipeline. Uh, we usually, you know, advocate for people to think about how you can distribute parsing, um, and think about scaling out to improve your parsing pipeline, right? Um, there are some metrics that I would advocate for people to think about, you know, if they're relevant to your pipeline and, uh, kind of how you can optimize or, uh, improve your pipeline for each of those dimensions. So if you dump your data into a traditional OCR pipeline like Textract, uh, or you use a
- 14:43
OCR engine like Tesseract, um, does it actually capture the tables-- the table representation? How faithful is it to act- the actual content layout, right? 'Cause if you get those wrong, your LLM on its semantic reading side later down the line is going to pay for it. Uh, it might be fine at the start, but if your LLM doesn't understand the location and spatial reasoning, it falls apart, um, at some point.
- 15:08
All right, so before we get to the demo, uh, there are just a few caveats to think about with building out a actual production pipeline, right? Which is when you have actual customers that care about using it and scaling it out, there are a few things that will always be brought up. Um, there's like multi-tenancy. You know, can one person see another person's data? How does permissioning metadata get pulled in, right? If I'm using Google Drive or if I'm using SharePoint, how do I pull in and yank that permission metadata without causing the whole system to grind to a halt or without re-indexing and re-parsing the pipelines each time? Um, and
- 15:38
that's related to the last one, which is freshness, right? This is kind of why Anthropic, a lot of other companies have kind of given up on trying to pre-index data, uh, because data freshness is really hard to achieve if your datas are complex, your data is complex, and you also have metadata that needs to be up to date.
- 15:56
All right. So, uh, in order to break it down more easily, um, we like to think about the synchronization process as three steps, right? O-one is pure ingestion. You know, can we get all your documents into a synchronized system? Uh, the second step is going to be parsing, which is how do we actually convert your document to a standardized format? Uh, for our use cases today, we'll talk through converting all documents into a Markdown or a text format, but there are also other formats you can use if you only care about multimodal representations. Um, and last stage is obviously, uh, re-munging that
- 16:25
data and then exporting it into a pre-indexed state, um, in order to make it available for all the tools and steps that we talked about before.
- 16:39
All right, so, uh, I'm gonna skip through this, uh, for the sake of time. Um, one major thing that we wanna highlight here is, you know, storing your data is actually quite important. Um, in order to distribute your data out, you usually want some type of layered storage system. Um, a lot of vector storage, uh, services store everything on memory. Um, usually you're playing with some trade-off of do I want it on disk and then load it into a faster memory store, or do I want everything just in the memory store? Um, for us, you know, we've dealt with this problem
- 17:09
at the multi-tenant level to, you know, millions, tens of millions of documents. Our recommendation here is to try to persist your data onto disk and then have a, a shielded memory layer in order to retrieve the data more efficiently. Uh, for our use cases today, we'll be looking at turbopuffer. Um, and turbopuffer, as well as most vector stores, will also handle, uh, permissioning, right? So what's the metadata that you have? How you pull the data in, uh, is actually quite an important part of your system success because when you have too many documents, you need to filter by metadata. Um,
- 17:39
so we push that into the vector storage layer. You know, as the data is being ingested, let's just make sure we tag that into the vector stores. All right. So, uh, we're going to jump through a demo. Um, and for the demo today, the, the process to think about is, you know, how is the agent going to use the subcomponents in order to actually give you a response that's grounded, that actually can traverse multiple documents, and actually gets y- gets you a result that you're pretty happy with? Um,
- 18:09
all right, so we'll go through there here. Uh, let's see. Here we go. Oh
- 18:21
All right, so in this demo, um, we will have a pre-indexed set of documents, um, that is available from a data source that we've defined ahead of time. This is going to be one hundred and thirty-five files, uh, that are derived from Alphabet's financial reports, right? So we have all these, uh, HTM files that are mixed in with raw PDF files, and if you actually look at them, uh, we've pulled in the metadata for where this information can be found.
- 18:51
So if you have file metadata available along with, uh, you know-- Later on, you can tag this with, uh, the year that it came from or the specific company it came from. That gives your agents the capabilities to filter down, uh, the data. Um, and what we'll see here is, uh, in your hybrid re-retrieval flow, right, um, you'll have the option of if you want to be more keyword-based or more semantic-based in your retrieval flow. Um, let's just do a even balance for now, and we're gonna search for, uh, just any
- 19:21
cash flow statements. And we'll talk through kind of the tools that your agent would wanna use along this process, right? So, um, what you've seen here on the right is the set of chunks that are returned from our query, as well as page screenshots that correspond to, uh, what is essentially, uh, the page that it came from, right? So these two together gives your agents the capability to s-- uh, filter down to a specific chunk and request context around that specific chunk in a much cleaner way than before. Um,
- 19:52
you then might want to think about how you would grep for a specific, uh, piece of text in the file, right? So what happens if you have a regex pattern that you're looking for, uh, you know, now you're looking specifically for cash flow, um, and you can find that easily. Uh, in order to find that in a specific file, all you need to do is issue a grep command and then read from that specific command kind of what that offset is and then, uh, kind of what the max length is that you want. So, uh, this is a different file, so I'll just give an example. You know, if we offset five
- 20:22
hundred characters and pull in the next thousand, we'll end up with this text being available, right? So with these two tools, this suite of tools together, your agent doesn't really need the files in a local file state anymore. Um, you can interact with what is effectively a remote file storage without needing to spin up or download all the data that would need to be streamed to your system. Um, and that, uh, gives the agent a much more capable, uh, set of capabilities without actually downloading all the information and burning the tokens. So we will,
- 20:51
uh, go to a chat example, and I'm gonna pull it up here. So, uh, give me the cash flow from-- as a table from twenty twenty-one through twenty twenty-five. Cash flow income as a table.
- 21:14
All right, so give me a second here. I'm gonna run this again.
- 21:27
Give it a few seconds. Uh, while that's running, I'm gonna show you guys the end results. Uh, let's go to the chat. So this is gonna be in the chat flow timeline. Um, let's see. Still thinking there. So as the-- as an agent is running over the, over the process of, you know, retrieving your answers from the corpus, it's gonna look through, uh, your corpus by issuing search commands, right? And as it's issuing the search commands, uh, you want your agent to be able to review the screenshots, review the context, and then
- 21:56
determine what is actually relevant for its search in a much more holistic way, right? So you get to think through the process. You get to read specific files. You get to search through files again and re-scope your queries, and in the end, that gives you a much cleaner result, right? So, uh, in this agent response, uh, you can see the cash flow summary, uh, by, you know-- And you can see that it's been pulled in from multiple files as part of its thinking process. So when we do go into the
- 22:26
traversal flow, um, you can usually find and ground your answers in, uh, in a specific file, and that gives you much more confidence that your agent's found the right component or the right chunk. Um, you know, that's what we're looking for at the end of the day for agent-based retrieval. All right. So we're gonna jump back here. We'll take a look at, uh, that session once it finishes streaming. Uh, but, you know, the idea here and the takeaway is that, you know, search helps you find information, and the harness gives you the results that you can trust. So,
- 22:57
okay, if you're interested, um, feel free to, uh, visit the QR code here. Thank you.