AI Engineer World's Fair 2026
Building the Document Context Layer for AI Agents — Jerry Liu, LlamaIndex
Read the talk
Building the Document Context Layer for AI Agents
Jerry Liu explains how RAG separates into agent reasoning and document context, why reading a PDF requires reconstructing its structure, and how fast parsing plus selective visual inspection can balance accuracy, cost, and latency.
From a talk by Jerry Liu
At a glance
Ideas worth remembering
Modern RAG can separate agent reasoning from context access: the agent chooses queries and iterates, while the document layer supplies readable, searchable information.
Document parsing must recover relationships as well as characters. Positioned glyphs, drawn table borders, and multicolumn layouts do not guarantee a useful reading order or semantic structure.
Choose parsing effort by workload. Regulated extraction, large-scale indexing, and interactive uploads place different demands on accuracy, cost, and latency.
A fast collection-wide pass can precede selective VLM inspection. This gives the agent broad context before it spends more time reading particular tables or charts.
Structured extraction needs a path back to the source and a decision about uncertain values before they enter downstream systems.
From fixed retrieval to an agent that chooses its searches
The early RAG chatbot followed a fixed recipe: split a private document collection into chunks, embed them, store them in a vector database, retrieve the top k matches, and give those matches to an LLM to generate an answer. In Jerry Liu’s account, this was enough to make a basic application that could chat over a corpus. Every question still passed through the same sequence of steps.
Liu, co-founder and CEO of LlamaIndex, frames RAG in 2026 as an agent harness plus a context layer. Better tool use and agent loops make that separation useful: the harness handles reasoning and repeated actions, while the context layer gives the agent material it can read and search. LlamaIndex’s focus has consequently moved from its origins as a RAG framework toward document infrastructure for agents.
The concrete change is who chooses the search query. A fixed retrieval pipeline takes its query and returns matches. An agent can reason about which keyword or search term is likely to find the information it needs, inspect the results, and search again. Retrieval complexity moves into this loop. Even a basic search tool becomes more useful when the agent can supply a better query rather than accepting the first result set as its only context.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Context access and programs move up the stack
As long contexts and compaction improve, Liu sees more attention moving from managing window overflow toward connecting the right MCP servers, skills, and tasks. These connections determine what an agent can reach and do. The context problem expands from fitting text into a prompt to giving a reasoning system useful access to organizational information.
Program definition is changing alongside access. Python and TypeScript remain part of the earlier development picture, but English increasingly expresses tasks, goals, and runbooks for both engineers and nontechnical teams. The progression runs from asking a simple question, to assigning an end-to-end task, to producing a repeatable program that executes at scale. Liu’s further forecast is that longer-running agents may eventually need a goal and scoring rubric rather than a fully described task; that is a proposed direction, not a capability demonstrated here.
That makes document access a persistent bottleneck even as models improve. Useful context includes web information, tools and skills, warehouses such as Snowflake and Databricks, and files held in SharePoint, Box, Dropbox, or S3. Liu estimates that more than 10 trillion pages of human knowledge live in PDFs, PowerPoints, Word documents, and Excel sheets. The scale explains LlamaIndex’s focus: much of the information an agent needs already exists, but its container does not make it easy to consume. Agents also produce more readily readable formats such as Markdown and HTML.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Three layers turn files into usable work
An agent-native document platform needs three distinct capabilities. Parsing makes a file readable; storage makes it available to operate on; workflows turn recurring document tasks into repeatable execution. Keeping these concerns separate also leaves room to tune a frequent task more carefully than a general-purpose agent would.
- Parsing: Convert PDFs, PowerPoints, and Word documents into accurate, token-efficient context, including Markdown and metadata. The output should preserve the information an agent needs without making it ingest all of the source format’s machinery.
- Semantic storage: Capture and store documents behind an interface that lets agents manage and operate on them. This is document management for humans and agents, extending the role previously served by a file system and applications such as Microsoft Word.
- Repeatable workflows: Specialize recurring jobs such as invoice processing, KYC, and claims. A purpose-built workflow gives the system a place to tune cost and accuracy against a known goal.
This platform is still incomplete. Agent-native formats, document versioning, editing, and collaboration remain open areas in Liu’s framing. He names them as work ahead; the recording develops parsing, extraction, and search rather than explaining implementations for those future capabilities.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A PDF draws a page; the parser reconstructs its meaning
Why does document OCR remain hard after more than 20 years? A PDF is designed for printing and display. Text can appear as individual glyphs with coordinates. A table can consist of drawn line segments and text positioned inside apparent cells. The page looks structured to a person, but those drawing instructions do not by themselves give an agent a usable table.
The table example exposes the reconstruction task. First, the parser encounters characters and shapes at positions. It must determine which pieces belong together and recover a representation that humans and agents can interpret. Multicolumn text introduces another problem: the stored sequence of characters does not guarantee the order a person would read the page. Recognizing characters is only part of the job; arranging them into the right relationships is what makes the result useful.
Word documents and PowerPoints offer more structural information, but feeding their native XML directly to an agent creates its own burden. Much of the tagging is unnecessary for understanding the content. A useful parser lifts relevant formatting and semantic metadata, removes irrelevant markup, and can render the page so its overall layout remains visible. More source structure helps, but it still needs to become a representation suited to the reader.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Combine file-aware parsing with visual models
Two approaches attack the reconstruction problem from different directions. Pipeline methods use rules and file information to group text, identify tables and paragraphs, and produce an output representation. A vision-language model, or VLM, can instead read the document visually and turn it into text in one shot. Visual reading can recover layout that a simpler pipeline misses, but Liu identifies three costs: hallucinations even on text-only pages, expensive inference, and insufficient semantics and grounding.
LlamaParse combines file-aware processing with visual understanding. Its described ingredients include optimized PDF, Word, and PowerPoint engines; an agentic harness that routes between cheaper specialized models and frontier models; and parameter-efficient document VLMs focused on particular document classes or elements such as tables and charts. The cost argument follows from that routing: use specialized processing where it suffices, and spend on stronger visual reasoning where the document requires it.
Liu describes this hybrid as lying on the cost–accuracy Pareto frontier: a set of choices where improving one objective requires giving up something on another. He also expects specialized document workflows to outperform general frontier-model use on this tradeoff. These comparative claims are his assessment; the recording does not establish a universal performance advantage across document types.
ParseBench gives this problem a broader evaluation target. Liu describes 2,000 human-verified pages and approximately 50 evaluated frontier models, open-weight models, and specialized OCR solutions. It measures tables, charts, content faithfulness, and semantic formatting, with emphasis on whether agents can understand the result rather than merely whether its syntax is correct. That choice matches the table problem: plausible-looking output is less useful if it loses the relationships needed to interpret the information.
The desired direction is lower cost and higher accuracy, but the distribution of documents complicates any single choice. Simple pages and difficult enterprise documents demand different processing. Liu’s conclusion from the benchmark is that document understanding remains unsolved; covering a range of operating points matters because a real collection contains a changing mix of document types and complexity.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Read everything quickly, then inspect selected pages deeply
Latency adds a third objective. A parser suitable for an offline indexing job may make an interactive agent wait too long, while a cheap first pass may be inadequate for regulated extraction. Liu distinguishes three workloads by what a mistake or a delay costs.
- High accuracy: Financial services and insurance may require 99% to almost 100% accuracy because an incorrect extraction can damage a financial model or trigger a fraud flag. Those figures describe requirements, not demonstrated achievement. Paying more per page for deeper reasoning can be justified by the consequence of an error.
- Low cost: Indexing a million-plus documents per day from a continually updated SharePoint collection favors a scalable offline pipeline. Some imperfections can be acceptable if an agent can later revisit the source and recover the needed information with citations.
- Low latency: Uploading a thousand documents and wanting them processed within a minute puts pressure on VLM-based processing. Liu explicitly includes his own service among those that find this scenario difficult.
LightParse addresses the fast-pass role. Liu describes a free, open-source Rust parser that produces Markdown without a VLM or deeper model, and calls it the fastest open-source parser. That speed superlative is his claim rather than a quantified comparison in the recording. Its practical role is clear: make an initial reading of many documents cheap and quick, while keeping a slower visual parser available as a tool.
Follow the thousand-PDF example through the change in processing. The agent first gets a fast pass over the whole collection and scans the resulting context. If its task then requires understanding values on a page containing a table or chart, it calls a VLM-based tool for that page. The collection becomes broadly readable before the system spends time on deeper visual interpretation. The table that began as positioned characters and lines now receives the additional inspection needed to read its values correctly. This is a proposed operating pattern, not a timed demonstration of processing a thousand PDFs within a minute.
Where does the slower visual work enter the agent loop? The flow below shows it after the collection-wide pass, on a branch selected by the task. The important relationship is the scope of work: the fast parser covers all documents, while the visual tool handles the pages that need deeper interpretation. Liu describes LightParse as an installable skill that complements other VLM-based OCR tools.
The example contains a thousand PDFs.
The agent scans all documents first, then uses slower visual processing when its task requires deeper understanding of a table or chart.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Extraction connects document reading to downstream systems
The closing technical section moves from readable context to structured information. Processing a million invoices, expense reports, receipts, or claims requires outputs that a downstream database or system can accept. This automates the familiar human task of scanning paperwork and entering data. Parsing supplies a usable representation; extraction turns the relevant information into structured outputs.
LlamaParse’s described extraction capabilities attach granular source citations to every extracted output and provide confidence scores and uncertainty flags. Those serve different practical purposes: a citation lets someone return to the source, while an uncertainty flag helps decide whether a value should enter another system. The decision occurs before insertion, so extraction does not have to treat every produced value as equally trustworthy. The recording gives no confidence threshold or calibration method.
Document search also becomes a larger toolset: BM25, grep, vector search, reading, and scrolling. That returns to the opening change in RAG. An agent can choose a search term, retrieve material, and continue reading rather than relying on one fixed top-k response. The document context layer earns its place by supporting both discovery across a collection and closer inspection of the source when the task demands it.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Read the complete timestamped transcript
- 0:01
[music]
- 0:12
>> I think we can get started. Hey
- 0:13
everyone, I'm Jerry, co-founder and CEO
- 0:15
of LlamaIndex, and today I'm excited to
- 0:18
uh give a talk called building the
- 0:19
document context layer for AI agents. Um
- 0:22
really uh big shoutout to the AI
- 0:24
Engineer World Fair for hosting. Um and
- 0:26
if you've seen some of my earlier talks
- 0:28
from the previous uh AI Engineer
- 0:30
conferences, uh we've kind of traced
- 0:31
through a lot of the evolution of how,
- 0:33
you know, um agent advances have
- 0:35
correlated with uh you know, how you
- 0:36
inject context into evolving
- 0:38
applications. Um and so today, in 2026,
- 0:40
we'll kind of talk about three main
- 0:42
topics. One is, you know, what is RAG in
- 0:44
2026? Um two, uh basically that
- 0:47
basically decomposes into an agent
- 0:49
harness plus a context layer. So kind of
- 0:51
going into the modern document context
- 0:53
layer for agents, um and giving you a
- 0:54
little bit of sense of kind of some of
- 0:56
the core capabilities today, um
- 0:57
especially uh unlocking, you know, the
- 0:59
vast trove of document-based data today,
- 1:01
and giving you a glimpse of what's next.
- 1:04
All right. RAG in 2026. But before that,
- 1:06
just a little bit about the company, and
- 1:07
I'll kind of get on with the main talk.
- 1:09
Um you might have seen us as a RAG
- 1:10
framework. We started in 2023, got
- 1:12
pretty popular, um kind of created a lot
- 1:14
of techniques around like advanced RAG,
- 1:16
that type of stuff. Um today, we're
- 1:18
basically the main document
- 1:19
infrastructure for AI agents. Um we
- 1:21
deliver the best platform for agents to
- 1:23
actually read and operate over
- 1:24
documents, and we see ourselves as
- 1:26
unlocking basically the vast troves of
- 1:27
document context even as agents and
- 1:29
models themselves get better.
- 1:33
All right. So, what does RAG actually
- 1:35
mean in 2026, and how are agents
- 1:37
actually inhaling and addressing context
- 1:39
today?
- 1:40
This is a snapshot from um actually 3
- 1:43
years ago, um when I first gave a talk
- 1:45
on this, and naive RAG was basically
- 1:46
just building a simple chatbot over your
- 1:48
private corpus of data. If you flash
- 1:50
back all the way to January of like
- 1:51
2023, um you know, this technique
- 1:54
basically consisted of you have some
- 1:56
corpus of documents, um you chunk it up,
- 1:58
embed it, and put it into a vector
- 2:00
database. You do some naive top K
- 2:02
retrieval, and then you generate some
- 2:03
stuff with an LLM. All the steps are
- 2:05
fixed. You use kind of like a fixed set
- 2:07
of techniques, and then with that you
- 2:09
actually get some basic application
- 2:10
results by just being able to chat over
- 2:12
a corpus of documents.
- 2:14
Of course, the agent landscape has
- 2:16
changed quite a bit and quite
- 2:18
dramatically since then.
- 2:19
One is agent loops and tool use have
- 2:22
gotten a lot a lot better, especially as
- 2:24
both the models and agent harnesses have
- 2:26
increased. And there's basically a
- 2:28
cleaner separation now between agent
- 2:30
reasoning and how they actually interact
- 2:31
with context. If you look at what I call
- 2:34
kind of the modern generalized agents,
- 2:36
which includes, you know, all your
- 2:37
favorite applications and tools out
- 2:38
there from Claude Code, Claude Code
- 2:40
Work, Open Claw, Codex, and a few
- 2:42
others,
- 2:43
you know, the retrieval complexity has
- 2:45
started to get baked into the agent
- 2:47
layer. So, instead of coming up with a
- 2:49
variety of hacks to really work around
- 2:51
the limitations of naive top K
- 2:52
retrieval, you can start getting the
- 2:54
agent to really reason about, for
- 2:56
instance, the best key best keyword to
- 2:57
search for, best like, you know, search
- 2:59
term to actually get back good results.
- 3:01
Even if the retrieval tools are basic,
- 3:03
the agent can input the right queries to
- 3:05
basically loop upon itself and help
- 3:07
solve the task at hand.
- 3:09
Number two, context is moving up the
- 3:11
stack. There's been a lot of
- 3:13
conversations, I think, for the first
- 3:14
two and a half years of Gen AI of how do
- 3:17
you actually manage the agent context
- 3:18
window, make sure it doesn't overflow,
- 3:20
that type of thing. But, I think as, you
- 3:22
know, a lot of the evolving techniques,
- 3:24
compaction, long contexts have evolved,
- 3:27
I think more and more of the
- 3:28
conversation is actually how do you just
- 3:30
hook up the right MCP servers and skills
- 3:32
and tasks to the agent to enable it to
- 3:35
do various types of tasks.
- 3:37
And so, that also applies to agent
- 3:38
orchestration, right? Coding agents have
- 3:40
gotten better.
- 3:42
Abstractions of actually defining what
- 3:44
types of tasks and programs you want to
- 3:45
build have moved a little bit upwards
- 3:47
towards English as opposed to through
- 3:49
code. And so, maybe in 2023 to 2025, you
- 3:52
define programs and you still build
- 3:54
stuff via like importing Python, using
- 3:56
TypeScript through code. It's pretty
- 3:58
clear these days more and more people
- 4:00
are just building stuff using English,
- 4:02
whether you're a software engineer or
- 4:04
you're a non-technical function like
- 4:06
go-to-market marketing. And you're
- 4:08
defining runbooks through English and
- 4:10
kind of like defining the right goals
- 4:12
and making sure the AI is aligned on the
- 4:14
right task. And we see that carrying
- 4:16
over going forward as well.
- 4:19
So, in terms of just like general
- 4:21
evolving knowledge work patterns, and
- 4:22
this kind of forms a foundation for what
- 4:24
we how we think about the context layer.
- 4:26
You know,
- 4:27
in the past maybe you could use AI to
- 4:29
just like ask simple questions, get it
- 4:31
resolved.
- 4:32
Today, you're starting to be able to
- 4:34
define and solve more end-to-end like
- 4:36
tasks just through English and also use
- 4:39
English to start
- 4:40
compiling repeatable programs to
- 4:43
actually execute something at scale.
- 4:45
Obviously, there's been a lot of
- 4:46
discussion on how this will evolve even
- 4:48
more, and we do see AI agents as
- 4:50
behaving even more autonomously through
- 4:53
solving long horizon tasks, looping upon
- 4:55
themselves to be able to achieve a goal.
- 4:57
So, the future will move towards a state
- 4:59
where you actually might not even have
- 5:00
to define the task in English, but
- 5:02
actually more of the goal and a scoring
- 5:04
rubric, and then the agent will use all
- 5:06
available context available to it to
- 5:08
actually solve the task.
- 5:10
So, with that overall framing, let's
- 5:13
talk about what kind of what is a pretty
- 5:15
core focus for us as a company, which is
- 5:17
basically helping to unlock unstructured
- 5:19
context to basically feed into these
- 5:21
agents even as they get really, really
- 5:23
good both in terms of the core model
- 5:24
capabilities as well as the agent
- 5:26
harness. And we'll focus a lot on our
- 5:28
core focus area, which is document
- 5:31
context, but of course there's plenty of
- 5:32
other sources of context as well. In
- 5:34
fact, we do think context really is
- 5:37
everything.
- 5:38
In the end, you could have a infinitely
- 5:40
smart agents, but your ability to
- 5:42
actually get value out of this
- 5:44
infinitely AI agent um, able to do more
- 5:47
and more stuff end to end, is actually,
- 5:49
you know, giving it the right things to
- 5:51
do. Um, and the right things to do
- 5:52
include the actual task, uh, the goal,
- 5:55
um, but also access to the vast troves
- 5:58
of organizational context available to
- 6:00
it. Whether that is web search, web
- 6:02
context, whether that is connectors
- 6:05
through, you know, like tools and skills
- 6:06
and MCP servers, whether it's connectors
- 6:09
through Snowflake or Databricks
- 6:10
warehouses, um, or for us, whether it's
- 6:12
access to the vast trove, like the 90%
- 6:15
of documents that are stored within
- 6:16
SharePoint, Box, Dropbox, S3, um, you
- 6:20
know, all the unstructured, uh, data
- 6:21
that's locked up within document
- 6:23
containers. Um, and so for us, you know,
- 6:25
we care a lot about how do you actually
- 6:27
build the right tools and unlock the
- 6:29
right context here.
- 6:31
We see documents as universal containers
- 6:34
for unstructured context. Uh, there's
- 6:35
over 10 trillion plus pages of human
- 6:37
native knowledge locked up within PDFs,
- 6:39
PowerPoints, Word documents, and Excel
- 6:41
sheets. And at the same time, agents are
- 6:43
also starting to generate exponentially
- 6:45
more data in terms of, you know, more
- 6:47
agent native formats in terms of
- 6:48
markdown and HTML.
- 6:52
For us, building an agent native
- 6:54
document platform contains three main
- 6:56
pieces. Uh, it contains the document
- 6:58
parsing layer, actually digitalizing,
- 7:01
um, the vast troves of PDFs,
- 7:02
PowerPoints, and Word documents into
- 7:05
kind of accurate token efficient
- 7:07
context, um, and feeding it, you know,
- 7:09
either markdown, metadata, and other
- 7:11
forms that actually enable agents to do
- 7:13
stuff over your documents.
- 7:15
The next is actually what we call the
- 7:16
semantic and storage layer, um, which is
- 7:19
kind of like this concept of document
- 7:21
management for humans and agents.
- 7:22
Instead of humans opening up Microsoft
- 7:24
Word, using your file system or file
- 7:26
storage, how do you create some sort of
- 7:28
agent interface that actually captures,
- 7:30
stores all your documents, and enables
- 7:32
agents to actually manage and operate
- 7:34
over them in various ways. Um, and then
- 7:37
there's still a room for actually kind
- 7:39
of repeatable document workflows is
- 7:41
agent layer they can actually encode
- 7:42
within the
- 7:44
uh
- 7:44
within like this type of platform, too.
- 7:46
If there is a repeatable workflow, like
- 7:48
invoice processing, KYC, or claims,
- 7:51
instead of always offloading it to a
- 7:53
generalized agent, how do you actually
- 7:54
develop some sort of specialized
- 7:56
workflow for it and really carefully
- 7:58
tune cost and accuracy to make sure that
- 8:00
you achieve your goal.
- 8:03
So, when we talk about like the concepts
- 8:05
today in in terms of these three layers,
- 8:07
uh we'll talk about this concept of
- 8:09
document OCR, which is one of the core
- 8:11
concepts of this track, um to document
- 8:13
extraction, document search, and
- 8:14
document workflows. There's still a
- 8:16
variety of topics yet left unsolved from
- 8:18
agent native document formats, document
- 8:20
versioning, document editing, you know,
- 8:22
hill climbing as a service, and just
- 8:24
like a wide set of uh kind of remaining
- 8:27
concepts to actually create a
- 8:28
comprehensive piece of software where
- 8:30
agents and humans can collaborate on
- 8:32
documents and actually do work over
- 8:33
them.
- 8:35
So, we'll start with the first section
- 8:37
um and maybe start kind of with uh the
- 8:39
document OCR RP.
- 8:41
Um first off, uh you know, document OCR
- 8:44
is hard. Um some of you might have seen
- 8:46
a blog post that we put out a few months
- 8:48
ago related to this topic. But, the
- 8:50
reason this problem even exists is an
- 8:52
agent cannot actually take the like the
- 8:54
raw PDF like file binary and make sense
- 8:57
of it. Um that's because the like the
- 8:59
way PDFs are actually stored as a format
- 9:02
is it's rendered for kind of like
- 9:03
display purposes um and not really for
- 9:06
uh you know, machine consumption in an
- 9:08
interpretable manner.
- 9:09
Um
- 9:10
they're designed for printing. So,
- 9:12
basically text are represented as almost
- 9:14
like individual glyphs with coordinates.
- 9:16
Um tables are not represented as tables,
- 9:19
they're represented as like typically
- 9:21
line segments uh drawn with various
- 9:23
types of borders, and also text drawn at
- 9:25
certain cell positions. And so, if
- 9:28
you're kind of an agent that's trying to
- 9:29
make sense of this like document, you're
- 9:31
going to have a really hard time
- 9:33
actually trying to reason about what
- 9:34
character and what shapes like map to
- 9:36
what. Um so the whole point of document
- 9:38
OCR, um, you know, it's been around for
- 9:40
like 20 plus years, is to really try to
- 9:43
create some sort of, uh, digitalized
- 9:45
well-interpretable, uh, representation
- 9:47
um, that's both interpretable to humans,
- 9:49
um, as well as AI agents.
- 9:51
Also, reading order itself, if you have
- 9:52
a multi-column layout, there is no, uh,
- 9:55
guarantee that the way it's represented
- 9:57
in the PDF, um, actually corresponds to
- 9:59
the typical ways that humans would read
- 10:01
it. Cuz again, it's basically just an
- 10:02
arbitrary sequence of characters drawn
- 10:04
with uh, coordinate positions.
- 10:07
Related to this, you know, even Word
- 10:09
doc, uh, parsing is hard. Um, they're a
- 10:11
little bit more structured than PDFs,
- 10:12
but they're in kind of like this, uh,
- 10:14
custom bespoke XML format, and this
- 10:16
applies to PowerPoints as well. Um, it
- 10:19
contains more structural information,
- 10:20
but still there's a ton of like fluff.
- 10:22
Like you don't actually need to ingest
- 10:24
all the tags to actually have the agent
- 10:26
make sense of the Word document. It
- 10:28
still needs to infer a bunch of
- 10:29
structure from it. It needs like the
- 10:31
right, uh, kind of to lift the right
- 10:33
metadata around like formatting,
- 10:35
semantics, that type of stuff. But also
- 10:37
being able to ignore the tags and
- 10:38
actually be able to render the document
- 10:40
so that the agent can see the overall
- 10:42
structure of the page. Um, it's still
- 10:44
generally hard problem, and there's a
- 10:45
lot of uplift you can get by actually
- 10:47
parsing it into a more interpretable
- 10:49
format instead of just using the native,
- 10:51
uh, OXML and feeding that to an agent.
- 10:55
Um, if you're familiar with document
- 10:56
understanding, um, you know, it's been
- 10:58
around for quite a bit of time. Um,
- 11:00
there's a lot of these like heuristic
- 11:02
and pipeline-based approaches, which,
- 11:04
uh, focused on kind of more, uh, I guess
- 11:06
like human-driven, hand-handwritten like
- 11:08
techniques to analyze like kind of
- 11:10
various pieces of text, group them into
- 11:12
clusters, and identify tables,
- 11:14
paragraphs, and be able to kind of like,
- 11:16
uh, generate some sort of output
- 11:18
representation. Um, a lot of basic
- 11:20
techniques if you use open-source
- 11:21
libraries like PyPDF, PyMuPDF, um, and
- 11:24
of course like some of the,
- 11:26
uh, more recent approaches, um,
- 11:27
basically use this type of approach.
- 11:29
There's also, of course, using a VLM to
- 11:31
one-shot a document um, into text, uh,
- 11:34
that's what we call like a vision-based
- 11:36
approach.
- 11:37
This works decently well. Obviously, you
- 11:39
know, I think there's kind of a lot of
- 11:41
uplift you can get by being able to read
- 11:42
the visual structure, but there's a lot
- 11:45
of sub-optimal pieces about it. It can
- 11:48
hallucinate on text-only pages. It costs
- 11:50
a ton of money,
- 11:52
and also it still lacks a lot of the
- 11:53
semantics and grounding that you
- 11:55
typically expect with some sort of
- 11:56
document processing tool.
- 11:58
And so for us, we really think about
- 12:00
combining both the pipeline-based
- 12:02
approaches of deeply understanding the
- 12:04
file containers and binaries with the
- 12:06
vision-based approaches to help
- 12:08
generate, you know, kind of
- 12:10
a hybrid approach that we think is at
- 12:11
the Pareto frontier of cost and
- 12:12
accuracy.
- 12:15
A quick note on this, I'll probably just
- 12:18
like skim the high level, but the Pareto
- 12:20
frontier for document OCR
- 12:23
will always be like much more accurate
- 12:24
and cheap compared to the Pareto
- 12:26
frontier for, you know, wherever the
- 12:27
frontier models are in terms of document
- 12:29
understanding. It's because it's a very
- 12:31
specific data type, and there's always
- 12:33
ways to kind of like distill the latest
- 12:35
visual understanding capabilities from,
- 12:37
you know, Gemini, GPT,
- 12:39
Opus into kind of a carefully tailored
- 12:42
workflow that's able to process your
- 12:43
documents at scale and in a highly
- 12:45
accurate manner.
- 12:47
To some extent, that's exactly what we
- 12:48
do. You know, we both optimize
- 12:50
underlying like PDF engines plus like
- 12:53
Word, PowerPoint, and others.
- 12:55
We have an agentic harness that's like
- 12:56
carefully tuned for auto routing between
- 12:59
cheaper specialized models to frontier
- 13:01
models,
- 13:02
and also kind of specialized fine-tuned
- 13:04
document VLMs that are parameter
- 13:06
efficient and focus on specific classes
- 13:08
of documents,
- 13:09
elements like tables, charts, and
- 13:11
others.
- 13:14
That's our commercial service called
- 13:16
LlamaParse, which I'll talk about in a
- 13:18
bit. But I think in general, you know,
- 13:20
we're extremely committed to advancing
- 13:22
the frontier of just document
- 13:24
understanding cuz basically if you're
- 13:25
within an enterprise organization and
- 13:27
you have a massive long tail of
- 13:29
documents across financial services,
- 13:31
insurance, manufacturing, legal,
- 13:33
government, and a bunch of others,
- 13:36
there's just a lot of complexity in a
- 13:37
lot of these document types. And if
- 13:39
you're actually trying to unlock context
- 13:41
at scale,
- 13:42
most of the models are not up for the
- 13:43
task. And so, we've created this thing
- 13:45
called ParseBench, which is a
- 13:47
comprehensive enterprise document
- 13:48
benchmark for agents. We we think about
- 13:51
it as the most comprehensive enterprise
- 13:53
document benchmark. You know, it
- 13:54
contains 2,000 human-verified pages. It
- 13:57
measures tables, charts, content
- 13:59
faithfulness, semantic formatting.
- 14:01
And it's optimized for
- 14:03
just how AI agents are actually able to
- 14:05
understand these documents instead of
- 14:07
like syntactic correctness.
- 14:09
We've If you look at parsebench.ai,
- 14:11
which is, you know, kind of the it's a
- 14:13
fully public page on the internet. It's
- 14:15
also available on Hugging Face and
- 14:17
Kaggle.
- 14:18
We benchmark probably like 50 different
- 14:20
frontier models, open weight models,
- 14:22
specialized OCR solutions. And you
- 14:24
really want the Pareto curve to kind of
- 14:26
be towards the left and up in terms of
- 14:29
accuracy, like extremely high accuracy,
- 14:31
but also extremely low cost. And there's
- 14:34
just so many different types of
- 14:35
documents, where some are a little bit
- 14:37
simpler, maybe some are a little bit
- 14:39
more complex. And ideally, you want to
- 14:41
cover all the points on the curve to
- 14:43
deal with the dynamic distribution of
- 14:45
various types of documents out there.
- 14:46
It's pretty clear, even if you look at
- 14:48
this graph, that it's
- 14:50
like you can increase the complexity of
- 14:52
the benchmark, and that document
- 14:53
understanding is definitely not a 100%
- 14:55
solved. But, you know, you fundamentally
- 14:58
need to kind of advance a lot of the
- 15:00
core capabilities to make sure they're
- 15:01
able to process, unlock the vast trove
- 15:03
of enterprise context out there.
- 15:08
There's kind of like a few different
- 15:10
points on this accuracy cost latency
- 15:12
Pareto curve. There's what I call like
- 15:14
the high accuracy regime, where like,
- 15:16
you know, some institutions basically
- 15:17
need like 99 to almost 100% accuracy,
- 15:20
because basically, the downside of an
- 15:22
incorrect extraction is you completely
- 15:23
mess up your financial model, you
- 15:25
completely, you know, you basically get
- 15:27
flagged for fraud or a bunch of other
- 15:29
really really bad things.
- 15:30
In regulated industries like insurance
- 15:32
and financial services, we see this a
- 15:34
decent amount. This typically means
- 15:36
you're willing to pay a little bit more
- 15:37
money per page for like deeper agent
- 15:39
tech reasoning to at least make sure
- 15:41
that you get back the the information in
- 15:43
the right format. There's also like the
- 15:45
low cost regime. Let's say you're just
- 15:47
trying to index, you know, the million
- 15:48
plus documents per day within your
- 15:51
that's you know, being continually
- 15:52
updated within your SharePoint just for
- 15:54
like rag knowledge base search. You
- 15:56
know, in these cases, you obviously want
- 15:57
it to be not like terribly inaccurate,
- 15:59
but even if it's a little bit messed up,
- 16:01
it's okay, too. Because in the end if
- 16:03
you have a sufficiently good agent, it
- 16:04
can always dive deeper into the document
- 16:06
and surface the right information with
- 16:08
the right citations and grounding.
- 16:10
So for these, you know, being able to
- 16:11
create some sort of scalable offline
- 16:13
indexing pipeline that has the best like
- 16:15
cost constraints
- 16:17
is something that is optimal.
- 16:19
One thing about VLMs based approaches
- 16:21
though is that, you know, they're
- 16:23
typically not very fast. And I think a
- 16:26
lot of times if you have like real-time
- 16:28
file uploads, let's say you're using
- 16:29
Cloud Co-work, you upload a thousand
- 16:31
documents and you need to process it
- 16:33
within a minute a minute, like having a
- 16:35
bunch of VLMs process that at scale is
- 16:38
really tough for basically every single
- 16:39
OCR service out there. And and to be
- 16:41
fair, that includes ours, too. I think
- 16:43
in general, there's also some sort of
- 16:46
need for an extremely low latency
- 16:48
solution so that you can actually
- 16:50
process stuff in real-time even if you
- 16:52
have like deeper VLM enabled processing
- 16:55
for kind of like deeper visual
- 16:56
inspection and analysis.
- 16:58
And so that's what I call kind of being
- 17:00
in the agent loop. So besides Lama
- 17:02
Parse, which is kind of our commercial
- 17:03
service around like document processing
- 17:05
and extraction, we also created this
- 17:07
tool called Light Parse.
- 17:09
It is surprisingly really really good. I
- 17:11
don't know if you've been following some
- 17:12
of the Twitter threads, but it is
- 17:15
Rust-based. It is the fastest
- 17:16
open-source parser out there. It is
- 17:18
completely free, um, and there is
- 17:20
basically no strings attached. I think
- 17:21
it's like MIT or Apache license. Um, and
- 17:24
it basically is the most accurate like
- 17:26
markdown parser out there that doesn't
- 17:28
use a VLM or any sort of kind of like
- 17:29
deeper model. Um, and so this is kind of
- 17:33
nice because you can use it as a default
- 17:35
in the assistive agent loop. Um, let's
- 17:38
say you're uploading a a bunch of
- 17:39
documents to Quad Code, Quad Code Work,
- 17:41
Codex, and you want it to process like a
- 17:43
thousand PDFs extremely quickly. You can
- 17:46
always do that, um, and then, you know,
- 17:48
equip a VLM-based parser like Llama
- 17:50
Parser or other frontier models as a
- 17:51
tool. So, what these agents will do is
- 17:54
they'll do like a fast pass over all the
- 17:56
documents first, uh, uh, kind of like
- 17:58
just scan through all the context
- 18:00
extremely efficiently, um, and then if
- 18:02
actually needs to dive into a page with
- 18:04
like tables, with like charts, and
- 18:06
actually needs to more deeply understand
- 18:07
the values, um, it will use a VLM-based
- 18:09
tool, slower processing, to actually
- 18:11
make sure it reads the information
- 18:12
correctly. This is available as a
- 18:14
one-click installable skill, um,
- 18:16
complements kind of any other deeper
- 18:18
VLM-based OCR tool you want to use, um,
- 18:20
and we kind of designed to make it as
- 18:22
fast as possible and also easy to plug
- 18:24
in to your favorite AI agent.
- 18:28
So, I kind of speed ran through a bunch
- 18:30
of this stuff, but basically, you know,
- 18:32
uh, we spent a bunch of time on the, uh,
- 18:35
parsing layer. There's also other, um,
- 18:37
general components around like the
- 18:38
semantic and storage layer in terms of
- 18:40
document extraction and search, and of
- 18:42
course like document workflows. Um, and
- 18:45
due to time, I'll probably kind of just,
- 18:46
uh, skip some of the, uh, unexplored
- 18:48
areas like agent native document
- 18:50
formats, hill climbing as a service, and
- 18:52
others, but I'll share the full set of
- 18:53
slides online. In terms of the semantic
- 18:56
and storage layer, you know, besides
- 18:57
document parsing, um, a lot of use cases
- 19:00
also require actually getting back, uh,
- 19:02
structured information at scale from
- 19:04
documents. Whether you're processing,
- 19:06
you know, a million invoices or expense
- 19:07
reports or receipts or claims, you need
- 19:10
to make sure that, you know, you want to
- 19:11
actually get back structured outputs
- 19:13
that you can put into a downstream
- 19:15
database or system.
- 19:16
And so, a lot of these use cases
- 19:18
basically revolve around the form of,
- 19:20
you know, how do you automate a lot of
- 19:22
workloads that humans typically do in
- 19:24
scanning a lot of paperwork and doing
- 19:26
data entry. Whether it is kind of,
- 19:28
again, invoices, claims, contracts,
- 19:30
receipts, or others.
- 19:32
We kind of created these capabilities
- 19:33
within Llama Parse as well.
- 19:35
A lot of our capabilities are actually
- 19:37
tuned towards like low cost while
- 19:39
extremely high accuracy. And you get
- 19:42
back granular citations all the way back
- 19:44
to the source document for every
- 19:45
extracted output. And of course, you can
- 19:47
run this in a pipeline at scale with
- 19:49
confidence scores
- 19:50
and also, you know, being able to
- 19:52
actually flag whether or not we're
- 19:54
certain about a certain value before
- 19:56
deciding to put it into some sort of
- 19:58
system.
- 20:00
There's also document search, which is a
- 20:02
basically expanded tool set as I
- 20:03
mentioned around retrieval, BM25, grep,
- 20:06
vector search, reading, and scrolling.
- 20:08
And so, all these capabilities are
- 20:10
available within some of our commercial
- 20:12
platforms as well as open source
- 20:13
offerings. But I also just wanted to
- 20:15
paint a picture of the general concepts
- 20:16
out there today.
- 20:18
So, I'll skip this section about kind of
- 20:19
what's next and then maybe just go all
- 20:22
the way to the end. I know I'm a little
- 20:23
bit over time. So, really appreciate you
- 20:25
all spending time today and then let me
- 20:27
just how do I get to this part really
- 20:30
quick? Oh, right.
- 20:33
I'm going to skip this piece. I'll put
- 20:35
put this online.
- 20:36
Our booth is at LG 47. If you guys are
- 20:39
interested in stopping by and we're
- 20:41
hosting a giant pickleball tournament
- 20:42
today. So, thank you for your time.
- 20:45
>> [applause]