AI Engineer World's Fair 2026

Building the Document Context Layer for AI Agents — Jerry Liu, LlamaIndex

Read the talk

Building the Document Context Layer for AI Agents

Jerry Liu explains how RAG separates into agent reasoning and document context, why reading a PDF requires reconstructing its structure, and how fast parsing plus selective visual inspection can balance accuracy, cost, and latency.

From a talk by Jerry Liu

At a glance

Ideas worth remembering

  • Modern RAG can separate agent reasoning from context access: the agent chooses queries and iterates, while the document layer supplies readable, searchable information.

  • Document parsing must recover relationships as well as characters. Positioned glyphs, drawn table borders, and multicolumn layouts do not guarantee a useful reading order or semantic structure.

  • Choose parsing effort by workload. Regulated extraction, large-scale indexing, and interactive uploads place different demands on accuracy, cost, and latency.

  • A fast collection-wide pass can precede selective VLM inspection. This gives the agent broad context before it spends more time reading particular tables or charts.

  • Structured extraction needs a path back to the source and a decision about uncertain values before they enter downstream systems.

From fixed retrieval to an agent that chooses its searches

The early RAG chatbot followed a fixed recipe: split a private document collection into chunks, embed them, store them in a vector database, retrieve the top k matches, and give those matches to an LLM to generate an answer. In Jerry Liu’s account, this was enough to make a basic application that could chat over a corpus. Every question still passed through the same sequence of steps.

Source frame: From fixed retrieval to an agent that chooses its searches
Source frame: From fixed retrieval to an agent that chooses its searches

Liu, co-founder and CEO of LlamaIndex, frames RAG in 2026 as an agent harness plus a context layer. Better tool use and agent loops make that separation useful: the harness handles reasoning and repeated actions, while the context layer gives the agent material it can read and search. LlamaIndex’s focus has consequently moved from its origins as a RAG framework toward document infrastructure for agents.

The concrete change is who chooses the search query. A fixed retrieval pipeline takes its query and returns matches. An agent can reason about which keyword or search term is likely to find the information it needs, inspect the results, and search again. Retrieval complexity moves into this loop. Even a basic search tool becomes more useful when the agent can supply a better query rather than accepting the first result set as its only context.

1:401:45
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:12 · section reference included

Context access and programs move up the stack

As long contexts and compaction improve, Liu sees more attention moving from managing window overflow toward connecting the right MCP servers, skills, and tasks. These connections determine what an agent can reach and do. The context problem expands from fitting text into a prompt to giving a reasoning system useful access to organizational information.

Source frame: Context access and programs move up the stack
Source frame: Context access and programs move up the stack

Program definition is changing alongside access. Python and TypeScript remain part of the earlier development picture, but English increasingly expresses tasks, goals, and runbooks for both engineers and nontechnical teams. The progression runs from asking a simple question, to assigning an end-to-end task, to producing a repeatable program that executes at scale. Liu’s further forecast is that longer-running agents may eventually need a goal and scoring rubric rather than a fully described task; that is a proposed direction, not a capability demonstrated here.

That makes document access a persistent bottleneck even as models improve. Useful context includes web information, tools and skills, warehouses such as Snowflake and Databricks, and files held in SharePoint, Box, Dropbox, or S3. Liu estimates that more than 10 trillion pages of human knowledge live in PDFs, PowerPoints, Word documents, and Excel sheets. The scale explains LlamaIndex’s focus: much of the information an agent needs already exists, but its container does not make it easy to consume. Agents also produce more readily readable formats such as Markdown and HTML.

3:093:14
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

3:09 · section reference included

Three layers turn files into usable work

An agent-native document platform needs three distinct capabilities. Parsing makes a file readable; storage makes it available to operate on; workflows turn recurring document tasks into repeatable execution. Keeping these concerns separate also leaves room to tune a frequent task more carefully than a general-purpose agent would.

Source frame: Three layers turn files into usable work
Source frame: Three layers turn files into usable work
  • Parsing: Convert PDFs, PowerPoints, and Word documents into accurate, token-efficient context, including Markdown and metadata. The output should preserve the information an agent needs without making it ingest all of the source format’s machinery.
  • Semantic storage: Capture and store documents behind an interface that lets agents manage and operate on them. This is document management for humans and agents, extending the role previously served by a file system and applications such as Microsoft Word.
  • Repeatable workflows: Specialize recurring jobs such as invoice processing, KYC, and claims. A purpose-built workflow gives the system a place to tune cost and accuracy against a known goal.

This platform is still incomplete. Agent-native formats, document versioning, editing, and collaboration remain open areas in Liu’s framing. He names them as work ahead; the recording develops parsing, extraction, and search rather than explaining implementations for those future capabilities.

6:526:54
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

6:52 · section reference included

A PDF draws a page; the parser reconstructs its meaning

Why does document OCR remain hard after more than 20 years? A PDF is designed for printing and display. Text can appear as individual glyphs with coordinates. A table can consist of drawn line segments and text positioned inside apparent cells. The page looks structured to a person, but those drawing instructions do not by themselves give an agent a usable table.

Source frame: A PDF draws a page; the parser reconstructs its meaning
Source frame: A PDF draws a page; the parser reconstructs its meaning

The table example exposes the reconstruction task. First, the parser encounters characters and shapes at positions. It must determine which pieces belong together and recover a representation that humans and agents can interpret. Multicolumn text introduces another problem: the stored sequence of characters does not guarantee the order a person would read the page. Recognizing characters is only part of the job; arranging them into the right relationships is what makes the result useful.

Word documents and PowerPoints offer more structural information, but feeding their native XML directly to an agent creates its own burden. Much of the tagging is unnecessary for understanding the content. A useful parser lifts relevant formatting and semantic metadata, removes irrelevant markup, and can render the page so its overall layout remains visible. More source structure helps, but it still needs to become a representation suited to the reader.

8:418:44
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

8:35 · section reference included

Combine file-aware parsing with visual models

Two approaches attack the reconstruction problem from different directions. Pipeline methods use rules and file information to group text, identify tables and paragraphs, and produce an output representation. A vision-language model, or VLM, can instead read the document visually and turn it into text in one shot. Visual reading can recover layout that a simpler pipeline misses, but Liu identifies three costs: hallucinations even on text-only pages, expensive inference, and insufficient semantics and grounding.

Source frame: Combine file-aware parsing with visual models
Source frame: Combine file-aware parsing with visual models

LlamaParse combines file-aware processing with visual understanding. Its described ingredients include optimized PDF, Word, and PowerPoint engines; an agentic harness that routes between cheaper specialized models and frontier models; and parameter-efficient document VLMs focused on particular document classes or elements such as tables and charts. The cost argument follows from that routing: use specialized processing where it suffices, and spend on stronger visual reasoning where the document requires it.

Liu describes this hybrid as lying on the cost–accuracy Pareto frontier: a set of choices where improving one objective requires giving up something on another. He also expects specialized document workflows to outperform general frontier-model use on this tradeoff. These comparative claims are his assessment; the recording does not establish a universal performance advantage across document types.

ParseBench gives this problem a broader evaluation target. Liu describes 2,000 human-verified pages and approximately 50 evaluated frontier models, open-weight models, and specialized OCR solutions. It measures tables, charts, content faithfulness, and semantic formatting, with emphasis on whether agents can understand the result rather than merely whether its syntax is correct. That choice matches the table problem: plausible-looking output is less useful if it loses the relationships needed to interpret the information.

The desired direction is lower cost and higher accuracy, but the distribution of documents complicates any single choice. Simple pages and difficult enterprise documents demand different processing. Liu’s conclusion from the benchmark is that document understanding remains unsolved; covering a range of operating points matters because a real collection contains a changing mix of document types and complexity.

11:0011:02
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

10:55 · section reference included

Read everything quickly, then inspect selected pages deeply

Latency adds a third objective. A parser suitable for an offline indexing job may make an interactive agent wait too long, while a cheap first pass may be inadequate for regulated extraction. Liu distinguishes three workloads by what a mistake or a delay costs.

Source frame: Read everything quickly, then inspect selected pages deeply
Source frame: Read everything quickly, then inspect selected pages deeply
  • High accuracy: Financial services and insurance may require 99% to almost 100% accuracy because an incorrect extraction can damage a financial model or trigger a fraud flag. Those figures describe requirements, not demonstrated achievement. Paying more per page for deeper reasoning can be justified by the consequence of an error.
  • Low cost: Indexing a million-plus documents per day from a continually updated SharePoint collection favors a scalable offline pipeline. Some imperfections can be acceptable if an agent can later revisit the source and recover the needed information with citations.
  • Low latency: Uploading a thousand documents and wanting them processed within a minute puts pressure on VLM-based processing. Liu explicitly includes his own service among those that find this scenario difficult.

LightParse addresses the fast-pass role. Liu describes a free, open-source Rust parser that produces Markdown without a VLM or deeper model, and calls it the fastest open-source parser. That speed superlative is his claim rather than a quantified comparison in the recording. Its practical role is clear: make an initial reading of many documents cheap and quick, while keeping a slower visual parser available as a tool.

Follow the thousand-PDF example through the change in processing. The agent first gets a fast pass over the whole collection and scans the resulting context. If its task then requires understanding values on a page containing a table or chart, it calls a VLM-based tool for that page. The collection becomes broadly readable before the system spends time on deeper visual interpretation. The table that began as positioned characters and lines now receives the additional inspection needed to read its values correctly. This is a proposed operating pattern, not a timed demonstration of processing a thousand PDFs within a minute.

Where does the slower visual work enter the agent loop? The flow below shows it after the collection-wide pass, on a branch selected by the task. The important relationship is the scope of work: the fast parser covers all documents, while the visual tool handles the pages that need deeper interpretation. Liu describes LightParse as an installable skill that complements other VLM-based OCR tools.

How it fits togetherA fast collection pass with selective visual inspection

The example contains a thousand PDFs.

The agent scans all documents first, then uses slower visual processing when its task requires deeper understanding of a table or chart.

15:0815:10
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

15:08 · section reference included

Read the complete timestamped transcript
  1. 0:01

    [music]

  2. 0:12

    >> I think we can get started. Hey

  3. 0:13

    everyone, I'm Jerry, co-founder and CEO

  4. 0:15

    of LlamaIndex, and today I'm excited to

  5. 0:18

    uh give a talk called building the

  6. 0:19

    document context layer for AI agents. Um

  7. 0:22

    really uh big shoutout to the AI

  8. 0:24

    Engineer World Fair for hosting. Um and

  9. 0:26

    if you've seen some of my earlier talks

  10. 0:28

    from the previous uh AI Engineer

  11. 0:30

    conferences, uh we've kind of traced

  12. 0:31

    through a lot of the evolution of how,

  13. 0:33

    you know, um agent advances have

  14. 0:35

    correlated with uh you know, how you

  15. 0:36

    inject context into evolving

  16. 0:38

    applications. Um and so today, in 2026,

  17. 0:40

    we'll kind of talk about three main

  18. 0:42

    topics. One is, you know, what is RAG in

  19. 0:44

    2026? Um two, uh basically that

  20. 0:47

    basically decomposes into an agent

  21. 0:49

    harness plus a context layer. So kind of

  22. 0:51

    going into the modern document context

  23. 0:53

    layer for agents, um and giving you a

  24. 0:54

    little bit of sense of kind of some of

  25. 0:56

    the core capabilities today, um

  26. 0:57

    especially uh unlocking, you know, the

  27. 0:59

    vast trove of document-based data today,

  28. 1:01

    and giving you a glimpse of what's next.

  29. 1:04

    All right. RAG in 2026. But before that,

  30. 1:06

    just a little bit about the company, and

  31. 1:07

    I'll kind of get on with the main talk.

  32. 1:09

    Um you might have seen us as a RAG

  33. 1:10

    framework. We started in 2023, got

  34. 1:12

    pretty popular, um kind of created a lot

  35. 1:14

    of techniques around like advanced RAG,

  36. 1:16

    that type of stuff. Um today, we're

  37. 1:18

    basically the main document

  38. 1:19

    infrastructure for AI agents. Um we

  39. 1:21

    deliver the best platform for agents to

  40. 1:23

    actually read and operate over

  41. 1:24

    documents, and we see ourselves as

  42. 1:26

    unlocking basically the vast troves of

  43. 1:27

    document context even as agents and

  44. 1:29

    models themselves get better.

  45. 1:33

    All right. So, what does RAG actually

  46. 1:35

    mean in 2026, and how are agents

  47. 1:37

    actually inhaling and addressing context

  48. 1:39

    today?

  49. 1:40

    This is a snapshot from um actually 3

  50. 1:43

    years ago, um when I first gave a talk

  51. 1:45

    on this, and naive RAG was basically

  52. 1:46

    just building a simple chatbot over your

  53. 1:48

    private corpus of data. If you flash

  54. 1:50

    back all the way to January of like

  55. 1:51

    2023, um you know, this technique

  56. 1:54

    basically consisted of you have some

  57. 1:56

    corpus of documents, um you chunk it up,

  58. 1:58

    embed it, and put it into a vector

  59. 2:00

    database. You do some naive top K

  60. 2:02

    retrieval, and then you generate some

  61. 2:03

    stuff with an LLM. All the steps are

  62. 2:05

    fixed. You use kind of like a fixed set

  63. 2:07

    of techniques, and then with that you

  64. 2:09

    actually get some basic application

  65. 2:10

    results by just being able to chat over

  66. 2:12

    a corpus of documents.

  67. 2:14

    Of course, the agent landscape has

  68. 2:16

    changed quite a bit and quite

  69. 2:18

    dramatically since then.

  70. 2:19

    One is agent loops and tool use have

  71. 2:22

    gotten a lot a lot better, especially as

  72. 2:24

    both the models and agent harnesses have

  73. 2:26

    increased. And there's basically a

  74. 2:28

    cleaner separation now between agent

  75. 2:30

    reasoning and how they actually interact

  76. 2:31

    with context. If you look at what I call

  77. 2:34

    kind of the modern generalized agents,

  78. 2:36

    which includes, you know, all your

  79. 2:37

    favorite applications and tools out

  80. 2:38

    there from Claude Code, Claude Code

  81. 2:40

    Work, Open Claw, Codex, and a few

  82. 2:42

    others,

  83. 2:43

    you know, the retrieval complexity has

  84. 2:45

    started to get baked into the agent

  85. 2:47

    layer. So, instead of coming up with a

  86. 2:49

    variety of hacks to really work around

  87. 2:51

    the limitations of naive top K

  88. 2:52

    retrieval, you can start getting the

  89. 2:54

    agent to really reason about, for

  90. 2:56

    instance, the best key best keyword to

  91. 2:57

    search for, best like, you know, search

  92. 2:59

    term to actually get back good results.

  93. 3:01

    Even if the retrieval tools are basic,

  94. 3:03

    the agent can input the right queries to

  95. 3:05

    basically loop upon itself and help

  96. 3:07

    solve the task at hand.

  97. 3:09

    Number two, context is moving up the

  98. 3:11

    stack. There's been a lot of

  99. 3:13

    conversations, I think, for the first

  100. 3:14

    two and a half years of Gen AI of how do

  101. 3:17

    you actually manage the agent context

  102. 3:18

    window, make sure it doesn't overflow,

  103. 3:20

    that type of thing. But, I think as, you

  104. 3:22

    know, a lot of the evolving techniques,

  105. 3:24

    compaction, long contexts have evolved,

  106. 3:27

    I think more and more of the

  107. 3:28

    conversation is actually how do you just

  108. 3:30

    hook up the right MCP servers and skills

  109. 3:32

    and tasks to the agent to enable it to

  110. 3:35

    do various types of tasks.

  111. 3:37

    And so, that also applies to agent

  112. 3:38

    orchestration, right? Coding agents have

  113. 3:40

    gotten better.

  114. 3:42

    Abstractions of actually defining what

  115. 3:44

    types of tasks and programs you want to

  116. 3:45

    build have moved a little bit upwards

  117. 3:47

    towards English as opposed to through

  118. 3:49

    code. And so, maybe in 2023 to 2025, you

  119. 3:52

    define programs and you still build

  120. 3:54

    stuff via like importing Python, using

  121. 3:56

    TypeScript through code. It's pretty

  122. 3:58

    clear these days more and more people

  123. 4:00

    are just building stuff using English,

  124. 4:02

    whether you're a software engineer or

  125. 4:04

    you're a non-technical function like

  126. 4:06

    go-to-market marketing. And you're

  127. 4:08

    defining runbooks through English and

  128. 4:10

    kind of like defining the right goals

  129. 4:12

    and making sure the AI is aligned on the

  130. 4:14

    right task. And we see that carrying

  131. 4:16

    over going forward as well.

  132. 4:19

    So, in terms of just like general

  133. 4:21

    evolving knowledge work patterns, and

  134. 4:22

    this kind of forms a foundation for what

  135. 4:24

    we how we think about the context layer.

  136. 4:26

    You know,

  137. 4:27

    in the past maybe you could use AI to

  138. 4:29

    just like ask simple questions, get it

  139. 4:31

    resolved.

  140. 4:32

    Today, you're starting to be able to

  141. 4:34

    define and solve more end-to-end like

  142. 4:36

    tasks just through English and also use

  143. 4:39

    English to start

  144. 4:40

    compiling repeatable programs to

  145. 4:43

    actually execute something at scale.

  146. 4:45

    Obviously, there's been a lot of

  147. 4:46

    discussion on how this will evolve even

  148. 4:48

    more, and we do see AI agents as

  149. 4:50

    behaving even more autonomously through

  150. 4:53

    solving long horizon tasks, looping upon

  151. 4:55

    themselves to be able to achieve a goal.

  152. 4:57

    So, the future will move towards a state

  153. 4:59

    where you actually might not even have

  154. 5:00

    to define the task in English, but

  155. 5:02

    actually more of the goal and a scoring

  156. 5:04

    rubric, and then the agent will use all

  157. 5:06

    available context available to it to

  158. 5:08

    actually solve the task.

  159. 5:10

    So, with that overall framing, let's

  160. 5:13

    talk about what kind of what is a pretty

  161. 5:15

    core focus for us as a company, which is

  162. 5:17

    basically helping to unlock unstructured

  163. 5:19

    context to basically feed into these

  164. 5:21

    agents even as they get really, really

  165. 5:23

    good both in terms of the core model

  166. 5:24

    capabilities as well as the agent

  167. 5:26

    harness. And we'll focus a lot on our

  168. 5:28

    core focus area, which is document

  169. 5:31

    context, but of course there's plenty of

  170. 5:32

    other sources of context as well. In

  171. 5:34

    fact, we do think context really is

  172. 5:37

    everything.

  173. 5:38

    In the end, you could have a infinitely

  174. 5:40

    smart agents, but your ability to

  175. 5:42

    actually get value out of this

  176. 5:44

    infinitely AI agent um, able to do more

  177. 5:47

    and more stuff end to end, is actually,

  178. 5:49

    you know, giving it the right things to

  179. 5:51

    do. Um, and the right things to do

  180. 5:52

    include the actual task, uh, the goal,

  181. 5:55

    um, but also access to the vast troves

  182. 5:58

    of organizational context available to

  183. 6:00

    it. Whether that is web search, web

  184. 6:02

    context, whether that is connectors

  185. 6:05

    through, you know, like tools and skills

  186. 6:06

    and MCP servers, whether it's connectors

  187. 6:09

    through Snowflake or Databricks

  188. 6:10

    warehouses, um, or for us, whether it's

  189. 6:12

    access to the vast trove, like the 90%

  190. 6:15

    of documents that are stored within

  191. 6:16

    SharePoint, Box, Dropbox, S3, um, you

  192. 6:20

    know, all the unstructured, uh, data

  193. 6:21

    that's locked up within document

  194. 6:23

    containers. Um, and so for us, you know,

  195. 6:25

    we care a lot about how do you actually

  196. 6:27

    build the right tools and unlock the

  197. 6:29

    right context here.

  198. 6:31

    We see documents as universal containers

  199. 6:34

    for unstructured context. Uh, there's

  200. 6:35

    over 10 trillion plus pages of human

  201. 6:37

    native knowledge locked up within PDFs,

  202. 6:39

    PowerPoints, Word documents, and Excel

  203. 6:41

    sheets. And at the same time, agents are

  204. 6:43

    also starting to generate exponentially

  205. 6:45

    more data in terms of, you know, more

  206. 6:47

    agent native formats in terms of

  207. 6:48

    markdown and HTML.

  208. 6:52

    For us, building an agent native

  209. 6:54

    document platform contains three main

  210. 6:56

    pieces. Uh, it contains the document

  211. 6:58

    parsing layer, actually digitalizing,

  212. 7:01

    um, the vast troves of PDFs,

  213. 7:02

    PowerPoints, and Word documents into

  214. 7:05

    kind of accurate token efficient

  215. 7:07

    context, um, and feeding it, you know,

  216. 7:09

    either markdown, metadata, and other

  217. 7:11

    forms that actually enable agents to do

  218. 7:13

    stuff over your documents.

  219. 7:15

    The next is actually what we call the

  220. 7:16

    semantic and storage layer, um, which is

  221. 7:19

    kind of like this concept of document

  222. 7:21

    management for humans and agents.

  223. 7:22

    Instead of humans opening up Microsoft

  224. 7:24

    Word, using your file system or file

  225. 7:26

    storage, how do you create some sort of

  226. 7:28

    agent interface that actually captures,

  227. 7:30

    stores all your documents, and enables

  228. 7:32

    agents to actually manage and operate

  229. 7:34

    over them in various ways. Um, and then

  230. 7:37

    there's still a room for actually kind

  231. 7:39

    of repeatable document workflows is

  232. 7:41

    agent layer they can actually encode

  233. 7:42

    within the

  234. 7:44

    uh

  235. 7:44

    within like this type of platform, too.

  236. 7:46

    If there is a repeatable workflow, like

  237. 7:48

    invoice processing, KYC, or claims,

  238. 7:51

    instead of always offloading it to a

  239. 7:53

    generalized agent, how do you actually

  240. 7:54

    develop some sort of specialized

  241. 7:56

    workflow for it and really carefully

  242. 7:58

    tune cost and accuracy to make sure that

  243. 8:00

    you achieve your goal.

  244. 8:03

    So, when we talk about like the concepts

  245. 8:05

    today in in terms of these three layers,

  246. 8:07

    uh we'll talk about this concept of

  247. 8:09

    document OCR, which is one of the core

  248. 8:11

    concepts of this track, um to document

  249. 8:13

    extraction, document search, and

  250. 8:14

    document workflows. There's still a

  251. 8:16

    variety of topics yet left unsolved from

  252. 8:18

    agent native document formats, document

  253. 8:20

    versioning, document editing, you know,

  254. 8:22

    hill climbing as a service, and just

  255. 8:24

    like a wide set of uh kind of remaining

  256. 8:27

    concepts to actually create a

  257. 8:28

    comprehensive piece of software where

  258. 8:30

    agents and humans can collaborate on

  259. 8:32

    documents and actually do work over

  260. 8:33

    them.

  261. 8:35

    So, we'll start with the first section

  262. 8:37

    um and maybe start kind of with uh the

  263. 8:39

    document OCR RP.

  264. 8:41

    Um first off, uh you know, document OCR

  265. 8:44

    is hard. Um some of you might have seen

  266. 8:46

    a blog post that we put out a few months

  267. 8:48

    ago related to this topic. But, the

  268. 8:50

    reason this problem even exists is an

  269. 8:52

    agent cannot actually take the like the

  270. 8:54

    raw PDF like file binary and make sense

  271. 8:57

    of it. Um that's because the like the

  272. 8:59

    way PDFs are actually stored as a format

  273. 9:02

    is it's rendered for kind of like

  274. 9:03

    display purposes um and not really for

  275. 9:06

    uh you know, machine consumption in an

  276. 9:08

    interpretable manner.

  277. 9:09

    Um

  278. 9:10

    they're designed for printing. So,

  279. 9:12

    basically text are represented as almost

  280. 9:14

    like individual glyphs with coordinates.

  281. 9:16

    Um tables are not represented as tables,

  282. 9:19

    they're represented as like typically

  283. 9:21

    line segments uh drawn with various

  284. 9:23

    types of borders, and also text drawn at

  285. 9:25

    certain cell positions. And so, if

  286. 9:28

    you're kind of an agent that's trying to

  287. 9:29

    make sense of this like document, you're

  288. 9:31

    going to have a really hard time

  289. 9:33

    actually trying to reason about what

  290. 9:34

    character and what shapes like map to

  291. 9:36

    what. Um so the whole point of document

  292. 9:38

    OCR, um, you know, it's been around for

  293. 9:40

    like 20 plus years, is to really try to

  294. 9:43

    create some sort of, uh, digitalized

  295. 9:45

    well-interpretable, uh, representation

  296. 9:47

    um, that's both interpretable to humans,

  297. 9:49

    um, as well as AI agents.

  298. 9:51

    Also, reading order itself, if you have

  299. 9:52

    a multi-column layout, there is no, uh,

  300. 9:55

    guarantee that the way it's represented

  301. 9:57

    in the PDF, um, actually corresponds to

  302. 9:59

    the typical ways that humans would read

  303. 10:01

    it. Cuz again, it's basically just an

  304. 10:02

    arbitrary sequence of characters drawn

  305. 10:04

    with uh, coordinate positions.

  306. 10:07

    Related to this, you know, even Word

  307. 10:09

    doc, uh, parsing is hard. Um, they're a

  308. 10:11

    little bit more structured than PDFs,

  309. 10:12

    but they're in kind of like this, uh,

  310. 10:14

    custom bespoke XML format, and this

  311. 10:16

    applies to PowerPoints as well. Um, it

  312. 10:19

    contains more structural information,

  313. 10:20

    but still there's a ton of like fluff.

  314. 10:22

    Like you don't actually need to ingest

  315. 10:24

    all the tags to actually have the agent

  316. 10:26

    make sense of the Word document. It

  317. 10:28

    still needs to infer a bunch of

  318. 10:29

    structure from it. It needs like the

  319. 10:31

    right, uh, kind of to lift the right

  320. 10:33

    metadata around like formatting,

  321. 10:35

    semantics, that type of stuff. But also

  322. 10:37

    being able to ignore the tags and

  323. 10:38

    actually be able to render the document

  324. 10:40

    so that the agent can see the overall

  325. 10:42

    structure of the page. Um, it's still

  326. 10:44

    generally hard problem, and there's a

  327. 10:45

    lot of uplift you can get by actually

  328. 10:47

    parsing it into a more interpretable

  329. 10:49

    format instead of just using the native,

  330. 10:51

    uh, OXML and feeding that to an agent.

  331. 10:55

    Um, if you're familiar with document

  332. 10:56

    understanding, um, you know, it's been

  333. 10:58

    around for quite a bit of time. Um,

  334. 11:00

    there's a lot of these like heuristic

  335. 11:02

    and pipeline-based approaches, which,

  336. 11:04

    uh, focused on kind of more, uh, I guess

  337. 11:06

    like human-driven, hand-handwritten like

  338. 11:08

    techniques to analyze like kind of

  339. 11:10

    various pieces of text, group them into

  340. 11:12

    clusters, and identify tables,

  341. 11:14

    paragraphs, and be able to kind of like,

  342. 11:16

    uh, generate some sort of output

  343. 11:18

    representation. Um, a lot of basic

  344. 11:20

    techniques if you use open-source

  345. 11:21

    libraries like PyPDF, PyMuPDF, um, and

  346. 11:24

    of course like some of the,

  347. 11:26

    uh, more recent approaches, um,

  348. 11:27

    basically use this type of approach.

  349. 11:29

    There's also, of course, using a VLM to

  350. 11:31

    one-shot a document um, into text, uh,

  351. 11:34

    that's what we call like a vision-based

  352. 11:36

    approach.

  353. 11:37

    This works decently well. Obviously, you

  354. 11:39

    know, I think there's kind of a lot of

  355. 11:41

    uplift you can get by being able to read

  356. 11:42

    the visual structure, but there's a lot

  357. 11:45

    of sub-optimal pieces about it. It can

  358. 11:48

    hallucinate on text-only pages. It costs

  359. 11:50

    a ton of money,

  360. 11:52

    and also it still lacks a lot of the

  361. 11:53

    semantics and grounding that you

  362. 11:55

    typically expect with some sort of

  363. 11:56

    document processing tool.

  364. 11:58

    And so for us, we really think about

  365. 12:00

    combining both the pipeline-based

  366. 12:02

    approaches of deeply understanding the

  367. 12:04

    file containers and binaries with the

  368. 12:06

    vision-based approaches to help

  369. 12:08

    generate, you know, kind of

  370. 12:10

    a hybrid approach that we think is at

  371. 12:11

    the Pareto frontier of cost and

  372. 12:12

    accuracy.

  373. 12:15

    A quick note on this, I'll probably just

  374. 12:18

    like skim the high level, but the Pareto

  375. 12:20

    frontier for document OCR

  376. 12:23

    will always be like much more accurate

  377. 12:24

    and cheap compared to the Pareto

  378. 12:26

    frontier for, you know, wherever the

  379. 12:27

    frontier models are in terms of document

  380. 12:29

    understanding. It's because it's a very

  381. 12:31

    specific data type, and there's always

  382. 12:33

    ways to kind of like distill the latest

  383. 12:35

    visual understanding capabilities from,

  384. 12:37

    you know, Gemini, GPT,

  385. 12:39

    Opus into kind of a carefully tailored

  386. 12:42

    workflow that's able to process your

  387. 12:43

    documents at scale and in a highly

  388. 12:45

    accurate manner.

  389. 12:47

    To some extent, that's exactly what we

  390. 12:48

    do. You know, we both optimize

  391. 12:50

    underlying like PDF engines plus like

  392. 12:53

    Word, PowerPoint, and others.

  393. 12:55

    We have an agentic harness that's like

  394. 12:56

    carefully tuned for auto routing between

  395. 12:59

    cheaper specialized models to frontier

  396. 13:01

    models,

  397. 13:02

    and also kind of specialized fine-tuned

  398. 13:04

    document VLMs that are parameter

  399. 13:06

    efficient and focus on specific classes

  400. 13:08

    of documents,

  401. 13:09

    elements like tables, charts, and

  402. 13:11

    others.

  403. 13:14

    That's our commercial service called

  404. 13:16

    LlamaParse, which I'll talk about in a

  405. 13:18

    bit. But I think in general, you know,

  406. 13:20

    we're extremely committed to advancing

  407. 13:22

    the frontier of just document

  408. 13:24

    understanding cuz basically if you're

  409. 13:25

    within an enterprise organization and

  410. 13:27

    you have a massive long tail of

  411. 13:29

    documents across financial services,

  412. 13:31

    insurance, manufacturing, legal,

  413. 13:33

    government, and a bunch of others,

  414. 13:36

    there's just a lot of complexity in a

  415. 13:37

    lot of these document types. And if

  416. 13:39

    you're actually trying to unlock context

  417. 13:41

    at scale,

  418. 13:42

    most of the models are not up for the

  419. 13:43

    task. And so, we've created this thing

  420. 13:45

    called ParseBench, which is a

  421. 13:47

    comprehensive enterprise document

  422. 13:48

    benchmark for agents. We we think about

  423. 13:51

    it as the most comprehensive enterprise

  424. 13:53

    document benchmark. You know, it

  425. 13:54

    contains 2,000 human-verified pages. It

  426. 13:57

    measures tables, charts, content

  427. 13:59

    faithfulness, semantic formatting.

  428. 14:01

    And it's optimized for

  429. 14:03

    just how AI agents are actually able to

  430. 14:05

    understand these documents instead of

  431. 14:07

    like syntactic correctness.

  432. 14:09

    We've If you look at parsebench.ai,

  433. 14:11

    which is, you know, kind of the it's a

  434. 14:13

    fully public page on the internet. It's

  435. 14:15

    also available on Hugging Face and

  436. 14:17

    Kaggle.

  437. 14:18

    We benchmark probably like 50 different

  438. 14:20

    frontier models, open weight models,

  439. 14:22

    specialized OCR solutions. And you

  440. 14:24

    really want the Pareto curve to kind of

  441. 14:26

    be towards the left and up in terms of

  442. 14:29

    accuracy, like extremely high accuracy,

  443. 14:31

    but also extremely low cost. And there's

  444. 14:34

    just so many different types of

  445. 14:35

    documents, where some are a little bit

  446. 14:37

    simpler, maybe some are a little bit

  447. 14:39

    more complex. And ideally, you want to

  448. 14:41

    cover all the points on the curve to

  449. 14:43

    deal with the dynamic distribution of

  450. 14:45

    various types of documents out there.

  451. 14:46

    It's pretty clear, even if you look at

  452. 14:48

    this graph, that it's

  453. 14:50

    like you can increase the complexity of

  454. 14:52

    the benchmark, and that document

  455. 14:53

    understanding is definitely not a 100%

  456. 14:55

    solved. But, you know, you fundamentally

  457. 14:58

    need to kind of advance a lot of the

  458. 15:00

    core capabilities to make sure they're

  459. 15:01

    able to process, unlock the vast trove

  460. 15:03

    of enterprise context out there.

  461. 15:08

    There's kind of like a few different

  462. 15:10

    points on this accuracy cost latency

  463. 15:12

    Pareto curve. There's what I call like

  464. 15:14

    the high accuracy regime, where like,

  465. 15:16

    you know, some institutions basically

  466. 15:17

    need like 99 to almost 100% accuracy,

  467. 15:20

    because basically, the downside of an

  468. 15:22

    incorrect extraction is you completely

  469. 15:23

    mess up your financial model, you

  470. 15:25

    completely, you know, you basically get

  471. 15:27

    flagged for fraud or a bunch of other

  472. 15:29

    really really bad things.

  473. 15:30

    In regulated industries like insurance

  474. 15:32

    and financial services, we see this a

  475. 15:34

    decent amount. This typically means

  476. 15:36

    you're willing to pay a little bit more

  477. 15:37

    money per page for like deeper agent

  478. 15:39

    tech reasoning to at least make sure

  479. 15:41

    that you get back the the information in

  480. 15:43

    the right format. There's also like the

  481. 15:45

    low cost regime. Let's say you're just

  482. 15:47

    trying to index, you know, the million

  483. 15:48

    plus documents per day within your

  484. 15:51

    that's you know, being continually

  485. 15:52

    updated within your SharePoint just for

  486. 15:54

    like rag knowledge base search. You

  487. 15:56

    know, in these cases, you obviously want

  488. 15:57

    it to be not like terribly inaccurate,

  489. 15:59

    but even if it's a little bit messed up,

  490. 16:01

    it's okay, too. Because in the end if

  491. 16:03

    you have a sufficiently good agent, it

  492. 16:04

    can always dive deeper into the document

  493. 16:06

    and surface the right information with

  494. 16:08

    the right citations and grounding.

  495. 16:10

    So for these, you know, being able to

  496. 16:11

    create some sort of scalable offline

  497. 16:13

    indexing pipeline that has the best like

  498. 16:15

    cost constraints

  499. 16:17

    is something that is optimal.

  500. 16:19

    One thing about VLMs based approaches

  501. 16:21

    though is that, you know, they're

  502. 16:23

    typically not very fast. And I think a

  503. 16:26

    lot of times if you have like real-time

  504. 16:28

    file uploads, let's say you're using

  505. 16:29

    Cloud Co-work, you upload a thousand

  506. 16:31

    documents and you need to process it

  507. 16:33

    within a minute a minute, like having a

  508. 16:35

    bunch of VLMs process that at scale is

  509. 16:38

    really tough for basically every single

  510. 16:39

    OCR service out there. And and to be

  511. 16:41

    fair, that includes ours, too. I think

  512. 16:43

    in general, there's also some sort of

  513. 16:46

    need for an extremely low latency

  514. 16:48

    solution so that you can actually

  515. 16:50

    process stuff in real-time even if you

  516. 16:52

    have like deeper VLM enabled processing

  517. 16:55

    for kind of like deeper visual

  518. 16:56

    inspection and analysis.

  519. 16:58

    And so that's what I call kind of being

  520. 17:00

    in the agent loop. So besides Lama

  521. 17:02

    Parse, which is kind of our commercial

  522. 17:03

    service around like document processing

  523. 17:05

    and extraction, we also created this

  524. 17:07

    tool called Light Parse.

  525. 17:09

    It is surprisingly really really good. I

  526. 17:11

    don't know if you've been following some

  527. 17:12

    of the Twitter threads, but it is

  528. 17:15

    Rust-based. It is the fastest

  529. 17:16

    open-source parser out there. It is

  530. 17:18

    completely free, um, and there is

  531. 17:20

    basically no strings attached. I think

  532. 17:21

    it's like MIT or Apache license. Um, and

  533. 17:24

    it basically is the most accurate like

  534. 17:26

    markdown parser out there that doesn't

  535. 17:28

    use a VLM or any sort of kind of like

  536. 17:29

    deeper model. Um, and so this is kind of

  537. 17:33

    nice because you can use it as a default

  538. 17:35

    in the assistive agent loop. Um, let's

  539. 17:38

    say you're uploading a a bunch of

  540. 17:39

    documents to Quad Code, Quad Code Work,

  541. 17:41

    Codex, and you want it to process like a

  542. 17:43

    thousand PDFs extremely quickly. You can

  543. 17:46

    always do that, um, and then, you know,

  544. 17:48

    equip a VLM-based parser like Llama

  545. 17:50

    Parser or other frontier models as a

  546. 17:51

    tool. So, what these agents will do is

  547. 17:54

    they'll do like a fast pass over all the

  548. 17:56

    documents first, uh, uh, kind of like

  549. 17:58

    just scan through all the context

  550. 18:00

    extremely efficiently, um, and then if

  551. 18:02

    actually needs to dive into a page with

  552. 18:04

    like tables, with like charts, and

  553. 18:06

    actually needs to more deeply understand

  554. 18:07

    the values, um, it will use a VLM-based

  555. 18:09

    tool, slower processing, to actually

  556. 18:11

    make sure it reads the information

  557. 18:12

    correctly. This is available as a

  558. 18:14

    one-click installable skill, um,

  559. 18:16

    complements kind of any other deeper

  560. 18:18

    VLM-based OCR tool you want to use, um,

  561. 18:20

    and we kind of designed to make it as

  562. 18:22

    fast as possible and also easy to plug

  563. 18:24

    in to your favorite AI agent.

  564. 18:28

    So, I kind of speed ran through a bunch

  565. 18:30

    of this stuff, but basically, you know,

  566. 18:32

    uh, we spent a bunch of time on the, uh,

  567. 18:35

    parsing layer. There's also other, um,

  568. 18:37

    general components around like the

  569. 18:38

    semantic and storage layer in terms of

  570. 18:40

    document extraction and search, and of

  571. 18:42

    course like document workflows. Um, and

  572. 18:45

    due to time, I'll probably kind of just,

  573. 18:46

    uh, skip some of the, uh, unexplored

  574. 18:48

    areas like agent native document

  575. 18:50

    formats, hill climbing as a service, and

  576. 18:52

    others, but I'll share the full set of

  577. 18:53

    slides online. In terms of the semantic

  578. 18:56

    and storage layer, you know, besides

  579. 18:57

    document parsing, um, a lot of use cases

  580. 19:00

    also require actually getting back, uh,

  581. 19:02

    structured information at scale from

  582. 19:04

    documents. Whether you're processing,

  583. 19:06

    you know, a million invoices or expense

  584. 19:07

    reports or receipts or claims, you need

  585. 19:10

    to make sure that, you know, you want to

  586. 19:11

    actually get back structured outputs

  587. 19:13

    that you can put into a downstream

  588. 19:15

    database or system.

  589. 19:16

    And so, a lot of these use cases

  590. 19:18

    basically revolve around the form of,

  591. 19:20

    you know, how do you automate a lot of

  592. 19:22

    workloads that humans typically do in

  593. 19:24

    scanning a lot of paperwork and doing

  594. 19:26

    data entry. Whether it is kind of,

  595. 19:28

    again, invoices, claims, contracts,

  596. 19:30

    receipts, or others.

  597. 19:32

    We kind of created these capabilities

  598. 19:33

    within Llama Parse as well.

  599. 19:35

    A lot of our capabilities are actually

  600. 19:37

    tuned towards like low cost while

  601. 19:39

    extremely high accuracy. And you get

  602. 19:42

    back granular citations all the way back

  603. 19:44

    to the source document for every

  604. 19:45

    extracted output. And of course, you can

  605. 19:47

    run this in a pipeline at scale with

  606. 19:49

    confidence scores

  607. 19:50

    and also, you know, being able to

  608. 19:52

    actually flag whether or not we're

  609. 19:54

    certain about a certain value before

  610. 19:56

    deciding to put it into some sort of

  611. 19:58

    system.

  612. 20:00

    There's also document search, which is a

  613. 20:02

    basically expanded tool set as I

  614. 20:03

    mentioned around retrieval, BM25, grep,

  615. 20:06

    vector search, reading, and scrolling.

  616. 20:08

    And so, all these capabilities are

  617. 20:10

    available within some of our commercial

  618. 20:12

    platforms as well as open source

  619. 20:13

    offerings. But I also just wanted to

  620. 20:15

    paint a picture of the general concepts

  621. 20:16

    out there today.

  622. 20:18

    So, I'll skip this section about kind of

  623. 20:19

    what's next and then maybe just go all

  624. 20:22

    the way to the end. I know I'm a little

  625. 20:23

    bit over time. So, really appreciate you

  626. 20:25

    all spending time today and then let me

  627. 20:27

    just how do I get to this part really

  628. 20:30

    quick? Oh, right.

  629. 20:33

    I'm going to skip this piece. I'll put

  630. 20:35

    put this online.

  631. 20:36

    Our booth is at LG 47. If you guys are

  632. 20:39

    interested in stopping by and we're

  633. 20:41

    hosting a giant pickleball tournament

  634. 20:42

    today. So, thank you for your time.

  635. 20:45

    >> [applause]