AI Engineer World's Fair 2026

From Ingestion to Agents: How AI Teams Build on Document Intelligence — Adit Abraham, Reducto

Read the talk

From Ingestion to Agents: How AI Teams Build on Document Intelligence

Adit Abraham of Reducto explains how to turn visually encoded documents into useful agent inputs: combine specialized models, correct OCR without rewriting the source, separate retrieval from reasoning, and give difficult extraction tasks tools and a verification loop.

From a talk by Adit Abraham

At a glance

Ideas worth remembering

  • Combine efficient layout detection with VLM interpretation where visual complexity requires it; large models need not process every part of every page.

  • OCR correction should preserve what the document says, including its mistakes. Targeted token edits reduce the opportunity for a model to rewrite the underlying facts.

  • Format data for each consumer: natural-language table renderings for retrieval, and structure-preserving representations for reasoning.

  • Tool use and repeated verification can make difficult extraction tasks tractable, but extraction quality must include both correctness and completeness.

  • Evaluate individual stages and completed agent work, using production monitoring as well as test datasets. Agents choosing tools still need good inputs, relevant context, and usable outputs.

When documents become inputs to decisions

A bad document parse can spoil a chatbot answer. In an agent workflow, the same mistake can travel through several decisions and into the finished work. That is the practical problem behind Adit Abraham’s talk. As Reducto’s co-founder and CEO, he draws on processing many billions of customer documents, including historical material that enterprises have struggled to use beyond a demo.

Source frame: When documents become inputs to decisions
Source frame: When documents become inputs to decisions

The earlier retrieval-augmented generation pattern mostly synthesized information: retrieve context, then answer a question through enterprise search or a chatbot. Agents expand the job. They make decisions, produce work, and generate or modify documents. Input quality matters throughout that longer sequence because each step can carry forward an earlier misunderstanding while adding information from another source.

Enterprise information also arrives without a tidy inventory. Different teams put files in Google Drive, Box, and other repositories; the contents are scattered, unstructured, and multimodal. A useful document system therefore has several jobs: recover the contents accurately, retrieve the relevant context, and support the interactions or modifications the task requires. Parsing is the first problem, but it does not settle the others.

0:210:57
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:13 · section reference included

PDFs preserve appearance; agents need meaning

PDFs were designed to reproduce a document’s appearance faithfully and make it printable. The present task asks for something different: a representation an agent can reason over. Abraham introduces the difficulty with a document-decision benchmark on which a frontier model scored about 30%. That figure describes the benchmark he cites, rather than a general PDF accuracy rate.

Source frame: PDFs preserve appearance; agents need meaning
Source frame: PDFs preserve appearance; agents need meaning

Humans encode meaning visually. A financial deck’s author is thinking about readers, rather than whether an agent can consume the slide. Abraham’s memorable example is SoftBank’s imagery of a goose laying eggs: a picture can carry part of the explanation. More routinely, merged cells establish relationships in tables, lines encode numerical trends, and handwriting holds information that even a human may struggle to read. Extracting characters alone can leave those relationships behind.

5:285:33
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

5:28 · section reference included

Use vision models selectively, then correct the transcription

A traditional document pipeline ran OCR and then processed the resulting text. That approach could work well when the layout stayed consistent: a W-2 presents a constrained problem that templates can address. Vision-language models (VLMs) widened the possibilities by reading visual content across unfamiliar layouts, with handwriting as a particularly useful case. Their generality makes the long tail more approachable.

Source frame: Use vision models selectively, then correct the transcription
Source frame: Use vision models selectively, then correct the transcription

At hundreds of millions of documents, however, efficiency and determinism also shape the design. The work can be divided between complementary capabilities:

  • Layout detection: Small computer-vision models identify document regions and can run on CPUs at scale. They help locate the difficult parts before more expensive processing.
  • Semantic interpretation: VLMs handle challenging visual content and help identify mistakes that simpler processing leaves behind.

This division gives each task a suitable tool instead of making every region depend on a large foundation model.

Once the page has been segmented and difficult regions located, a verification layer can play a role similar to a human reviewer. Reducto calls this agentic OCR. The important decision is to apply targeted token edits to an initial transcription, rather than ask a model to regenerate the whole page. Abraham compares the approach to fast edits in an IDE: preserve the existing output and change the small pieces that need correction.

Why restrict an intelligent model’s freedom? Consider a table with a total that its human author calculated incorrectly. A model rewriting the OCR can recognize the word “total,” add the entries, and replace the printed value with the correct arithmetic result. The output becomes less faithful precisely because the model understood the table. The desired correction is narrower: distinguish 0 from O, or a period from a comma, while retaining the source’s own mistake. Reading a document and checking its arithmetic are separate jobs.

7:307:32
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

7:23 · section reference included

The table needs different forms for retrieval and reasoning

A correct parse still needs a useful representation. Simple tables fit naturally in Markdown. Complex tables may need HTML because merged cells encode relationships that must survive conversion. HTML preserves that structure, but its tags consume tokens; applying it to every simple table adds overhead without a corresponding benefit. Reducto chooses the representation according to the table’s complexity.

Source frame: The table needs different forms for retrieval and reasoning
Source frame: The table needs different forms for retrieval and reasoning

Now follow that complex table into retrieval. A user asks, “how did revenue change over time,” without enumerating the table’s values. The reasoning model may understand the HTML perfectly once it receives it. The embedding model first has to find it among many documents, though, and Reducto finds that matching an ordinary language query to a mass of tags and numbers is difficult. A representation that serves reasoning can therefore obstruct the step that supplies the reasoning context.

The change is to create a second, natural-language rendering of the same table. Retrieval uses that block; reasoning uses the structured HTML. What does each consumer receive? The diagram makes the division visible: the retrieval-friendly text helps select the relevant table, while the reasoning-friendly form preserves its relationships. This addresses the matching problem without requiring the reasoning model to work from the natural-language rendering alone.

How it fits togetherTwo representations of one complex table

Merged cells carry meaning.

Natural-language content serves retrieval; HTML structure serves reasoning over the selected table.

10:4510:49
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

10:45 · section reference included

Better inputs reduce reconstruction work; routing reduces distraction

Returning to the document-decision benchmark, Abraham reports that supplying structured parse results alongside the original PDF improved answers across models from Google, Anthropic, and OpenAI. Some tested models surpassed the earlier frontier model’s out-of-the-box result. The comparison does not establish what that earlier model would do with the same improved inputs: Reducto could not test it after losing access.

Source frame: Better inputs reduce reconstruction work; routing reduces distraction
Source frame: Better inputs reduce reconstruction work; routing reduces distraction

The reported benefit also included fewer reasoning tokens and lower latency. Structured inputs let the model spend less effort reconstructing the data and more effort producing the answer. Improving the input pipeline can thus change both the quality and the amount of downstream work, without changing the model itself.

The next improvement concerns which documents enter the task. A large context window does not make every page helpful; Abraham describes quality erosion from excessive context as well as token cost. Two operations address different parts of that problem:

  • Classification: Route the right document to the appropriate processing pipeline.
  • Splitting: Select the relevant portions of a large document instead of passing the entire packet.

A paper-mail packet can run to hundreds of pages with uncertain contents. Making the final decision model sort through that packet adds a separate job that distracts from extracting information, reasoning over it, or making the intended decision.

13:2713:31
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

13:12 · section reference included

A line chart becomes a table through repeated reconstruction

Line charts expose a harder extraction problem. A vision encoder may retain the broad trend in revenue while losing the pixel-level detail needed to recover individual data points. The chart contains a potentially large table, but a rough visual reading does not recover that table. Abraham presents a case his team could not solve with a single model response.

Source frame: A line chart becomes a table through repeated reconstruction
Source frame: A line chart becomes a table through repeated reconstruction

The agent harness changes the task from a one-shot reading into an iterative reconstruction. The agent has a code interpreter and can visualize the chart it generates from extracted data. It examines that reconstruction, finds mistakes, and revises the extraction repeatedly. The reported result is a Markdown table recovered from the original line chart, with a plotted reconstruction shown in the presentation. The talk does not give a numerical error tolerance for the recovered points, so the example demonstrates the method rather than establishing exact numerical recovery.

Where does the improvement come from? The loop below makes the generated chart inspectable before the table is accepted. Tools give the agent a way to expose errors in its proposed data, and repeated checks let it change that proposal. The useful capability is the whole loop—extraction, execution, visualization, and correction.

Structured extraction can use a similar harness. A parent agent sets validation criteria for subagents, which matters when a form contains tens of thousands of fields and omissions can remain silent. Abraham describes a benchmark tradeoff: frontier models at maximum reasoning were precise about the rows they returned, yet dropped substantial content; dedicated document services recovered more content while making more errors in what they extracted. Precision asks whether returned content is correct. Recall asks how much required content was recovered. Reducto reports that its harness improved both for this task, though the talk supplies no scores to quantify the gain.

How it fits togetherChart extraction with a verification loop

Visual lines encode individual data points.

The agent visualizes its proposed data, finds mistakes, and revises the extraction before producing the final table.

16:3816:42
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

16:10 · section reference included

Evaluate each stage, then let agents choose the document tools

Evaluations should guide these design choices at two levels. Off-the-shelf datasets provide a test bed, while production monitoring catches the differences between a constructed dataset and the documents users actually submit. Within the pipeline, parsing, retrieval, and final formatting each need attention. Perfect parsing cannot help an agent that receives the wrong context. The final question remains whether the changes improve the agent’s completed work.

Source frame: Evaluate each stage, then let agents choose the document tools
Source frame: Evaluate each stage, then let agents choose the document tools

The closing architecture moves some orchestration into the agent itself. Abraham describes customers giving agents a file system to navigate and a CLI through which they can choose document tools as needed. A document no longer has to follow one predetermined end-to-end flow. The agent decides whether it needs to read a particular document and which operation to apply.

That interface separates readable content from metadata. The content field supplies what the agent needs to read; metadata retains information such as bounding boxes for citations. This gives the agent useful text while preserving a way to locate the supporting material in the original document. Editing receives only a brief mention, and document generation is presented as a direction for the coming months rather than an explained capability in this talk.

Abraham’s final recommendation is to reconsider the workflow as capabilities change. Decompose parsing to balance accuracy, cost, and latency; use verification to catch errors; format data for its consumer; classify and split before downstream work; and evaluate every stage. The file-system example extends those choices: tools and document representations can remain well defined even when the agent chooses the sequence in which to use them.

19:0119:06
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

19:01 · section reference included

Read the complete timestamped transcript
  1. 0:01

    [music]

  2. 0:13

    Hi. Can everyone hear me? Sweet. We have

  3. 0:16

    a short 20 minutes here and so I wanted

  4. 0:19

    to jump in and get right to it. Uh my

  5. 0:21

    name is De. I'm the co-founder and CEO

  6. 0:22

    of Redducto. Uh and today we wanted to

  7. 0:25

    talk about one of the I think really

  8. 0:28

    practical but maybe less sexy parts of

  9. 0:31

    building agents that actually work in

  10. 0:33

    the real world uh which is data. Um I'm

  11. 0:35

    sure you've seen plenty of talks about

  12. 0:37

    data today. We primarily have focused on

  13. 0:39

    building infrastructure for anybody

  14. 0:41

    working with some of the hardest sources

  15. 0:43

    of data which is unstructured images,

  16. 0:45

    PDFs, spreadsheets, everything that

  17. 0:47

    humans are used to using day-to-day. Uh,

  18. 0:50

    have any of you used Reduct already or

  19. 0:53

    trial it? Cool. Um, I guess helpful

  20. 0:57

    context maybe to start is we're an

  21. 0:59

    agentic document processing platform.

  22. 1:00

    Um, so we help a lot of the world's

  23. 1:03

    leading AI teams build uh both AI

  24. 1:06

    applications and also workflows

  25. 1:07

    depending on what they're trying to do.

  26. 1:09

    Uh, that includes a lot of the AI

  27. 1:10

    natives that you've probably seen today.

  28. 1:12

    It includes the Harveys, the Lores, the

  29. 1:13

    Rogos of the world. Uh, but also some of

  30. 1:16

    the largest enterprises in the world.

  31. 1:17

    And I think that's important context for

  32. 1:19

    what we're going to talk about. Um, we

  33. 1:21

    work with the largest tech companies,

  34. 1:23

    global financial institutions, insurance

  35. 1:25

    orgs, people that have decades worth of

  36. 1:28

    historical data that historically has

  37. 1:30

    been really, really hard to actually use

  38. 1:32

    outside of a demo context. And across

  39. 1:35

    all these companies at this point, we've

  40. 1:36

    processed many billions of documents for

  41. 1:38

    our customers. I've gotten kind of lazy

  42. 1:40

    with that plus symbol at the end, uh,

  43. 1:42

    but the number keeps changing and so

  44. 1:44

    we'll let it sit there. Uh but the main

  45. 1:47

    thing that we've learned across those

  46. 1:48

    and what I wanted to focus on today is

  47. 1:50

    not actually a reduct product itself.

  48. 1:52

    It's the intricacies of what we've

  49. 1:54

    learned that hopefully you can actually

  50. 1:56

    take home and implement in the work that

  51. 1:58

    you're doing as well. Um and I think a

  52. 2:00

    lot of the work that we've done and a

  53. 2:02

    lot of the learnings we've had are in

  54. 2:04

    this sort of broader thematic context of

  55. 2:06

    the scope of AI applications has changed

  56. 2:08

    a lot recently. Uh from just like

  57. 2:10

    information synthesis products to actual

  58. 2:12

    agents doing work. And with that we

  59. 2:15

    found that there are a few things that

  60. 2:16

    are worth talking through. One is

  61. 2:17

    framing the problem itself. Uh the

  62. 2:19

    bottleneck that people face. Uh why PDFs

  63. 2:22

    in particular are hard even though

  64. 2:23

    you've probably seen two dozen different

  65. 2:26

    PDF processing launches on Twitter. Uh

  66. 2:29

    where we find strengths and weaknesses

  67. 2:30

    with different tools. Uh we think

  68. 2:32

    there's a right place in time for

  69. 2:33

    traditional CV versus BLMs. Um and then

  70. 2:36

    more interestingly I think the latter

  71. 2:37

    half of this talk will actually be

  72. 2:38

    focused on the next frontier. uh what

  73. 2:41

    we've seen be possible as a result of

  74. 2:43

    having agents in the loop uh things that

  75. 2:45

    we found as a result of being able to

  76. 2:46

    use harnesses for different types of

  77. 2:48

    tasks uh and our learnings around things

  78. 2:50

    like evaluations uh as you go from rag

  79. 2:52

    to building agent products but I'll

  80. 2:55

    start with the first thing which is

  81. 2:57

    something that I assume has been harped

  82. 2:59

    on a lot if you went to AI engineer a

  83. 3:01

    few years ago the word that you would

  84. 3:03

    have heard in every single talk would

  85. 3:04

    have been rag um and everybody was

  86. 3:06

    building some form of rag application

  87. 3:08

    for a while all applications ations were

  88. 3:11

    really some form of information

  89. 3:12

    synthesis, right? You would pull in

  90. 3:15

    information from some context, whether

  91. 3:17

    that's a perfect prompt or a file that a

  92. 3:19

    user uploaded and you would have

  93. 3:21

    something like a search product. Um,

  94. 3:22

    you'd have enterprise search, you'd have

  95. 3:24

    a chatbot that would do simple question

  96. 3:26

    answer on top of the content and that

  97. 3:28

    was it. But today, the buzzword that

  98. 3:31

    you've probably heard a million times is

  99. 3:34

    agents. Uh, and for those of you that

  100. 3:36

    are engineers, you've probably had cloud

  101. 3:38

    code or similar tool do a lot of your

  102. 3:40

    endto-end work for many, many tasks. And

  103. 3:43

    that same sort of shift is starting to

  104. 3:44

    happen for all sorts of white collar

  105. 3:46

    work. Um, whether you're in finance or

  106. 3:48

    insurance or healthcare, people are

  107. 3:50

    starting to make autonomous decisions.

  108. 3:52

    They're trying to create end-to-end work

  109. 3:54

    products, not just answer questions from

  110. 3:56

    a PDF, but generate and modify PDFs as

  111. 3:58

    well. And that's a very different

  112. 4:00

    framing. the tools that you need, the

  113. 4:01

    problems that you face change a lot

  114. 4:03

    versus just trying to build a retrieval

  115. 4:06

    uh platform.

  116. 4:08

    And the common core for all of this is

  117. 4:12

    it actually ends up being an even more

  118. 4:14

    importance problem to solve for when you

  119. 4:17

    start having these multi-step pipelines.

  120. 4:19

    If you're just doing question answer,

  121. 4:20

    there is obviously a risk that comes

  122. 4:23

    with just answering the question. But

  123. 4:25

    when agents are making multiple

  124. 4:27

    decisions that compound over the course

  125. 4:29

    of a source of files, when they're

  126. 4:30

    pulling in multiple sources of data, the

  127. 4:32

    risk of bad inputs becomes really,

  128. 4:34

    really pervasive in your pipeline. And

  129. 4:37

    that's what we focus on because at the

  130. 4:40

    end of the day, a lot of the value of

  131. 4:42

    language model tools in the real world

  132. 4:44

    only applies in the context to which you

  133. 4:47

    apply that intelligence. And for a lot

  134. 4:49

    of enterprises, data is unstructured, it

  135. 4:52

    is scattered, and it is multimodal. uh

  136. 4:54

    it's not this like cleanly organized

  137. 4:56

    repository. You have random teams and

  138. 4:58

    different organizations that go through

  139. 5:00

    and put things in Google Drive and Box

  140. 5:02

    and wherever else you might have things.

  141. 5:04

    Um you're going to have data for formats

  142. 5:06

    that are unstructured by default. You

  143. 5:08

    don't necessarily know what is in your

  144. 5:10

    corpus of information. And you're going

  145. 5:12

    to have all sorts of downstream problems

  146. 5:14

    that come with that. Um some of the

  147. 5:16

    problems are going to be parsing

  148. 5:17

    extraction accuracy and that definitely

  149. 5:18

    matters. But it's also going to be about

  150. 5:20

    problems like are you retrieving the

  151. 5:22

    right context? It's going to be about

  152. 5:23

    how do you actually interact with that

  153. 5:25

    context and apply modifications to it.

  154. 5:28

    And the thing that you've probably heard

  155. 5:31

    harp on and others in the space is PDFs

  156. 5:33

    are surprisingly still a very hard

  157. 5:36

    problem. Um I don't know if any of you

  158. 5:38

    follow Serge which is a data lab that

  159. 5:40

    works with a lot of the foundation model

  160. 5:42

    companies. uh they have this really

  161. 5:43

    great benchmark called GDP PDF where

  162. 5:46

    even Fable uh like current frontier of

  163. 5:49

    model intelligence is at about 30% on

  164. 5:51

    their benchmark um and it's entirely

  165. 5:53

    predicated on this idea of can we have

  166. 5:55

    models go through and actually make

  167. 5:57

    determinations off of contents that

  168. 5:59

    would be in documents like PDFs uh and

  169. 6:02

    the reason why they're hard I'll come

  170. 6:03

    back to that benchmark in a second is

  171. 6:05

    fundamentally PDFs as a file format are

  172. 6:08

    both very old but were designed in a

  173. 6:10

    very different context I've genuinely

  174. 6:13

    met people that have worked on PDF

  175. 6:15

    processing longer than I've been alive.

  176. 6:16

    Uh I've met the people that worked on

  177. 6:18

    printer drivers for printers to print

  178. 6:20

    PDFs in the early 1990s. And that's what

  179. 6:22

    it was like. You wanted to be able to

  180. 6:24

    represent what was faithfully on the

  181. 6:26

    documents when you originally created it

  182. 6:28

    and have it be printable at the end of

  183. 6:30

    it. A lot of the considerations today

  184. 6:32

    are not around that. Um ultimately what

  185. 6:34

    we want is something like a markdown

  186. 6:37

    representation, something that agents

  187. 6:38

    will effectively reason on. And in the

  188. 6:41

    real world, humans encode so much

  189. 6:43

    context visually. Like the average

  190. 6:46

    financial analyst that is a new grad in

  191. 6:49

    their IB role uh is not going through

  192. 6:52

    and thinking in terms of will agents

  193. 6:54

    consume this deck. They're creating

  194. 6:55

    these really creative slides. I'm sure

  195. 6:58

    you've seen the softbank slides with uh

  196. 7:00

    goose laying eggs. Those details matter,

  197. 7:03

    right? Like a lot of the data that

  198. 7:04

    you're going to reason on is going to be

  199. 7:06

    tabular structures that maybe don't have

  200. 7:08

    clean grid lines and separate out the

  201. 7:10

    merge cells. You're going to have things

  202. 7:11

    like line charts and graphs. You're

  203. 7:13

    going to have messy handwriting that

  204. 7:15

    even I as a human would often struggle

  205. 7:17

    to read. And that's the sort of problem

  206. 7:19

    that you need to solve for if you're

  207. 7:20

    dealing with that long tail.

  208. 7:23

    And we started the company in 2023

  209. 7:26

    because we felt that there was a sort of

  210. 7:28

    step change in what was possible here.

  211. 7:30

    uh for a while as people would think

  212. 7:32

    about any sort of PDF processing

  213. 7:34

    problem, it used to be some modified

  214. 7:37

    version of an NLP pipeline. You do a

  215. 7:39

    simple OCR pass and then you try to

  216. 7:41

    post-process that text. And that worked

  217. 7:43

    when you would have really consistent

  218. 7:45

    layouts, right? Like if you knew that a

  219. 7:47

    W2 is always going to look like a W2,

  220. 7:49

    that's a constrained problem space that

  221. 7:51

    you can template your way around. But

  222. 7:53

    VLMs are interesting because they are

  223. 7:55

    fundamentally horizontal in nature. um

  224. 7:57

    for the first time you can have this

  225. 7:59

    premise of read the document the way

  226. 8:01

    that a human would have. You can have

  227. 8:02

    this premise of we want to address the

  228. 8:04

    longtail. Uh and so we found that they

  229. 8:07

    were incredible for all sorts of things

  230. 8:09

    like handwritten text in a way that

  231. 8:10

    traditional OCR just never was. But the

  232. 8:13

    flip side is that we don't think that

  233. 8:15

    they're a one-sizefits-all solution. And

  234. 8:17

    if you're solving this problem at scale,

  235. 8:19

    if you're a company dealing with

  236. 8:20

    hundreds of millions of documents, there

  237. 8:22

    are all sorts of secondary

  238. 8:23

    considerations like determinism. you

  239. 8:26

    care a lot about the efficiency of your

  240. 8:27

    processing. Uh, and so we find that

  241. 8:30

    there are some things where traditional

  242. 8:32

    CV is actually still really really

  243. 8:33

    strong. And this has been really

  244. 8:35

    underappreciated as autonomous vehicle

  245. 8:37

    research has gotten better. Techniques

  246. 8:38

    like object detection are more

  247. 8:40

    sophisticated than they were a decade

  248. 8:42

    ago. And so we find that subund million

  249. 8:44

    parameter models are really effective

  250. 8:46

    for things like detecting the layout of

  251. 8:48

    a document. Um, you can go really really

  252. 8:50

    far without even needing a large

  253. 8:52

    foundation model. And these are models

  254. 8:54

    that can actually run on CPU. You can

  255. 8:55

    run them at scale. You can make sure

  256. 8:57

    that you understand on a region level

  257. 8:59

    what are the hard things for you to

  258. 9:00

    process. And VLMs introduce this notion

  259. 9:03

    of semantics. You can go through and

  260. 9:04

    actually identify and make correct the

  261. 9:06

    sorts of mistakes that you were likely

  262. 9:07

    to have in your pipeline.

  263. 9:10

    And off of that idea of semantics, uh,

  264. 9:14

    when you're deconstructing this problem

  265. 9:15

    and you have a clear sense of where the

  266. 9:17

    more nuanced things are, if you've

  267. 9:19

    segmented the text on your page and you

  268. 9:20

    understand where the handwriting is, you

  269. 9:22

    can introduce this notion of a sort of

  270. 9:24

    agent in the loop. Whereas historically,

  271. 9:26

    you would have had a human review team

  272. 9:28

    go through and annotate and correct

  273. 9:29

    mistakes. Uh, VLMs can now present this

  274. 9:32

    sort of idea of what we call agentic

  275. 9:33

    OCR. Really for us what that looks like

  276. 9:36

    is if you've ever used you know a tool

  277. 9:39

    like cursor that's applying fast edits

  278. 9:40

    in your IDE there's this notion of

  279. 9:42

    speculative decoding where you're

  280. 9:44

    applying token level edits to your

  281. 9:45

    output. You can apply a similar sort of

  282. 9:47

    principle here uh that is not just you

  283. 9:49

    know sending OCR to Gemini writing a

  284. 9:52

    really pretty prompt asking it nicely to

  285. 9:55

    not deviate too much from the original

  286. 9:56

    because we find that when you're doing

  287. 9:58

    that sort of next token prediction you

  288. 9:59

    introduce net new loss cases where

  289. 10:01

    models that are really intelligent will

  290. 10:03

    start actually correcting things not

  291. 10:05

    faithfully to what was in the document.

  292. 10:06

    They'll see the word total and if the

  293. 10:08

    human made a mistake in that table

  294. 10:09

    models will actually sometimes go

  295. 10:10

    through and add up the values in the

  296. 10:13

    table themselves. What you want is to

  297. 10:15

    correct the token level edits that you

  298. 10:16

    want. Maybe you messed up a period

  299. 10:18

    versus a comma, a zero versus an O.

  300. 10:20

    Those sorts of details really, really

  301. 10:22

    matter. And that's a question of how do

  302. 10:24

    we actually represent what a human would

  303. 10:25

    have seen if they had read that

  304. 10:26

    document.

  305. 10:29

    So the way that we see this is agentic

  306. 10:31

    OCR is almost like that human loop

  307. 10:34

    analogy where you have the first inputs

  308. 10:37

    uh go through with a CD plus VLM parse

  309. 10:39

    but then you have a verification

  310. 10:41

    correction layer that ends up leading to

  311. 10:42

    a high confidence output.

  312. 10:45

    But I mentioned earlier that we don't

  313. 10:47

    see the range of problems as purely just

  314. 10:49

    parsing and extraction. And a lot of

  315. 10:52

    what this looks like is I think it's

  316. 10:54

    important to think through the details

  317. 10:56

    of your pipeline. Even if you have a

  318. 10:58

    great documents to markdown pipeline and

  319. 11:01

    a great example of this is if you've

  320. 11:03

    built any sort of rag platform um you've

  321. 11:06

    probably had some consideration around

  322. 11:07

    things like tables.

  323. 11:09

    There are a lot of things that you can

  324. 11:11

    encode well in something like markdown.

  325. 11:14

    uh but things like this table where the

  326. 11:16

    merge cells actually encode a lot of

  327. 11:18

    meaning. It matters that you're

  328. 11:20

    preserving that sort of structure and

  329. 11:21

    it's not a model limitation. LMS are

  330. 11:23

    incredible at reasoning through like an

  331. 11:25

    HTML structure of the same table, but

  332. 11:27

    you're also wasting a lot of tokens and

  333. 11:28

    that gets expensive quickly. And

  334. 11:30

    obviously on the other end of the

  335. 11:31

    spectrum, you probably don't want to go

  336. 11:34

    through and encode simple tables in HTML

  337. 11:36

    because then you have a lot of HTML tags

  338. 11:38

    that are erroneous. And so what we ended

  339. 11:40

    up doing was looking at this as sort of

  340. 11:42

    like a dynamic problem of when you have

  341. 11:43

    a simple table, great, we can

  342. 11:45

    approximate that data in markdown. When

  343. 11:48

    you have a more complex table, you may

  344. 11:50

    want to use something like HTML. But

  345. 11:52

    it's not only language models that

  346. 11:54

    should be a consideration in your

  347. 11:55

    pipeline. If you're doing anything

  348. 11:57

    related to embedding, you're also going

  349. 11:59

    to have the secondary problem of

  350. 12:00

    retrieval of that context. Right? That

  351. 12:02

    same table that I showed you earlier, if

  352. 12:05

    you look at the HTML representation is

  353. 12:08

    really really messy. The vast majority

  354. 12:10

    of that snippets is just HTML tags. It's

  355. 12:12

    just classifying the structure of the

  356. 12:14

    document. And the unfortunate thing is

  357. 12:16

    whereas in some blanket evals, you might

  358. 12:18

    have contents that is really trying to

  359. 12:22

    do the work for the model and say

  360. 12:24

    exactly what you're looking for. A real

  361. 12:26

    world person does not enumerate the

  362. 12:28

    values in the table. They just say how

  363. 12:30

    did revenue change over time and they

  364. 12:32

    assume that you're going to retrieve the

  365. 12:33

    right table when it's relevant. And

  366. 12:35

    whereas language models can reason

  367. 12:37

    through that text effectively if you're

  368. 12:38

    pulling from a large corpus, we find

  369. 12:40

    that embedding models really struggle to

  370. 12:42

    correlate that natural language human

  371. 12:44

    prompt with this messy blob of HTML tags

  372. 12:47

    and numbers. And so a big thing that you

  373. 12:49

    can do that actually takes very little

  374. 12:51

    effort is creating a representation

  375. 12:53

    that's more so designed for the

  376. 12:54

    embedding model itself. taking the same

  377. 12:56

    table, we're creating a natural language

  378. 12:58

    representation of that table. So you

  379. 13:00

    have the best of both worlds. When

  380. 13:01

    you're actually passing this into the

  381. 13:02

    model for reasoning, you're using the

  382. 13:04

    HTML table representation. And when

  383. 13:06

    you're trying to make sure that you

  384. 13:07

    retrieve the right snippets, you're

  385. 13:08

    using that natural language block.

  386. 13:12

    Off of that, uh there's also this idea

  387. 13:15

    of there's a lot to do that is not just

  388. 13:17

    parsing and extraction. Um, and I think

  389. 13:19

    a lot of the industry's focus has been

  390. 13:21

    on parsing and extraction historically

  391. 13:23

    because we do think that there's a

  392. 13:24

    massive uplift there. And I talked

  393. 13:27

    earlier about the GDF GDP PDF benchmark

  394. 13:31

    which I think is a great illustrative

  395. 13:33

    example of what you can see uh as a

  396. 13:35

    result of improving your data pipeline.

  397. 13:38

    So what we found is that if you take the

  398. 13:40

    same exact benchmark that I mentioned

  399. 13:42

    earlier, uh, unfortunately we couldn't

  400. 13:43

    test it on Fable because our access was

  401. 13:46

    cut. uh but if you test on other models

  402. 13:48

    and you give it both the original PDF

  403. 13:50

    but also a structured representation of

  404. 13:52

    the PDF like the parse results here

  405. 13:55

    across models whether it's Gemini

  406. 13:57

    whether it's anthropic or openi um you

  407. 14:00

    find that you actually improve end LLM

  408. 14:03

    performance just from better inputs and

  409. 14:05

    it's to an extreme where models like GPT

  410. 14:08

    5.5 and opus actually outperform

  411. 14:11

    something like Fable out of the box not

  412. 14:13

    just on an accuracy basis but as a

  413. 14:16

    result of giving better inputs, the

  414. 14:18

    models end up needing to use fewer

  415. 14:20

    reasoning tokens as well. Um, they're

  416. 14:21

    focused less on representing the data

  417. 14:24

    and more on the actual outputs and as a

  418. 14:26

    result they end up driving down latency

  419. 14:28

    and also getting to the correct answer

  420. 14:30

    more quickly. But even once you have

  421. 14:33

    that sort of pipeline and you've gone

  422. 14:35

    through and you've actually inspected

  423. 14:38

    everything in your parsing layer, a lot

  424. 14:41

    of human work is going to require

  425. 14:42

    actually understanding the range of what

  426. 14:44

    you have in your corpus, routing it to

  427. 14:46

    the appropriate pipeline and sort of

  428. 14:48

    decomposing that problem or even at the

  429. 14:50

    end editing and modifying your document.

  430. 14:52

    And so what we tried to do is look at

  431. 14:54

    this as this problem of how do you make

  432. 14:56

    sure that every interaction that a

  433. 14:57

    language model has with a document is as

  434. 14:59

    effective as if a human would have done

  435. 15:01

    it. Um if you're filling out a form, how

  436. 15:03

    do you make sure that you have precision

  437. 15:04

    in where you fill out fields? And a

  438. 15:06

    really good example of this uh on the

  439. 15:08

    orchestration side is I think

  440. 15:10

    classification splitting are a very

  441. 15:12

    underappreciated way to have an LM do

  442. 15:14

    its best work. Obviously, you can just

  443. 15:16

    go through and dump as much context as

  444. 15:18

    you want. And if you're doing a sort of

  445. 15:19

    needle in the haststack test, that might

  446. 15:21

    be fine. But in practice, there is

  447. 15:23

    erosion that you find in quality outside

  448. 15:25

    of just the token economics as a result

  449. 15:27

    of passing in too much. And instead,

  450. 15:30

    what we find is you can get a lot of

  451. 15:33

    headroom by thinking through things like

  452. 15:35

    how do you classify the right documents

  453. 15:37

    to the right sort of pipeline? And even

  454. 15:39

    for large documents, how do you make

  455. 15:41

    sure that you're passing in the snippets

  456. 15:43

    uh that are actually relevant? We see

  457. 15:45

    use cases where people will have things

  458. 15:46

    like paper mail. Uh and these paper mail

  459. 15:49

    packets can be hundreds of pages long.

  460. 15:51

    You don't necessarily know what is going

  461. 15:53

    to be contained within it. Um you might

  462. 15:55

    have issues like a person interle the

  463. 15:57

    content and having the model do that

  464. 15:59

    sort of work is almost like a

  465. 16:01

    distraction from the work that you're

  466. 16:03

    actually trying to achieve which might

  467. 16:04

    be extracting the data from the paper

  468. 16:06

    mail reasoning on it or making a

  469. 16:07

    decision.

  470. 16:10

    Again, uh I mentioned earlier that the

  471. 16:13

    second half I think is the more

  472. 16:14

    interesting piece. Um which is once you

  473. 16:17

    have the sort of initial classification

  474. 16:20

    and splitting layer, you've figured out

  475. 16:22

    sorting. Uh I think the thing that we

  476. 16:25

    are really excited about as a team is

  477. 16:27

    agent harnesses have been this really

  478. 16:29

    really interesting frontier to push past

  479. 16:31

    what canonically used to be hard

  480. 16:34

    unsolved problems. Um one good example

  481. 16:37

    of this that I'll talk about in a second

  482. 16:38

    is things like line charts. uh we work

  483. 16:40

    with many of the largest hedge funds in

  484. 16:42

    the world and things like line charts

  485. 16:44

    historically have been really really

  486. 16:45

    difficult because one they're an imaged

  487. 16:46

    format but two there's a lot of pixel

  488. 16:48

    level granularity that if you're doing

  489. 16:50

    anything with a traditional vision

  490. 16:52

    encoder you're probably going to lose

  491. 16:54

    you're going to get a rough plot of how

  492. 16:56

    revenue trended but you're not going to

  493. 16:57

    get the individual data points and so

  494. 16:59

    we've been thinking through how do we

  495. 17:00

    give agents the ability to have the

  496. 17:01

    right tools to solve for the specific

  497. 17:04

    type of problem that you're looking at

  498. 17:06

    in the case of chart extraction the

  499. 17:08

    chart on the left encodes codes a

  500. 17:10

    massive table of data. If you actually

  501. 17:12

    went through and tried to plot every

  502. 17:13

    single pixel, it would be really really

  503. 17:15

    difficult. U but it would also be hard

  504. 17:18

    for a model to even approximate the

  505. 17:20

    intricacies of the lines in between. And

  506. 17:22

    there's no model that out of the box can

  507. 17:24

    do this as a singleshot problem. What

  508. 17:26

    you're seeing on the right is a

  509. 17:28

    reconstruction of the markdown table

  510. 17:30

    that we're able to generate off of the

  511. 17:31

    initial line chart. And the only way

  512. 17:33

    that we were able to get there was to

  513. 17:35

    have an agent with all sorts of tools.

  514. 17:37

    It has its own code interpreter. It has

  515. 17:38

    the ability to visualize the chart that

  516. 17:40

    it's generating. And it's iteratively

  517. 17:42

    going through. It's finding mistakes in

  518. 17:43

    the line chart again and again and again

  519. 17:46

    until it's able to get to the final

  520. 17:47

    output. That applies for problems like

  521. 17:50

    structured extraction as well where for

  522. 17:52

    a while we've had this documents to

  523. 17:54

    structured output feature. Uh but you

  524. 17:56

    can really take it a step further by

  525. 17:58

    having an agent harness around that same

  526. 18:00

    sort of task. You can have a parent

  527. 18:01

    agent go through set validation criteria

  528. 18:04

    for sub agents to follow. Uh and this

  529. 18:07

    means that if you have something like a

  530. 18:09

    CBP form with tens of thousands of

  531. 18:11

    fields, that's the sort of problem where

  532. 18:13

    you end up finding a lot of issues that

  533. 18:15

    are silent in nature like you drop

  534. 18:17

    content, you drop rows and micro one

  535. 18:19

    actually released a really good

  536. 18:20

    benchmark in the space this morning

  537. 18:22

    where there's this like bifurcation in

  538. 18:25

    the market. Uh Frontier models with Max

  539. 18:27

    Reasoning are really really precise.

  540. 18:29

    Like provided that they extracted a row,

  541. 18:30

    odds are it's not a hallucinated row

  542. 18:32

    like they actually got it correct. but

  543. 18:34

    they silently drop a lot of the contents

  544. 18:36

    across the benchmark. Uh recall really

  545. 18:38

    really struggles. On the flip side, a

  546. 18:40

    lot of dedicated document processing

  547. 18:42

    services are actually behind frontier

  548. 18:44

    models from a precision perspective, but

  549. 18:47

    close that gap on a recall perspective.

  550. 18:49

    And so there's always been this sort of

  551. 18:51

    trade-off and it was only with an agent

  552. 18:53

    harness that we were able to find that

  553. 18:54

    sort of local maximum of both precision

  554. 18:57

    and also recall for this sort of task.

  555. 19:01

    The last thing and maybe the most

  556. 19:03

    important thing from this talk uh is

  557. 19:06

    that I think at the end of the day eval

  558. 19:09

    should underpin all of your decisions

  559. 19:11

    and it's been a big part of how we think

  560. 19:12

    about our product. Um that applies both

  561. 19:14

    to off-the-shelf data sets that you eval

  562. 19:17

    against but also to things like

  563. 19:18

    real-time production monitoring because

  564. 19:20

    your production data is going to differ

  565. 19:22

    from whatever else you have in your

  566. 19:24

    contrived set. And I really think it's

  567. 19:27

    important to think of eval not as just

  568. 19:29

    this like macrolevel view, but also the

  569. 19:33

    best teams that we work with look at

  570. 19:34

    eval on a granular level for each step

  571. 19:36

    of their pipeline. Uh the first thing

  572. 19:38

    might be that you want to make sure that

  573. 19:39

    the inputs to your pipeline are great.

  574. 19:41

    And of course, you should eval things

  575. 19:43

    like your parsing pipeline. But even

  576. 19:45

    perfect parsing with a horrible

  577. 19:47

    retrieval pipeline is not going to help

  578. 19:49

    if you're not passing the right context.

  579. 19:50

    And so it's important that you're

  580. 19:51

    thinking through details like your

  581. 19:53

    retrieval pipeline, your formatting at

  582. 19:56

    the end of the pipeline, and also

  583. 19:57

    ultimately the most important thing is

  584. 19:59

    are you able to improve end agent

  585. 20:01

    performance.

  586. 20:03

    I'll close off just with a a sense of

  587. 20:06

    where we are headed and where we've seen

  588. 20:07

    the industry head. Uh the most important

  589. 20:10

    thing I think is as agents get better

  590. 20:12

    and better, you can deviate from the

  591. 20:15

    sort of deterministic pipeline that you

  592. 20:16

    would have had a few years ago. Um, a

  593. 20:18

    lot of our customers will actually

  594. 20:20

    create effectively a file system for

  595. 20:21

    their agent to go through and navigate

  596. 20:23

    and let the agent decide what sorts of

  597. 20:25

    tools it wants to use. So, we create a

  598. 20:27

    CLI where instead of people creating a

  599. 20:29

    endto-end pipeline where documents

  600. 20:31

    always follow one specific flow, um, the

  601. 20:33

    agent will decide if it needs to read a

  602. 20:35

    certain type of document and they'll

  603. 20:37

    split that into two sets. Uh, one is a

  604. 20:39

    content field which the agent can read

  605. 20:41

    as it needs to. Uh, but the other is all

  606. 20:44

    the metadata that you would have wanted.

  607. 20:45

    If you're doing things like citations,

  608. 20:47

    you may want bounding boxes, so on and

  609. 20:49

    so forth.

  610. 20:51

    Um, I'll skip this part on editing. I

  611. 20:53

    think there's a lot of interesting work

  612. 20:54

    being done here. We've already released

  613. 20:56

    some of it, but in the next few months,

  614. 20:58

    you will see us look more and more

  615. 21:00

    towards things like document generation.

  616. 21:02

    But the recap for today, and I really

  617. 21:04

    appreciate your time, is one, I highly

  618. 21:07

    recommend that you decompose the parsing

  619. 21:09

    problem. To the extent possible, you

  620. 21:11

    should think of it as the right tool for

  621. 21:12

    the right task so that you can hit that

  622. 21:13

    perimeter frontier of accuracy, cost,

  623. 21:15

    and latency. Two, I think agentic

  624. 21:18

    verification is the biggest step change

  625. 21:20

    that the industry has had for a while

  626. 21:21

    and it's a really good opportunity for

  627. 21:23

    you to make sure that you're building

  628. 21:25

    pipelines that work in production.

  629. 21:27

    Three, I think it takes very little

  630. 21:29

    effort, but there's a lot of headroom

  631. 21:31

    from details like formatting the data

  632. 21:33

    for its consumer. or similar vein, I

  633. 21:36

    think it's really really important to

  634. 21:38

    think about not just the data processing

  635. 21:40

    but the the data orchestration and so

  636. 21:42

    you should always think about tools like

  637. 21:43

    classify and split as a way to augment

  638. 21:46

    your pipeline. Five uh make sure that

  639. 21:49

    you eval at every stage and six think

  640. 21:51

    through what that next frontier looks

  641. 21:53

    like for you because I think most

  642. 21:54

    successful companies in today's era have

  643. 21:57

    deviated a lot from what we used to do

  644. 21:58

    two three years ago. But if you have any

  645. 22:01

    questions uh please feel free to reach

  646. 22:03

    out at any point. My email is just first

  647. 22:05

    namered reductto.ai. Uh and you can also

  648. 22:08

    reach out on our website if we can be

  649. 22:09

    helpful for your use case. Thank you.