AI Engineer World's Fair 2026
From Ingestion to Agents: How AI Teams Build on Document Intelligence — Adit Abraham, Reducto
Read the talk
From Ingestion to Agents: How AI Teams Build on Document Intelligence
Adit Abraham of Reducto explains how to turn visually encoded documents into useful agent inputs: combine specialized models, correct OCR without rewriting the source, separate retrieval from reasoning, and give difficult extraction tasks tools and a verification loop.
From a talk by Adit Abraham
At a glance
Ideas worth remembering
Combine efficient layout detection with VLM interpretation where visual complexity requires it; large models need not process every part of every page.
OCR correction should preserve what the document says, including its mistakes. Targeted token edits reduce the opportunity for a model to rewrite the underlying facts.
Format data for each consumer: natural-language table renderings for retrieval, and structure-preserving representations for reasoning.
Tool use and repeated verification can make difficult extraction tasks tractable, but extraction quality must include both correctness and completeness.
Evaluate individual stages and completed agent work, using production monitoring as well as test datasets. Agents choosing tools still need good inputs, relevant context, and usable outputs.
When documents become inputs to decisions
A bad document parse can spoil a chatbot answer. In an agent workflow, the same mistake can travel through several decisions and into the finished work. That is the practical problem behind Adit Abraham’s talk. As Reducto’s co-founder and CEO, he draws on processing many billions of customer documents, including historical material that enterprises have struggled to use beyond a demo.
The earlier retrieval-augmented generation pattern mostly synthesized information: retrieve context, then answer a question through enterprise search or a chatbot. Agents expand the job. They make decisions, produce work, and generate or modify documents. Input quality matters throughout that longer sequence because each step can carry forward an earlier misunderstanding while adding information from another source.
Enterprise information also arrives without a tidy inventory. Different teams put files in Google Drive, Box, and other repositories; the contents are scattered, unstructured, and multimodal. A useful document system therefore has several jobs: recover the contents accurately, retrieve the relevant context, and support the interactions or modifications the task requires. Parsing is the first problem, but it does not settle the others.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
PDFs preserve appearance; agents need meaning
PDFs were designed to reproduce a document’s appearance faithfully and make it printable. The present task asks for something different: a representation an agent can reason over. Abraham introduces the difficulty with a document-decision benchmark on which a frontier model scored about 30%. That figure describes the benchmark he cites, rather than a general PDF accuracy rate.
Humans encode meaning visually. A financial deck’s author is thinking about readers, rather than whether an agent can consume the slide. Abraham’s memorable example is SoftBank’s imagery of a goose laying eggs: a picture can carry part of the explanation. More routinely, merged cells establish relationships in tables, lines encode numerical trends, and handwriting holds information that even a human may struggle to read. Extracting characters alone can leave those relationships behind.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Use vision models selectively, then correct the transcription
A traditional document pipeline ran OCR and then processed the resulting text. That approach could work well when the layout stayed consistent: a W-2 presents a constrained problem that templates can address. Vision-language models (VLMs) widened the possibilities by reading visual content across unfamiliar layouts, with handwriting as a particularly useful case. Their generality makes the long tail more approachable.
At hundreds of millions of documents, however, efficiency and determinism also shape the design. The work can be divided between complementary capabilities:
- Layout detection: Small computer-vision models identify document regions and can run on CPUs at scale. They help locate the difficult parts before more expensive processing.
- Semantic interpretation: VLMs handle challenging visual content and help identify mistakes that simpler processing leaves behind.
This division gives each task a suitable tool instead of making every region depend on a large foundation model.
Once the page has been segmented and difficult regions located, a verification layer can play a role similar to a human reviewer. Reducto calls this agentic OCR. The important decision is to apply targeted token edits to an initial transcription, rather than ask a model to regenerate the whole page. Abraham compares the approach to fast edits in an IDE: preserve the existing output and change the small pieces that need correction.
Why restrict an intelligent model’s freedom? Consider a table with a total that its human author calculated incorrectly. A model rewriting the OCR can recognize the word “total,” add the entries, and replace the printed value with the correct arithmetic result. The output becomes less faithful precisely because the model understood the table. The desired correction is narrower: distinguish 0 from O, or a period from a comma, while retaining the source’s own mistake. Reading a document and checking its arithmetic are separate jobs.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The table needs different forms for retrieval and reasoning
A correct parse still needs a useful representation. Simple tables fit naturally in Markdown. Complex tables may need HTML because merged cells encode relationships that must survive conversion. HTML preserves that structure, but its tags consume tokens; applying it to every simple table adds overhead without a corresponding benefit. Reducto chooses the representation according to the table’s complexity.
Now follow that complex table into retrieval. A user asks, “how did revenue change over time,” without enumerating the table’s values. The reasoning model may understand the HTML perfectly once it receives it. The embedding model first has to find it among many documents, though, and Reducto finds that matching an ordinary language query to a mass of tags and numbers is difficult. A representation that serves reasoning can therefore obstruct the step that supplies the reasoning context.
The change is to create a second, natural-language rendering of the same table. Retrieval uses that block; reasoning uses the structured HTML. What does each consumer receive? The diagram makes the division visible: the retrieval-friendly text helps select the relevant table, while the reasoning-friendly form preserves its relationships. This addresses the matching problem without requiring the reasoning model to work from the natural-language rendering alone.
Merged cells carry meaning.
Natural-language content serves retrieval; HTML structure serves reasoning over the selected table.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Better inputs reduce reconstruction work; routing reduces distraction
Returning to the document-decision benchmark, Abraham reports that supplying structured parse results alongside the original PDF improved answers across models from Google, Anthropic, and OpenAI. Some tested models surpassed the earlier frontier model’s out-of-the-box result. The comparison does not establish what that earlier model would do with the same improved inputs: Reducto could not test it after losing access.
The reported benefit also included fewer reasoning tokens and lower latency. Structured inputs let the model spend less effort reconstructing the data and more effort producing the answer. Improving the input pipeline can thus change both the quality and the amount of downstream work, without changing the model itself.
The next improvement concerns which documents enter the task. A large context window does not make every page helpful; Abraham describes quality erosion from excessive context as well as token cost. Two operations address different parts of that problem:
- Classification: Route the right document to the appropriate processing pipeline.
- Splitting: Select the relevant portions of a large document instead of passing the entire packet.
A paper-mail packet can run to hundreds of pages with uncertain contents. Making the final decision model sort through that packet adds a separate job that distracts from extracting information, reasoning over it, or making the intended decision.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A line chart becomes a table through repeated reconstruction
Line charts expose a harder extraction problem. A vision encoder may retain the broad trend in revenue while losing the pixel-level detail needed to recover individual data points. The chart contains a potentially large table, but a rough visual reading does not recover that table. Abraham presents a case his team could not solve with a single model response.
The agent harness changes the task from a one-shot reading into an iterative reconstruction. The agent has a code interpreter and can visualize the chart it generates from extracted data. It examines that reconstruction, finds mistakes, and revises the extraction repeatedly. The reported result is a Markdown table recovered from the original line chart, with a plotted reconstruction shown in the presentation. The talk does not give a numerical error tolerance for the recovered points, so the example demonstrates the method rather than establishing exact numerical recovery.
Where does the improvement come from? The loop below makes the generated chart inspectable before the table is accepted. Tools give the agent a way to expose errors in its proposed data, and repeated checks let it change that proposal. The useful capability is the whole loop—extraction, execution, visualization, and correction.
Structured extraction can use a similar harness. A parent agent sets validation criteria for subagents, which matters when a form contains tens of thousands of fields and omissions can remain silent. Abraham describes a benchmark tradeoff: frontier models at maximum reasoning were precise about the rows they returned, yet dropped substantial content; dedicated document services recovered more content while making more errors in what they extracted. Precision asks whether returned content is correct. Recall asks how much required content was recovered. Reducto reports that its harness improved both for this task, though the talk supplies no scores to quantify the gain.
Visual lines encode individual data points.
The agent visualizes its proposed data, finds mistakes, and revises the extraction before producing the final table.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Evaluate each stage, then let agents choose the document tools
Evaluations should guide these design choices at two levels. Off-the-shelf datasets provide a test bed, while production monitoring catches the differences between a constructed dataset and the documents users actually submit. Within the pipeline, parsing, retrieval, and final formatting each need attention. Perfect parsing cannot help an agent that receives the wrong context. The final question remains whether the changes improve the agent’s completed work.
The closing architecture moves some orchestration into the agent itself. Abraham describes customers giving agents a file system to navigate and a CLI through which they can choose document tools as needed. A document no longer has to follow one predetermined end-to-end flow. The agent decides whether it needs to read a particular document and which operation to apply.
That interface separates readable content from metadata. The content field supplies what the agent needs to read; metadata retains information such as bounding boxes for citations. This gives the agent useful text while preserving a way to locate the supporting material in the original document. Editing receives only a brief mention, and document generation is presented as a direction for the coming months rather than an explained capability in this talk.
Abraham’s final recommendation is to reconsider the workflow as capabilities change. Decompose parsing to balance accuracy, cost, and latency; use verification to catch errors; format data for its consumer; classify and split before downstream work; and evaluate every stage. The file-system example extends those choices: tools and document representations can remain well defined even when the agent chooses the sequence in which to use them.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Read the complete timestamped transcript
- 0:01
[music]
- 0:13
Hi. Can everyone hear me? Sweet. We have
- 0:16
a short 20 minutes here and so I wanted
- 0:19
to jump in and get right to it. Uh my
- 0:21
name is De. I'm the co-founder and CEO
- 0:22
of Redducto. Uh and today we wanted to
- 0:25
talk about one of the I think really
- 0:28
practical but maybe less sexy parts of
- 0:31
building agents that actually work in
- 0:33
the real world uh which is data. Um I'm
- 0:35
sure you've seen plenty of talks about
- 0:37
data today. We primarily have focused on
- 0:39
building infrastructure for anybody
- 0:41
working with some of the hardest sources
- 0:43
of data which is unstructured images,
- 0:45
PDFs, spreadsheets, everything that
- 0:47
humans are used to using day-to-day. Uh,
- 0:50
have any of you used Reduct already or
- 0:53
trial it? Cool. Um, I guess helpful
- 0:57
context maybe to start is we're an
- 0:59
agentic document processing platform.
- 1:00
Um, so we help a lot of the world's
- 1:03
leading AI teams build uh both AI
- 1:06
applications and also workflows
- 1:07
depending on what they're trying to do.
- 1:09
Uh, that includes a lot of the AI
- 1:10
natives that you've probably seen today.
- 1:12
It includes the Harveys, the Lores, the
- 1:13
Rogos of the world. Uh, but also some of
- 1:16
the largest enterprises in the world.
- 1:17
And I think that's important context for
- 1:19
what we're going to talk about. Um, we
- 1:21
work with the largest tech companies,
- 1:23
global financial institutions, insurance
- 1:25
orgs, people that have decades worth of
- 1:28
historical data that historically has
- 1:30
been really, really hard to actually use
- 1:32
outside of a demo context. And across
- 1:35
all these companies at this point, we've
- 1:36
processed many billions of documents for
- 1:38
our customers. I've gotten kind of lazy
- 1:40
with that plus symbol at the end, uh,
- 1:42
but the number keeps changing and so
- 1:44
we'll let it sit there. Uh but the main
- 1:47
thing that we've learned across those
- 1:48
and what I wanted to focus on today is
- 1:50
not actually a reduct product itself.
- 1:52
It's the intricacies of what we've
- 1:54
learned that hopefully you can actually
- 1:56
take home and implement in the work that
- 1:58
you're doing as well. Um and I think a
- 2:00
lot of the work that we've done and a
- 2:02
lot of the learnings we've had are in
- 2:04
this sort of broader thematic context of
- 2:06
the scope of AI applications has changed
- 2:08
a lot recently. Uh from just like
- 2:10
information synthesis products to actual
- 2:12
agents doing work. And with that we
- 2:15
found that there are a few things that
- 2:16
are worth talking through. One is
- 2:17
framing the problem itself. Uh the
- 2:19
bottleneck that people face. Uh why PDFs
- 2:22
in particular are hard even though
- 2:23
you've probably seen two dozen different
- 2:26
PDF processing launches on Twitter. Uh
- 2:29
where we find strengths and weaknesses
- 2:30
with different tools. Uh we think
- 2:32
there's a right place in time for
- 2:33
traditional CV versus BLMs. Um and then
- 2:36
more interestingly I think the latter
- 2:37
half of this talk will actually be
- 2:38
focused on the next frontier. uh what
- 2:41
we've seen be possible as a result of
- 2:43
having agents in the loop uh things that
- 2:45
we found as a result of being able to
- 2:46
use harnesses for different types of
- 2:48
tasks uh and our learnings around things
- 2:50
like evaluations uh as you go from rag
- 2:52
to building agent products but I'll
- 2:55
start with the first thing which is
- 2:57
something that I assume has been harped
- 2:59
on a lot if you went to AI engineer a
- 3:01
few years ago the word that you would
- 3:03
have heard in every single talk would
- 3:04
have been rag um and everybody was
- 3:06
building some form of rag application
- 3:08
for a while all applications ations were
- 3:11
really some form of information
- 3:12
synthesis, right? You would pull in
- 3:15
information from some context, whether
- 3:17
that's a perfect prompt or a file that a
- 3:19
user uploaded and you would have
- 3:21
something like a search product. Um,
- 3:22
you'd have enterprise search, you'd have
- 3:24
a chatbot that would do simple question
- 3:26
answer on top of the content and that
- 3:28
was it. But today, the buzzword that
- 3:31
you've probably heard a million times is
- 3:34
agents. Uh, and for those of you that
- 3:36
are engineers, you've probably had cloud
- 3:38
code or similar tool do a lot of your
- 3:40
endto-end work for many, many tasks. And
- 3:43
that same sort of shift is starting to
- 3:44
happen for all sorts of white collar
- 3:46
work. Um, whether you're in finance or
- 3:48
insurance or healthcare, people are
- 3:50
starting to make autonomous decisions.
- 3:52
They're trying to create end-to-end work
- 3:54
products, not just answer questions from
- 3:56
a PDF, but generate and modify PDFs as
- 3:58
well. And that's a very different
- 4:00
framing. the tools that you need, the
- 4:01
problems that you face change a lot
- 4:03
versus just trying to build a retrieval
- 4:06
uh platform.
- 4:08
And the common core for all of this is
- 4:12
it actually ends up being an even more
- 4:14
importance problem to solve for when you
- 4:17
start having these multi-step pipelines.
- 4:19
If you're just doing question answer,
- 4:20
there is obviously a risk that comes
- 4:23
with just answering the question. But
- 4:25
when agents are making multiple
- 4:27
decisions that compound over the course
- 4:29
of a source of files, when they're
- 4:30
pulling in multiple sources of data, the
- 4:32
risk of bad inputs becomes really,
- 4:34
really pervasive in your pipeline. And
- 4:37
that's what we focus on because at the
- 4:40
end of the day, a lot of the value of
- 4:42
language model tools in the real world
- 4:44
only applies in the context to which you
- 4:47
apply that intelligence. And for a lot
- 4:49
of enterprises, data is unstructured, it
- 4:52
is scattered, and it is multimodal. uh
- 4:54
it's not this like cleanly organized
- 4:56
repository. You have random teams and
- 4:58
different organizations that go through
- 5:00
and put things in Google Drive and Box
- 5:02
and wherever else you might have things.
- 5:04
Um you're going to have data for formats
- 5:06
that are unstructured by default. You
- 5:08
don't necessarily know what is in your
- 5:10
corpus of information. And you're going
- 5:12
to have all sorts of downstream problems
- 5:14
that come with that. Um some of the
- 5:16
problems are going to be parsing
- 5:17
extraction accuracy and that definitely
- 5:18
matters. But it's also going to be about
- 5:20
problems like are you retrieving the
- 5:22
right context? It's going to be about
- 5:23
how do you actually interact with that
- 5:25
context and apply modifications to it.
- 5:28
And the thing that you've probably heard
- 5:31
harp on and others in the space is PDFs
- 5:33
are surprisingly still a very hard
- 5:36
problem. Um I don't know if any of you
- 5:38
follow Serge which is a data lab that
- 5:40
works with a lot of the foundation model
- 5:42
companies. uh they have this really
- 5:43
great benchmark called GDP PDF where
- 5:46
even Fable uh like current frontier of
- 5:49
model intelligence is at about 30% on
- 5:51
their benchmark um and it's entirely
- 5:53
predicated on this idea of can we have
- 5:55
models go through and actually make
- 5:57
determinations off of contents that
- 5:59
would be in documents like PDFs uh and
- 6:02
the reason why they're hard I'll come
- 6:03
back to that benchmark in a second is
- 6:05
fundamentally PDFs as a file format are
- 6:08
both very old but were designed in a
- 6:10
very different context I've genuinely
- 6:13
met people that have worked on PDF
- 6:15
processing longer than I've been alive.
- 6:16
Uh I've met the people that worked on
- 6:18
printer drivers for printers to print
- 6:20
PDFs in the early 1990s. And that's what
- 6:22
it was like. You wanted to be able to
- 6:24
represent what was faithfully on the
- 6:26
documents when you originally created it
- 6:28
and have it be printable at the end of
- 6:30
it. A lot of the considerations today
- 6:32
are not around that. Um ultimately what
- 6:34
we want is something like a markdown
- 6:37
representation, something that agents
- 6:38
will effectively reason on. And in the
- 6:41
real world, humans encode so much
- 6:43
context visually. Like the average
- 6:46
financial analyst that is a new grad in
- 6:49
their IB role uh is not going through
- 6:52
and thinking in terms of will agents
- 6:54
consume this deck. They're creating
- 6:55
these really creative slides. I'm sure
- 6:58
you've seen the softbank slides with uh
- 7:00
goose laying eggs. Those details matter,
- 7:03
right? Like a lot of the data that
- 7:04
you're going to reason on is going to be
- 7:06
tabular structures that maybe don't have
- 7:08
clean grid lines and separate out the
- 7:10
merge cells. You're going to have things
- 7:11
like line charts and graphs. You're
- 7:13
going to have messy handwriting that
- 7:15
even I as a human would often struggle
- 7:17
to read. And that's the sort of problem
- 7:19
that you need to solve for if you're
- 7:20
dealing with that long tail.
- 7:23
And we started the company in 2023
- 7:26
because we felt that there was a sort of
- 7:28
step change in what was possible here.
- 7:30
uh for a while as people would think
- 7:32
about any sort of PDF processing
- 7:34
problem, it used to be some modified
- 7:37
version of an NLP pipeline. You do a
- 7:39
simple OCR pass and then you try to
- 7:41
post-process that text. And that worked
- 7:43
when you would have really consistent
- 7:45
layouts, right? Like if you knew that a
- 7:47
W2 is always going to look like a W2,
- 7:49
that's a constrained problem space that
- 7:51
you can template your way around. But
- 7:53
VLMs are interesting because they are
- 7:55
fundamentally horizontal in nature. um
- 7:57
for the first time you can have this
- 7:59
premise of read the document the way
- 8:01
that a human would have. You can have
- 8:02
this premise of we want to address the
- 8:04
longtail. Uh and so we found that they
- 8:07
were incredible for all sorts of things
- 8:09
like handwritten text in a way that
- 8:10
traditional OCR just never was. But the
- 8:13
flip side is that we don't think that
- 8:15
they're a one-sizefits-all solution. And
- 8:17
if you're solving this problem at scale,
- 8:19
if you're a company dealing with
- 8:20
hundreds of millions of documents, there
- 8:22
are all sorts of secondary
- 8:23
considerations like determinism. you
- 8:26
care a lot about the efficiency of your
- 8:27
processing. Uh, and so we find that
- 8:30
there are some things where traditional
- 8:32
CV is actually still really really
- 8:33
strong. And this has been really
- 8:35
underappreciated as autonomous vehicle
- 8:37
research has gotten better. Techniques
- 8:38
like object detection are more
- 8:40
sophisticated than they were a decade
- 8:42
ago. And so we find that subund million
- 8:44
parameter models are really effective
- 8:46
for things like detecting the layout of
- 8:48
a document. Um, you can go really really
- 8:50
far without even needing a large
- 8:52
foundation model. And these are models
- 8:54
that can actually run on CPU. You can
- 8:55
run them at scale. You can make sure
- 8:57
that you understand on a region level
- 8:59
what are the hard things for you to
- 9:00
process. And VLMs introduce this notion
- 9:03
of semantics. You can go through and
- 9:04
actually identify and make correct the
- 9:06
sorts of mistakes that you were likely
- 9:07
to have in your pipeline.
- 9:10
And off of that idea of semantics, uh,
- 9:14
when you're deconstructing this problem
- 9:15
and you have a clear sense of where the
- 9:17
more nuanced things are, if you've
- 9:19
segmented the text on your page and you
- 9:20
understand where the handwriting is, you
- 9:22
can introduce this notion of a sort of
- 9:24
agent in the loop. Whereas historically,
- 9:26
you would have had a human review team
- 9:28
go through and annotate and correct
- 9:29
mistakes. Uh, VLMs can now present this
- 9:32
sort of idea of what we call agentic
- 9:33
OCR. Really for us what that looks like
- 9:36
is if you've ever used you know a tool
- 9:39
like cursor that's applying fast edits
- 9:40
in your IDE there's this notion of
- 9:42
speculative decoding where you're
- 9:44
applying token level edits to your
- 9:45
output. You can apply a similar sort of
- 9:47
principle here uh that is not just you
- 9:49
know sending OCR to Gemini writing a
- 9:52
really pretty prompt asking it nicely to
- 9:55
not deviate too much from the original
- 9:56
because we find that when you're doing
- 9:58
that sort of next token prediction you
- 9:59
introduce net new loss cases where
- 10:01
models that are really intelligent will
- 10:03
start actually correcting things not
- 10:05
faithfully to what was in the document.
- 10:06
They'll see the word total and if the
- 10:08
human made a mistake in that table
- 10:09
models will actually sometimes go
- 10:10
through and add up the values in the
- 10:13
table themselves. What you want is to
- 10:15
correct the token level edits that you
- 10:16
want. Maybe you messed up a period
- 10:18
versus a comma, a zero versus an O.
- 10:20
Those sorts of details really, really
- 10:22
matter. And that's a question of how do
- 10:24
we actually represent what a human would
- 10:25
have seen if they had read that
- 10:26
document.
- 10:29
So the way that we see this is agentic
- 10:31
OCR is almost like that human loop
- 10:34
analogy where you have the first inputs
- 10:37
uh go through with a CD plus VLM parse
- 10:39
but then you have a verification
- 10:41
correction layer that ends up leading to
- 10:42
a high confidence output.
- 10:45
But I mentioned earlier that we don't
- 10:47
see the range of problems as purely just
- 10:49
parsing and extraction. And a lot of
- 10:52
what this looks like is I think it's
- 10:54
important to think through the details
- 10:56
of your pipeline. Even if you have a
- 10:58
great documents to markdown pipeline and
- 11:01
a great example of this is if you've
- 11:03
built any sort of rag platform um you've
- 11:06
probably had some consideration around
- 11:07
things like tables.
- 11:09
There are a lot of things that you can
- 11:11
encode well in something like markdown.
- 11:14
uh but things like this table where the
- 11:16
merge cells actually encode a lot of
- 11:18
meaning. It matters that you're
- 11:20
preserving that sort of structure and
- 11:21
it's not a model limitation. LMS are
- 11:23
incredible at reasoning through like an
- 11:25
HTML structure of the same table, but
- 11:27
you're also wasting a lot of tokens and
- 11:28
that gets expensive quickly. And
- 11:30
obviously on the other end of the
- 11:31
spectrum, you probably don't want to go
- 11:34
through and encode simple tables in HTML
- 11:36
because then you have a lot of HTML tags
- 11:38
that are erroneous. And so what we ended
- 11:40
up doing was looking at this as sort of
- 11:42
like a dynamic problem of when you have
- 11:43
a simple table, great, we can
- 11:45
approximate that data in markdown. When
- 11:48
you have a more complex table, you may
- 11:50
want to use something like HTML. But
- 11:52
it's not only language models that
- 11:54
should be a consideration in your
- 11:55
pipeline. If you're doing anything
- 11:57
related to embedding, you're also going
- 11:59
to have the secondary problem of
- 12:00
retrieval of that context. Right? That
- 12:02
same table that I showed you earlier, if
- 12:05
you look at the HTML representation is
- 12:08
really really messy. The vast majority
- 12:10
of that snippets is just HTML tags. It's
- 12:12
just classifying the structure of the
- 12:14
document. And the unfortunate thing is
- 12:16
whereas in some blanket evals, you might
- 12:18
have contents that is really trying to
- 12:22
do the work for the model and say
- 12:24
exactly what you're looking for. A real
- 12:26
world person does not enumerate the
- 12:28
values in the table. They just say how
- 12:30
did revenue change over time and they
- 12:32
assume that you're going to retrieve the
- 12:33
right table when it's relevant. And
- 12:35
whereas language models can reason
- 12:37
through that text effectively if you're
- 12:38
pulling from a large corpus, we find
- 12:40
that embedding models really struggle to
- 12:42
correlate that natural language human
- 12:44
prompt with this messy blob of HTML tags
- 12:47
and numbers. And so a big thing that you
- 12:49
can do that actually takes very little
- 12:51
effort is creating a representation
- 12:53
that's more so designed for the
- 12:54
embedding model itself. taking the same
- 12:56
table, we're creating a natural language
- 12:58
representation of that table. So you
- 13:00
have the best of both worlds. When
- 13:01
you're actually passing this into the
- 13:02
model for reasoning, you're using the
- 13:04
HTML table representation. And when
- 13:06
you're trying to make sure that you
- 13:07
retrieve the right snippets, you're
- 13:08
using that natural language block.
- 13:12
Off of that, uh there's also this idea
- 13:15
of there's a lot to do that is not just
- 13:17
parsing and extraction. Um, and I think
- 13:19
a lot of the industry's focus has been
- 13:21
on parsing and extraction historically
- 13:23
because we do think that there's a
- 13:24
massive uplift there. And I talked
- 13:27
earlier about the GDF GDP PDF benchmark
- 13:31
which I think is a great illustrative
- 13:33
example of what you can see uh as a
- 13:35
result of improving your data pipeline.
- 13:38
So what we found is that if you take the
- 13:40
same exact benchmark that I mentioned
- 13:42
earlier, uh, unfortunately we couldn't
- 13:43
test it on Fable because our access was
- 13:46
cut. uh but if you test on other models
- 13:48
and you give it both the original PDF
- 13:50
but also a structured representation of
- 13:52
the PDF like the parse results here
- 13:55
across models whether it's Gemini
- 13:57
whether it's anthropic or openi um you
- 14:00
find that you actually improve end LLM
- 14:03
performance just from better inputs and
- 14:05
it's to an extreme where models like GPT
- 14:08
5.5 and opus actually outperform
- 14:11
something like Fable out of the box not
- 14:13
just on an accuracy basis but as a
- 14:16
result of giving better inputs, the
- 14:18
models end up needing to use fewer
- 14:20
reasoning tokens as well. Um, they're
- 14:21
focused less on representing the data
- 14:24
and more on the actual outputs and as a
- 14:26
result they end up driving down latency
- 14:28
and also getting to the correct answer
- 14:30
more quickly. But even once you have
- 14:33
that sort of pipeline and you've gone
- 14:35
through and you've actually inspected
- 14:38
everything in your parsing layer, a lot
- 14:41
of human work is going to require
- 14:42
actually understanding the range of what
- 14:44
you have in your corpus, routing it to
- 14:46
the appropriate pipeline and sort of
- 14:48
decomposing that problem or even at the
- 14:50
end editing and modifying your document.
- 14:52
And so what we tried to do is look at
- 14:54
this as this problem of how do you make
- 14:56
sure that every interaction that a
- 14:57
language model has with a document is as
- 14:59
effective as if a human would have done
- 15:01
it. Um if you're filling out a form, how
- 15:03
do you make sure that you have precision
- 15:04
in where you fill out fields? And a
- 15:06
really good example of this uh on the
- 15:08
orchestration side is I think
- 15:10
classification splitting are a very
- 15:12
underappreciated way to have an LM do
- 15:14
its best work. Obviously, you can just
- 15:16
go through and dump as much context as
- 15:18
you want. And if you're doing a sort of
- 15:19
needle in the haststack test, that might
- 15:21
be fine. But in practice, there is
- 15:23
erosion that you find in quality outside
- 15:25
of just the token economics as a result
- 15:27
of passing in too much. And instead,
- 15:30
what we find is you can get a lot of
- 15:33
headroom by thinking through things like
- 15:35
how do you classify the right documents
- 15:37
to the right sort of pipeline? And even
- 15:39
for large documents, how do you make
- 15:41
sure that you're passing in the snippets
- 15:43
uh that are actually relevant? We see
- 15:45
use cases where people will have things
- 15:46
like paper mail. Uh and these paper mail
- 15:49
packets can be hundreds of pages long.
- 15:51
You don't necessarily know what is going
- 15:53
to be contained within it. Um you might
- 15:55
have issues like a person interle the
- 15:57
content and having the model do that
- 15:59
sort of work is almost like a
- 16:01
distraction from the work that you're
- 16:03
actually trying to achieve which might
- 16:04
be extracting the data from the paper
- 16:06
mail reasoning on it or making a
- 16:07
decision.
- 16:10
Again, uh I mentioned earlier that the
- 16:13
second half I think is the more
- 16:14
interesting piece. Um which is once you
- 16:17
have the sort of initial classification
- 16:20
and splitting layer, you've figured out
- 16:22
sorting. Uh I think the thing that we
- 16:25
are really excited about as a team is
- 16:27
agent harnesses have been this really
- 16:29
really interesting frontier to push past
- 16:31
what canonically used to be hard
- 16:34
unsolved problems. Um one good example
- 16:37
of this that I'll talk about in a second
- 16:38
is things like line charts. uh we work
- 16:40
with many of the largest hedge funds in
- 16:42
the world and things like line charts
- 16:44
historically have been really really
- 16:45
difficult because one they're an imaged
- 16:46
format but two there's a lot of pixel
- 16:48
level granularity that if you're doing
- 16:50
anything with a traditional vision
- 16:52
encoder you're probably going to lose
- 16:54
you're going to get a rough plot of how
- 16:56
revenue trended but you're not going to
- 16:57
get the individual data points and so
- 16:59
we've been thinking through how do we
- 17:00
give agents the ability to have the
- 17:01
right tools to solve for the specific
- 17:04
type of problem that you're looking at
- 17:06
in the case of chart extraction the
- 17:08
chart on the left encodes codes a
- 17:10
massive table of data. If you actually
- 17:12
went through and tried to plot every
- 17:13
single pixel, it would be really really
- 17:15
difficult. U but it would also be hard
- 17:18
for a model to even approximate the
- 17:20
intricacies of the lines in between. And
- 17:22
there's no model that out of the box can
- 17:24
do this as a singleshot problem. What
- 17:26
you're seeing on the right is a
- 17:28
reconstruction of the markdown table
- 17:30
that we're able to generate off of the
- 17:31
initial line chart. And the only way
- 17:33
that we were able to get there was to
- 17:35
have an agent with all sorts of tools.
- 17:37
It has its own code interpreter. It has
- 17:38
the ability to visualize the chart that
- 17:40
it's generating. And it's iteratively
- 17:42
going through. It's finding mistakes in
- 17:43
the line chart again and again and again
- 17:46
until it's able to get to the final
- 17:47
output. That applies for problems like
- 17:50
structured extraction as well where for
- 17:52
a while we've had this documents to
- 17:54
structured output feature. Uh but you
- 17:56
can really take it a step further by
- 17:58
having an agent harness around that same
- 18:00
sort of task. You can have a parent
- 18:01
agent go through set validation criteria
- 18:04
for sub agents to follow. Uh and this
- 18:07
means that if you have something like a
- 18:09
CBP form with tens of thousands of
- 18:11
fields, that's the sort of problem where
- 18:13
you end up finding a lot of issues that
- 18:15
are silent in nature like you drop
- 18:17
content, you drop rows and micro one
- 18:19
actually released a really good
- 18:20
benchmark in the space this morning
- 18:22
where there's this like bifurcation in
- 18:25
the market. Uh Frontier models with Max
- 18:27
Reasoning are really really precise.
- 18:29
Like provided that they extracted a row,
- 18:30
odds are it's not a hallucinated row
- 18:32
like they actually got it correct. but
- 18:34
they silently drop a lot of the contents
- 18:36
across the benchmark. Uh recall really
- 18:38
really struggles. On the flip side, a
- 18:40
lot of dedicated document processing
- 18:42
services are actually behind frontier
- 18:44
models from a precision perspective, but
- 18:47
close that gap on a recall perspective.
- 18:49
And so there's always been this sort of
- 18:51
trade-off and it was only with an agent
- 18:53
harness that we were able to find that
- 18:54
sort of local maximum of both precision
- 18:57
and also recall for this sort of task.
- 19:01
The last thing and maybe the most
- 19:03
important thing from this talk uh is
- 19:06
that I think at the end of the day eval
- 19:09
should underpin all of your decisions
- 19:11
and it's been a big part of how we think
- 19:12
about our product. Um that applies both
- 19:14
to off-the-shelf data sets that you eval
- 19:17
against but also to things like
- 19:18
real-time production monitoring because
- 19:20
your production data is going to differ
- 19:22
from whatever else you have in your
- 19:24
contrived set. And I really think it's
- 19:27
important to think of eval not as just
- 19:29
this like macrolevel view, but also the
- 19:33
best teams that we work with look at
- 19:34
eval on a granular level for each step
- 19:36
of their pipeline. Uh the first thing
- 19:38
might be that you want to make sure that
- 19:39
the inputs to your pipeline are great.
- 19:41
And of course, you should eval things
- 19:43
like your parsing pipeline. But even
- 19:45
perfect parsing with a horrible
- 19:47
retrieval pipeline is not going to help
- 19:49
if you're not passing the right context.
- 19:50
And so it's important that you're
- 19:51
thinking through details like your
- 19:53
retrieval pipeline, your formatting at
- 19:56
the end of the pipeline, and also
- 19:57
ultimately the most important thing is
- 19:59
are you able to improve end agent
- 20:01
performance.
- 20:03
I'll close off just with a a sense of
- 20:06
where we are headed and where we've seen
- 20:07
the industry head. Uh the most important
- 20:10
thing I think is as agents get better
- 20:12
and better, you can deviate from the
- 20:15
sort of deterministic pipeline that you
- 20:16
would have had a few years ago. Um, a
- 20:18
lot of our customers will actually
- 20:20
create effectively a file system for
- 20:21
their agent to go through and navigate
- 20:23
and let the agent decide what sorts of
- 20:25
tools it wants to use. So, we create a
- 20:27
CLI where instead of people creating a
- 20:29
endto-end pipeline where documents
- 20:31
always follow one specific flow, um, the
- 20:33
agent will decide if it needs to read a
- 20:35
certain type of document and they'll
- 20:37
split that into two sets. Uh, one is a
- 20:39
content field which the agent can read
- 20:41
as it needs to. Uh, but the other is all
- 20:44
the metadata that you would have wanted.
- 20:45
If you're doing things like citations,
- 20:47
you may want bounding boxes, so on and
- 20:49
so forth.
- 20:51
Um, I'll skip this part on editing. I
- 20:53
think there's a lot of interesting work
- 20:54
being done here. We've already released
- 20:56
some of it, but in the next few months,
- 20:58
you will see us look more and more
- 21:00
towards things like document generation.
- 21:02
But the recap for today, and I really
- 21:04
appreciate your time, is one, I highly
- 21:07
recommend that you decompose the parsing
- 21:09
problem. To the extent possible, you
- 21:11
should think of it as the right tool for
- 21:12
the right task so that you can hit that
- 21:13
perimeter frontier of accuracy, cost,
- 21:15
and latency. Two, I think agentic
- 21:18
verification is the biggest step change
- 21:20
that the industry has had for a while
- 21:21
and it's a really good opportunity for
- 21:23
you to make sure that you're building
- 21:25
pipelines that work in production.
- 21:27
Three, I think it takes very little
- 21:29
effort, but there's a lot of headroom
- 21:31
from details like formatting the data
- 21:33
for its consumer. or similar vein, I
- 21:36
think it's really really important to
- 21:38
think about not just the data processing
- 21:40
but the the data orchestration and so
- 21:42
you should always think about tools like
- 21:43
classify and split as a way to augment
- 21:46
your pipeline. Five uh make sure that
- 21:49
you eval at every stage and six think
- 21:51
through what that next frontier looks
- 21:53
like for you because I think most
- 21:54
successful companies in today's era have
- 21:57
deviated a lot from what we used to do
- 21:58
two three years ago. But if you have any
- 22:01
questions uh please feel free to reach
- 22:03
out at any point. My email is just first
- 22:05
namered reductto.ai. Uh and you can also
- 22:08
reach out on our website if we can be
- 22:09
helpful for your use case. Thank you.