AI Engineer World's Fair 2026

From Raw Documents to AI-Ready Data — Leo Platzer & Jeff Koss

Jeff Koss· Deasy Labs / CollibraLeo Platzer· Deasy Labs / Collibra21:37

Read the talk

From Raw Documents to AI-Ready Data

Jeff Koss and Leo Platzer show how document classification, sensitivity checks and scheduled curation turn a small legal-ops chatbot pilot into a maintainable retrieval system—and why duplicates can consume far more context than their share of the corpus suggests.

From a talk by Jeff Koss and Leo Platzer

At a glance

Ideas worth remembering

  • A successful 40-file pilot can hide the work of discovering relevant documents, excluding sensitive information and choosing trustworthy versions at production scale.

  • Conditional taxonomies classify a broad category first, then extract more specific metadata within its branch; evidence and human feedback make those labels reviewable.

  • A data slice combines selection criteria with scheduled maintenance, carrying source additions and deletions through to downstream retrieval data.

  • Stale and duplicated material can occupy a disproportionate share of retrieved context because top-K retrieval selects matches rather than sampling the corpus uniformly.

  • Folder-level context files reuse metadata to help agents choose where to explore before spending tokens inspecting individual documents.

Forty curated files hide the hard part

A legal-ops chatbot works well on around 40 hand-picked files. The pilot parses documents, performs optical character recognition, splits content into chunks, creates embeddings, and loads a vector database and a knowledge graph. Then the team expands to more than 80,000 files. Three questions stall the project: which files belong, where sensitive information lives, and which documents deserve trust. The retrieval pipeline already works; selecting its inputs has become the problem.

Jeff Koss, a staff customer engineer at Deasy Labs, develops this problem through the North River Manufacturing demo scenario. Leo Platzer, Deasy’s former CTO, later explains the retrieval consequences. Deasy supplies the unstructured-data capabilities discussed here following its acquisition by Collibra.

North River has 40,000 employees across 20 countries and millions of accumulated SharePoint documents. Its legal and procurement teams need current answers about supplier termination clauses, agreements expiring this or next quarter, and partnership exclusivity terms. Those questions require more than finding text that resembles a query: the system must find the right agreement and the version that still applies.

The scenario gives several teams different requirements:

  • Data and AI: answer accuracy.
  • AI engineering: a discoverable set of relevant files without sensitive content.
  • Knowledge management: prevention of sensitive-data leaks and redaction where required.
  • Legal operations: dependable answers for the business.
  • Enterprise transformation: a repeatable process that can support another 15 use cases.

A manually curated pilot satisfies these requirements once. Production needs a way to keep satisfying them as the source collection changes.

3:253:56
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:13 · section reference included

Document quality includes competing versions of the truth

Koss brings a structured-data background from Informatica and Databricks, where quality includes dimensions such as completeness, conformity and accuracy. Documents add a practical set of concerns: repeated information, conflicting information and freshness. An expired contract can be perfectly readable and easy to retrieve while still being the wrong basis for an answer.

Metadata connects curation to the chatbot. Tags describe what a document contains, allowing the downstream system to select relevant material and use those descriptions alongside its content. The intended input is therefore a collection of current, relevant, high-quality files enriched with metadata, with sensitive material excluded.

Manual maintenance becomes its own bottleneck. Koss describes a telecommunications company where six to seven subject-matter experts cannot keep up with curating 20,000 HTML files. The consequences extend beyond slow preparation: incorrect answers, sensitive information appearing in chatbots, and a collection that becomes brittle as it changes.

3:564:26
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

4:26 · section reference included

Turn 339 mixed files into a smaller legal collection

The demonstration starts with 339 ingested files from a mixed SharePoint collection: call transcripts, PDFs, spreadsheets and Word documents. Deasy exposes both file-level information and page- or chunk-level content. The first task is to classify this mixture so the legal material becomes selectable.

A taxonomy is a tagging tree. Its tags can come from several places:

  • Manual definitions: create a tag and supply a prompt describing how the model should classify a file.
  • Reusable definitions: select tags from a library or import an existing taxonomy and valid values through CSV.
  • AI suggestions: ask the system to propose tags and classification values when the categories are not already known.

Classification values are optional. Supplying them lets the team state its existing categories; generating them helps explore a collection whose structure is still unfamiliar.

The tree can make later classification conditional on an earlier result. First identify legal and contract files; then classify contract type within that branch, such as supplier contracts or NDAs. Other tags can apply unconditionally. This expresses a useful dependency: a question about contract type belongs to the contract branch rather than every engineering document or call transcript.

After selecting files and tags and generating metadata, the legal-and-contract category contains 68 files. Clicking that category filters the collection. Applying two filters brings the 339-file collection down to 20 files; the recording does not identify both filter criteria. The observable change is concrete: generated metadata turns an undifferentiated file collection into a smaller set the engineer can select.

Selected presentation frame from From Raw Documents to AI-Ready Data — Leo Platzer & Jeff Koss at 592 secondsOpen full source frame
The interface shows generated document categories, including 68 files labeled contracts and legal, alongside a filtered file list.

Why did a file enter that category? The interface exposes the passage that supported its classification. Thumbs-up and thumbs-down feedback can supply examples for few-shot learning. This puts a review step beside the generated label: a person can inspect the evidence and provide an example rather than accepting the category without explanation.

6:397:09
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

6:39 · section reference included

Check sensitivity, redundancy and freshness before delivery

Sensitivity detection supports both AI-based scanning and patterns interpreted with surrounding context. Koss’s example is a request to detect Finnish national IDs. The system generated a regular expression and added Finnish words that he believed were associated with the identifier. The intended mechanism is to combine a matching character pattern with contextual clues, reducing accidental matches; this anecdote does not establish the resulting false-positive rate.

Quality checks apply several distinct selection rules:

  • Duplicates and conflicts: group related material and use AI to suggest which items to retain.
  • Freshness: restrict files to a specified date range or number of days.
  • Required metadata: identify the files that satisfy the information requirements of the intended data product.

The retained collection becomes a data slice: files selected by explicit criteria for a particular use.

A second demonstration starts with a separate 14-file legal slice. A use-case description and the files’ contents guide suggestions for either a flat taxonomy or a tree with multiple levels. A person chooses which proposed tags belong, and the system also generates classification values. This makes the taxonomy specific to both the business question and the actual collection.

Generated metadata includes sensitivity findings, file- and chunk-level evidence, and model-returned confidence scores. Scores such as 95% or 100% describe the model’s reported confidence; the recording does not establish that they are calibrated probabilities of correctness. Their practical companion is the evidence: a reviewer can see the content behind the classification.

9:4910:19
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

9:49 · section reference included

A curated slice must follow changes in SharePoint

The data slice is maintained on a schedule. Workflows pick up files added to or deleted from SharePoint and update the selected collection. The SDK then makes that data product available to the chatbot, including metadata used when loading the vector database. Freshness therefore depends on recurring source updates reaching the downstream retrieval system; a scheduled refresh leaves an interval between a source change and its propagation.

Where does a source change travel? The flow below separates the maintained collection from its downstream use. Updating SharePoint alone does not update chatbot context: the scheduled slice and SDK delivery connect those stages.

Selected presentation frame from From Raw Documents to AI-Ready Data — Leo Platzer & Jeff Koss at 851 secondsOpen full source frame
A completed workflow diagram places a curated data slice between raw files, ingestion, and an AI application.

For the North River scenario, Koss reports reducing preparation from four months to days, improving chatbot accuracy through file selection and metadata, and reducing compliance risk by filtering sensitive material out of the slice. These are presented scenario outcomes rather than a documented customer benchmark. The mechanism is nevertheless specific: less manual discovery, additional metadata for retrieval, and exclusion of identified sensitive content.

How it fits togetherFrom source changes to refreshed retrieval data

Files are added or deleted.

Scheduled maintenance updates the selected data product; SDK delivery carries its metadata into the chatbot’s retrieval pipeline.

13:3314:03
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

13:33 · section reference included

Thirty percent bad data can occupy much more of the context

Platzer turns to the retrieval consequence. A query selects a limited number of results—the top K—from the search collection. Repeated information and conflicting versions can occupy those positions. His Slack example makes the temporal problem clear: someone states one thing today and changes it tomorrow, while both messages remain searchable. Similarity to the question does not by itself establish which statement still applies.

Duplicate handling must distinguish useful additional context from simple redundancy. Deasy retains material that contributes context and removes material considered redundant. That distinction matters when two versions look alike but describe a real change: deleting every near-duplicate could discard the information needed to understand the current state.

The share of poor-quality files in a corpus need not equal their share of retrieved context. Platzer gives a 1,000-file example in which 30% are stale or duplicative, yet up to 80% of an agent’s context can contain stale or redundant information. Retrieval selects relevant-looking results rather than sampling files uniformly, so repeated matches can crowd the limited context. The 80% figure is a reported possible outcome, not a fixed conversion from corpus contamination to context contamination.

On the same multi-hop RAG tasks, Platzer reports roughly doubled recall and a 10–15% improvement in accuracy or task completion after data-quality work. These were chatbot retrieval evaluations, not coding-agent evaluations; the recording does not specify the full protocol or whether the completion gain is relative or in percentage points. The results support testing curation as a retrieval intervention, without establishing the same numerical gain for every application.

Selected presentation frame from From Raw Documents to AI-Ready Data — Leo Platzer & Jeff Koss at 1000 secondsOpen full source frame
Two plotted retrieval-evaluation curves and an adjacent table are shown on the results slide.

Relevance criteria depend on the use case: a legal chatbot and another enterprise application need different files. Duplication and freshness remain concerns across those uses, although deciding what counts as useful history or obsolete information still requires care. Cleaning the collection changes what retrieval can supply before any change to the answering model.

15:0015:30
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

15:00 · section reference included

Give agents a folder map before they open the files

The final technical demonstration uses the SDK to fetch metadata for a SharePoint site, describe each folder, and create an index across folders. Claude Code or Codex can consult a Markdown context file to decide where to look before inspecting individual documents. The same metadata used for curation now supplies a navigation layer for an agent.

Selected presentation frame from From Raw Documents to AI-Ready Data — Leo Platzer & Jeff Koss at 1185 secondsOpen full source frame
A dark code editor is projected during the context-repository demonstration; the displayed text is not legible at this scale.

One folder contains 23 files. Without the generated description, the agent would need to inspect those files to decide whether the folder is relevant. The context repository mirrors the SharePoint structure and describes document types, key topics, folder purpose, when to use it, and example questions. This changes the first step from opening documents to choosing a promising folder using its summary.

How does metadata save repeated discovery? The diagram shows two stages: metadata first becomes a folder index and descriptions; the agent then uses that map to choose where to inspect content. The map helps route exploration rather than supplying every document’s answer. Platzer proposes improved performance and lower token cost from this approach, without reporting a measured coding-agent result.

How it fits togetherMetadata becomes a map for agent exploration

Fetch the metadata available for the site through the SDK.

A folder index and context descriptions let the agent choose a location before reading its individual files.

18:2018:50
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

18:20 · section reference included

Start with a business question and representative files

Koss closes with an offer of one-week proofs of concept built around a business problem and sample files. The proposed outputs are concrete: metadata for the files, a sensitivity report, a data-quality report, and a taxonomy shaped by the use case and the collection. That gives the initial experiment a broader scope than testing answer generation alone: it also examines whether the source documents can become a usable, maintainable input.

20:1220:42
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

20:09 · section reference included

Resources

Read the complete timestamped transcript
  1. 0:13

    All right. Hey, guys, let's get going. Can everybody hear me? All good? Perfect. All right. My name is Jeff Koss. I'm a staff customer engineer with Deasy Labs. Um, as you can see from the slide, it says, "Brought to you by Collibra." Collibra acquired Deasy Labs last summer. Uh, Collibra is more of a data governance tool for structured and AI, um, agents and models. So the gap that Collibra had was we didn't have unstructured, and that's really where,

  2. 0:43

    uh, Deasy shines. So today, uh, with me is Leo Platzer.

  3. 0:47

    Yeah. Hi, everyone. Uh, I'm Leo. I was previously the CTO at Deasy Labs, which, as Jeff has already mentioned, was acquired last year by Collibra.

  4. 0:56

    And you can see we're both twins as well. We have our hats that we're giving away at our booth as well. So come by and get them. First come, first served. So today, we're talking about, um, how we can go from a stalled POC to production with Deasy data curation. So we're really focused on unstructured data. So our demo scenario I typically talk about is a manufacturing, uh, company, North River Manufacturing. You can see that, you know, there's four-- forty thousand employees. They operate in twenty

  5. 1:26

    different countries. Uh, really what they're trying to do is they have millions of documents in SharePoint. They've accumulated them over years and years, and they need to figure out which ones are relevant for their chatbot. They're invested heavily in AI to improve operational efficiency. And their big objective this year is they wanna build a legal ops AI chatbot. Now, what they wanna do is they wanna help the legal and procurement teams answer questions around supplier contracts, partnership agreements,

  6. 1:56

    and it's gotta be real time. So some of the questions they're trying to answer are, what are the termination clauses for a specific supplier? Um, which agreements expire this quarter? Which agreements expire next quarter? And what partnerships have e- uh, exclusivity terms?

  7. 2:14

    So some of the folks in the teams that we've been working with, um, is Sarah. She's the head of data and AI. So she's responsible for the accuracy. Um, also, there's Jim. He's the AI engineer. He just needs to know which files do I need. I have tons and tons, thousands and thousands of files in SharePoint. I need to find just the relevant ones with no sensitivity. Priya is the head of knowledge management. She wants to make sure that no sensitivity is leaked, um, and make sure that if

  8. 2:44

    there, if there is specific data that's, uh, classified, that it gets redacted as well. Then there's Michael. He's the director of legal ops. He needs a chatbot to deliver accurate responses to his team. So he's from the business. And then finally, there's Jackson. So Jackson's enterprise transformation, uh, lead, and this is his flagship use case, but also he has fifteen other use cases that he needs to, to deal with, right? So he

  9. 3:14

    wants to have a framework, a repeatable, successful process. He doesn't just wanna do it for one time. He wants to do it for multiple different use cases.

  10. 3:25

    So what they're able to do so far, they were able to take forty, around forty different files, pre-curated files, and build a small little POC, small little mini pilot, where they can do, uh, parsing, chunking, OCR, embedding, and then put it in a vector database and a knowledge graph. And they got pretty good responses, right? Seems pretty easy. You can do that with forty files.

  11. 3:56

    But when they try to do it with many, many different files, with eighty thousand plus files, right, now all of a sudden, it's a little bit more challenging. And so some of the challenges they had were, "I can't discover the files that I need. I don't know where the sensitive data is in my-- on SharePoint. I can't trust the quality of the data." So my background, um, I used to work at both Informatica and Databricks. So for me,

  12. 4:26

    data quality is typically the six different dimensions of data quality: completeness, conformity, accuracy. But when you deal with documents, it's a little bit different. It's I wanna see is there duplicate information, conflicting information, right? Leo's gonna talk about-- he has a really good slide about showing why it's so important and also the freshness of the data. And also everybody's talking about context, right? So what we wanna do is when we tag the metadata, we wanna use that in our

  13. 4:56

    chatbot, in our vector database to make it more relevant, have better accuracy in the chatbot.

  14. 5:04

    So you can see some of the challenges they had, right? They're getting incorrect answers. There's sensitive information that's being leaked out in the chatbots. It takes a ton of time to go to production because they can't find the files. They're spending too much time looking for the files instead of writing code and doing things that they, they really like to do, right? It's difficult to maintain. I'm working with another telecommunications company and they have six to seven different SMEs and they're trying to curate

  15. 5:34

    twenty thousand HTML files and they can't keep up, right? So it's very brittle, very difficult. And what we do is we provide the ability to automate this entire process.

  16. 5:46

    So over here, now you can see they wanna, they wanna try to access and leverage over eighty thousand different files. So they went from forty files to over eighty thousand files and really what they wanna do is they wanna find which files are relevant for their chatbot, which have high quality- Which ones are up to date with knowledge. They don't want to use a, a version where the contract is, is old and expired, right? And they wanna enrich their chatbot with the metadata that

  17. 6:16

    they've tagged, and of course, they don't want any sensitive data. And so what we do over here is we provide the ability to do that, right? We can access the data, um, over here. We can understand what the data is. We can figure out which files are relevant, do the OCR, the chunking, and things like that.

  18. 6:39

    And so let me show you a little-- a brief demo. So over here, you see in SharePoint, much like what North River has, they have a whole SharePoint site full of different, you know, there's call transcripts, there's PDFs, there's Excel, there's Word documents, and I-- they don't know which ones are relevant, which ones are focused around legal. Now, over here, when we look at Deasy, we've already opened up the application right here, and I can see, you know, I've already ingested three hundred and thirty-nine files. And now I can

  19. 7:09

    see both the ch-- uh, page level or chunk level as well as the file level. I can see all the information right here.

  20. 7:18

    And then over here, this is what we call a taxonomy. Taxonomy is a grouping of tags or a tagging tree. So I've ingested the three hundred and thirty-nine files, and now I can either manually create a user-defined tag, I can leverage tags that are reusable from my tag library, or I can leverage AI to auto-suggest what tags I should even have, right? I can leverage AI to say, "You-- these are the tags are relevant." So here I just type in my tag name, data, uh, document categorization, give it a prompt, and this prompt is what

  21. 7:48

    we're gonna send to the LLM to determine, is this file-- how should we tag this file? Now, if I know what I'm looking for, I can put in the classification values. It's optional. And again, if I-- I can leverage AI to auto-suggest all these tags for me.

  22. 8:08

    Now, also the other thing to point out, if I have my, uh, my taxonomy here, my valid values, um, already, I can upload them through a CSV. And so once, once we do this now, I can see I started to build this, this tagging tree. Now, the other thing I can do is I can have conditions on my tagging tree. So remember I said I only want legal and contracts? That's what they're looking for. Now what I can do is I can ha-have, um, sub-tagging. So only legal and contract files, I wanna

  23. 8:38

    pull out more data around the contract type. Is it a supplier contract? Is it part of an NDA, right? Is it materials and, and time contract?

  24. 8:53

    And now what we can do when we add this, you can see we have different branches in our tagging tree.

  25. 9:01

    And I could have, um, unconditional tags as well. I can have any number of different tags. Now what we're doing is we're gonna generate the metadata. So I choose what files I want. I can choose what tags I want. I can say Generate metadata, and we'll run this, and you'll see

  26. 9:19

    the results over here. So now I can see the co-- document category. I can hover my mouse. There's engineering documents. There's my contracts and legal, right? I have sixty-eight different files. And if I click on any of those, it automatically filters for me. Now, I can also see the evidence. Why did AI tag this? Where in the document did it find evidence that it's a legal contract, right? And I can also do thumbs up and thumbs

  27. 9:49

    down, so there's few-shot learning as well. So I can give it an example if I want to. Now, you can s-- notice I ha-- also have filters. So I have two different filters. I quickly went from three hundred and thirty-nine files down to twenty files. So we also have the ability to scan for sensitivity as well. So whether it's PII, PHI, um, could be anything. I was in, uh, in London at a Gartner conference, uh, a couple months ago, and there were

  28. 10:19

    some folks from Finland that came up, and they wanted to know, you know, "Can we find Finnish national ID?" I don't speak Finnish, so I put it in there. I did this demo. And not only can we do AI-based, we can do pattern-based with context. And it ended up figuring out the, the regex, um, syntax for me, but it also added some Finnish words which I thought were probably associated with Finnish national ID. So it's gonna be less false positives as well.

  29. 10:52

    All right. Let me keep going here. So the other thing we can do is we can do quality. So again, for us, we have a dashboard in here for quality where we can go through and see what are the different groupings for, um, for duplication, for conflicting information, um, and then we'll, we'll leverage AI to auto-suggest which ones to keep. Now, the other thing we can do too is, again, for freshness, we can say, "I only want files within X number of days or within a specific

  30. 11:22

    date range." And then the last thing we can do for data quality is, you know, here's the, the pieces of metadata that I absolutely have to have, and then what we'll do is we'll figure out, you know, of the three hundred and thirty-nine files, maybe you really need two hundred files. And then from there you can create a data slice. And then we can also give it context as well, uh, over here. So I'll go through one more video, and we also have a live demo for you that Leo's gonna be doing in a

  31. 11:52

    second. So in this case now, what I can do is I can, I can leverage a, a data slice. So my data slice could be filtered based on the specific criteria that I want. So in this case, I have fourteen files. These are only the files that are relevant for legal. And now what I'm doing is I'm copying my use case description. What I can do is I can say, "Here's what I'm trying to build," and I can either do a flat, uh, taxonomy or multiple layer levels in my tagging tree.

  32. 12:23

    And then over there, when I run it, it's gonna auto-suggest these tags for me. There's human in the loop, so I can choose which ones I like, which ones I want, which ones are relevant. But again, it's all done for me based on what I'm trying to do. We gave a context and based on the content of the files themselves. And then over here, it even builds out the classification values for me.

  33. 12:50

    And then if we want, when we run-- when we generate the metadata, we're in the job, now we can see over here, here's all the sensitive information, right? Here's all the different information that we have over here.

  34. 13:04

    And then again, I can see the evidence at the file level and the chunk level. I can see the confidence score. So coming back from, from AI, what the model returned, I don't have to guess. You know, it's a hundred percent confident, ninety-five percent confidence, here's why. So you can have ver-- a lot of confidence in the, the, um, AI-ready data product that you guys are gonna be building.

  35. 13:30

    All right.

  36. 13:33

    And now finally, what we can do, uh, for delivery and maintenance. So, you know, one other thing that, that's really important is with our data slice technology, we have workflows where we can put it on a set schedule, so it's auto-refreshed. So we wanna make sure that your chatbot never gets stale. So if there's new records or, sorry, new files that get added to SharePoint or deleted, we'll automatically pick those up and add those on a set schedule to my, um, uh, to my data slice or my AI-ready data

  37. 14:03

    product. And then our SDK will pick that up in the chatbot, and then when we load the vector database, you'll always have the most freshest, uh, metadata.

  38. 14:14

    So what's it mean for North River? So basically, they were able to go from four months of data prep work to down to days. They were able to increase their data utilization for AI, they increased their chatbot accuracy because they not only were able to figure out which files to use, but also the metadata from the files, they were able to inject that and leverage that through our SDK in their chatbot, right? And then the other thing they did was they reduced their compliance risk. They knew where all the sensitive

  39. 14:44

    data is, and they made sure that they were not part of-- they were all filtered out in their, um, data slice. Okay, so I'm gonna turn it over to my brother from another mother- -Leo, and he's gonna talk about, um, context.

  40. 15:00

    Yeah. Hi, everyone. Um, I just wanted to basically, like, give the so what, right? Technically, okay, you deduplicate your data, you reduce your search space, right? You clean it up, you have it fresh. What does it actually mean, right? So if we were looking at the vector space, right, when you ask a question, it comes in, and you have duplicated data, the same information, or potentially even a different version of the document with conflicting information now, right? Think of

  41. 15:30

    your Slack messages. Today a person says this, tomorrow a person says that, right? The truth changes over time. If it cha-- if it stays, however, in your knowledge base, in your search base, your top K will always be full of pretty much irrelevant information, right? So in Deasy, for example, something which is considered duplicative is either another piece of context, like your database has changed and I can actually use it, or it is simply redundant,

  42. 16:00

    and we will remove it from the knowledge base.

  43. 16:04

    The second piece is, like, same as-- same reason as before. If your ma-- information is old, right? If you suddenly are working with data from before the two thousands, if you're working with information which was only relevant last week, is not relevant this week, in the future, some of the use cases when agents actually get into production for, let's call it, the most high-stakes use cases, that has detrimental effects. And we have done, um, a lot of

  44. 16:33

    evaluation around what is the effect of making sure that your unstructured data quality is maintained, is consistent, and is always handled deterministically in the same way, right? And what we're essentially seeing is that if your raw corpus... Okay, let's imagine you have a thousand files, and thirty percent of your files are stale or duplicative. It can happen that up

  45. 17:03

    to eighty percent of your actual context your agent uses to answer a question is filled up with stale information. Which means, like, it actually uses eighty percent of completely redundant, useless information and knowledge, and this has to be managed. Um, what we have seen is that by running on the exact same tasks, so this is a RAG chatbot. It's not, um, running an actual, like, let's say, Claude Code

  46. 17:33

    task or Codex on a coding task. It was simply, like, multi-hop RAG evaluation. And what we saw is that if we look at recall one, like, I mean, we pretty much double, uh, the recall. Um, and overall, improve our accuracy, so, like, task completion by o-- between ten and fifteen percent. And why is this so significant? This is so significant because this has nothing to do with the use case you're working with. This is true

  47. 18:04

    independent of what you're actually using the data for, right? 'Cause if you actually know what you wanna use your data for, you can apply relevancy and all nod other data quality dimensions, right? But duplication and freshness is true no matter what.

  48. 18:20

    And yesterday, right, there is a lot of context topic, a lot of context flags around here. Um, I wanted to give a very, very quick demo on how you would use the DZI SDK to prepare a context repository, and I'm gonna run this live. So what this is doing is essentially it's gonna fetch all of the metadata we have about a given SharePoint site, and we will then contextualize each folder within that

  49. 18:50

    SharePoint site, as well as create an index over all folders so that your Claude Code, for example, or your Codex, does not even have to look at the individual files to figure out what to look for, but simply looks at the context MD file. So what you're seeing here is essentially purely by using DZI, you can curate context files where you are directly aware what exactly is within each

  50. 19:20

    folder, right? What is its purpose, right? When do you wanna use it? So for example, this is one-to-one the mirror of what I actually have in SharePoint. But now, instead of actually looking at these 23 files to figure out whether I even should look at this folder, I know exactly what document types I have, what it is about, what are the key topics, when to use this folder, and what are example questions.

  51. 19:48

    So this is how-- this is all the kinds of use cases today, right? How you can use DZI for use cases like your RAG chatbot, or use DZI to actually improve the performance and reduce, for example, the token cost for your agents. Um, but yeah, basically I'll move over to Jeff again to summarize this whole presentation.

  52. 20:09

    Awesome.

  53. 20:12

    Thanks, Leo. All right, so guys, um, we're almost out of time. Our booth is way in the back by stage two. Please come by, bring any different use cases that you guys have. Um, we're happy to have discussions, do live demos, um, talk about what your challenges are around unstructured documents. Um, over here, um, before we leave, just wanna let you guys know we have one-week POCs, so we're happy to prove our technology.

  54. 20:42

    Uh, again, you know, you guys bring your, your business problem, um, you're trying to solve with AI, bring some sample data, a bunch of different files, and then what we can do is we can get-- generate metadata for each different files. Um, we could get a full sensitivity report on your, on your data, data quality report as well. And then we'll show you how we can build a full taxonomy based on your specific use cases, your specific files. So with that, everybody, appreciate you guys-

  55. 21:11

    Thank you so much

  56. 21:12

    ... spending time with us. And lastly, we only have one more case of these hats, so come on by and grab 'em. So we'll be out in the back. Thanks, guys. Appreciate you.