← All speakers

Bio, Work & Ideas

Leo Platzer

Conference affiliation: Collibra

On this page

Leo Platzer is the co-founder and former CTO of Deasy Labs, which Collibra acquired in 2025. His work centers on making enterprise documents usable by AI: organizing their contents, adding meaningful metadata, and helping teams select relevant, current information while filtering sensitive material.

From data-quality engineering to Deasy Labs

Platzer’s engineering background includes Amazon, Mercedes-Benz, and the chatbot startup E-Bot 7. At QuantumBlack, McKinsey’s AI business, he worked as a machine-learning software engineer on an AI-powered data-quality product. He subsequently co-founded Deasy Labs with Reece Griffiths and Mikko Peiponen, who had also worked on machine-learning data governance at McKinsey.

Founded in 2023 and initially called Deasie, the company joined Y Combinator’s Summer 2023 batch. It applied the founders’ data-governance experience to language-model applications built on internal knowledge, focusing on preparing unstructured content—documents, emails, and transcripts—for particular use cases.

As CTO, Platzer helped build a product around metadata orchestration: creating useful descriptions and classifications of content, then making those attributes available to downstream AI workflows. A document’s presence in a repository does not establish whether it belongs in an assistant’s answer. Teams also need to understand its relevance, sensitivity, quality, and freshness.

Turning document collections into usable knowledge

The platform Platzer co-founded connects document processing with the decisions teams need to make before deploying an AI application:

  • Use-case-specific data preparation: Deasy connects to sources such as SharePoint and S3, extracts and chunks their contents, and enriches files with metadata. Teams can assemble datasets using topic, time, relevance, quality, and sensitivity, giving them more control over what enters a retrieval pipeline.
  • Suggested taxonomies: Automated discovery proposes categories from the content itself, helping teams organize unfamiliar collections before applying labels across a repository. In Platzer and Jeff Koss’s joint demonstration, metadata tagging includes evidence and confidence scores, giving teams a basis for assessing the classifications.
  • Maintained data slices: Deasy monitors source systems and refreshes selected datasets as content changes. It can also write metadata back to the original systems, making enrichment useful beyond a single AI application. Checks for duplicates, conflicting information, and freshness help determine which documents should remain available to an assistant.

Platzer’s 2026 presentation with Jeff Koss illustrates why this preparation matters through a manufacturer’s legal-operations chatbot: a pilot works on 40 hand-picked files but struggles when expanded to 80,000 SharePoint documents. The example connects document governance to retrieval quality rather than treating labeling as an end in itself.

In his portion of the presentation, Platzer explains how stale and duplicated documents can crowd useful information out of an agent’s retrieved context. He reports that a corpus containing 30% stale or duplicated material could yield context that was up to 80% useless, and that cleaning the data roughly doubled recall in the multi-hop retrieval evaluations he presented. These results illustrate a specific failure mechanism: defects in the source collection can become disproportionately prominent in the information an agent receives.

Bringing unstructured-data preparation into Collibra

Collibra’s acquisition of Deasy Labs brought the startup’s automated discovery, classification, filtering, and enrichment capabilities into a broader data-governance platform. Collibra released Unstructured AI, incorporating Deasy Labs, in October 2025.

Platzer’s work connects data-quality engineering with the practical requirements of enterprise AI. The product he helped build gives teams a way to choose and maintain the knowledge their applications use; his retrieval analysis explains why those choices can materially affect the answers they produce.

Read the topics behind these talks

1 conference talk

Key ideas

Scroll to read ↓

Jeff Koss and Leo Platzer show how document classification, sensitivity checks and scheduled curation turn a small legal-ops chatbot pilot into a maintainable retrieval system—and why duplicates can consume far more context than their share of the corpus suggests.

  • A successful 40-file pilot can hide the work of discovering relevant documents, excluding sensitive information and choosing trustworthy versions at production scale.
    3:25 ↗
  • Conditional taxonomies classify a broad category first, then extract more specific metadata within its branch; evidence and human feedback make those labels reviewable.
    8:08 ↗
  • A data slice combines selection criteria with scheduled maintenance, carrying source additions and deletions through to downstream retrieval data.
    10:52 ↗
  • Stale and duplicated material can occupy a disproportionate share of retrieved context because top-K retrieval selects matches rather than sampling the corpus uniformly.
    15:00 ↗
  • Folder-level context files reuse metadata to help agents choose where to explore before spending tokens inspecting individual documents.
    18:20 ↗