← All speakers

Bio, Work & Ideas

Jeff Koss

Conference affiliation: Collibra

On this page

Jeff Koss is a customer engineer whose career connects enterprise data warehousing, integration platforms, and the preparation of business documents for AI. His work centers on helping organizations make information usable beyond a carefully selected pilot dataset, drawing on experience at Informatica, Databricks, and Deasy Labs.

From enterprise data platforms to AI

Koss’s early work included application development, data warehousing, and business intelligence. He joined Informatica in 2001 and spent more than 22 years there, in roles including Solutions Engineer, Senior Product Specialist, and Principal Product Specialist. He supported customers across financial services, healthcare, oil and gas, and manufacturing. By the middle of the 2010s, his focus included big data, virtualization, and metadata management: helping organizations connect information and understand what it contains.

He subsequently worked as a Solutions Architect at Databricks supporting regulated industries, extending his customer engineering work into a platform combining analytics, machine learning, and generative AI. His later work with Deasy Labs brought that experience to unstructured information—documents whose usefulness depends on more than whether a search system can find them. In 2026, he held the role of Staff Customer Engineer at Deasy Labs (Collibra).

His experience with PowerCenter informs a practical concern about enterprise architecture: disconnected tools complicate the work of bringing AI, analytics, and data warehousing together. In his reflection on joining Databricks, Koss argued for connecting those workloads rather than making customers assemble them across separate systems. That position gives his career transitions a technical throughline: the challenge is both preparing information and making it usable across the systems that depend on it.

Making documents usable for AI

Koss’s work with Deasy Labs addresses unstructured data for AI, where the contents and condition of a document collection shape what an application can reliably retrieve. A legal-operations chatbot example developed with Leo Platzer makes the problem concrete: a pilot using 40 hand-picked files worked, while expanding it to 80,000 SharePoint documents exposed sensitive material, duplicates, conflicting information, and outdated content. Moving to an enterprise collection required more than increasing the number of searchable files.

The unstructured-data technology at Deasy Labs and Collibra connects several forms of preparation:

  • Metadata with evidence: AI-assisted taxonomies and tagging organize documents into meaningful categories. Evidence and confidence scores accompany the metadata, helping users assess classifications rather than treating every generated label as equally reliable.
  • Sensitive-data filtering: Detecting sensitive content allows teams to exclude unsuitable material before it reaches a chatbot. The useful dataset depends on the application’s needs, rather than simply including every available file.
  • Document quality and freshness: Checks for duplicates, conflicts, and outdated content help teams select relevant material. Scheduled refreshes keep those data slices current as the underlying collection changes.

These mechanisms extend a concern running through Koss’s work in integration and metadata management: information must be organized and maintained so that downstream applications can use it appropriately.

Read the topics behind these talks

1 conference talk

Key ideas

Scroll to read ↓

Jeff Koss and Leo Platzer show how document classification, sensitivity checks and scheduled curation turn a small legal-ops chatbot pilot into a maintainable retrieval system—and why duplicates can consume far more context than their share of the corpus suggests.

  • A successful 40-file pilot can hide the work of discovering relevant documents, excluding sensitive information and choosing trustworthy versions at production scale.
    3:25 ↗
  • Conditional taxonomies classify a broad category first, then extract more specific metadata within its branch; evidence and human feedback make those labels reviewable.
    8:08 ↗
  • A data slice combines selection criteria with scheduled maintenance, carrying source additions and deletions through to downstream retrieval data.
    10:52 ↗
  • Stale and duplicated material can occupy a disproportionate share of retrieved context because top-K retrieval selects matches rather than sampling the corpus uniformly.
    15:00 ↗
  • Folder-level context files reuse metadata to help agents choose where to explore before spending tokens inspecting individual documents.
    18:20 ↗