← All organizations

Multimodal data infrastructure and vector search

LanceDB

LanceDB builds data infrastructure for AI teams to store, search and prepare multimodal datasets for model training and applications. Its platform brings raw data, metadata and embeddings together, supporting dataset curation, feature engineering, and vector, full-text and hybrid search with SQL filters. LanceDB OSS is an Apache 2.0 licensed embedded retrieval library; LanceDB Enterprise provides a commercial multimodal lakehouse deployable in a private cloud or customers’ own cloud environments. Applications include retrieval-augmented generation, recommendation systems and agent memory.

Founded in 2022, LanceDB is led by co-founders Chang She, CEO, and Lei Xu, CTO. Its technical foundation is Lance, a columnar format supporting random access and shuffling for multimodal training data, including images and point clouds. LanceDB builds disk-based search indexes on that format; both are written in Rust. Apache Arrow interoperability and automatic versioning help teams work across data tools and track dataset changes. Lance introduced community governance in 2025, separately from the LanceDB project.

In June 2025, the company reported more than 20 million downloads of its open-source packages over the preceding year, with LanceDB Enterprise customers including Runway, Midjourney and Character.ai searching tens of billions of vectors and managing petabytes of training data. That year, it closed a $30 million Series A led by Theory Ventures. By August 2026, LanceDB also shipped as CrewAI’s default vector backend for agent memory.

www.lancedb.com

2 talks

Newest first

1 speaker at AIE

Affiliations reflect their AIE appearances, not necessarily current employment.

Start here

  1. Scaling Enterprise-Grade RAG Systems: Lessons from the Legal Frontier

    Start here to learn how a multimodal lakehouse brings embeddings, text, and other AI data together for enterprise legal retrieval.

    Calvin Qi · Chang SheAI Engineer World's Fair 2025

  2. The Hierarchy of Needs for Training Dataset Development

    Read this for the research workflow around training data: Spark, Trino, GPU-backed synthetic-data workflows, and dataset materialization.

    Chang She · Noah ShpakAI Engineer World's Fair 2024

Messages from the stage

Legal retrieval under operational constraints

The discussion with Calvin Qi connects domain-specific evaluation, confidentiality, customer isolation, and retention requirements with the tradeoff between ingestion throughput and query latency.

Storage operations for training datasets

The discussion with Noah Shpak examines how Lance supports row-reference shuffling, random access, scans, large-binary streaming, and zero-copy schema evolution for multimodal dataset development.

Affiliations reflect each recorded session, not necessarily current employment.

Company sources · checked 2026-08-28