← All organizations

Video understanding AI

TwelveLabs

TwelveLabs builds AI models and tools that let developers and enterprises search, summarize, and analyze video libraries. Its Marengo embedding model makes video content searchable across visual, audio, and speech signals, while Pegasus converts footage into structured information such as scenes, entities, and temporal segments. Developers access capabilities through REST APIs and Python or Node.js SDKs. Jockey, available in research preview, extends this to questions spanning video and image collections: it plans multiple steps and returns answers grounded in cited moments.

Founded in 2021, TwelveLabs grew from work analyzing video during military service. Its founding team comprised CEO Jae Lee, Aiden Lee, SJ Kim, Dave Chung, and Soyoung Lee. Its research focuses on preserving the different signals within video: Marengo 2.7 introduced multiple specialized vectors for appearance, motion, on-screen text, and speech instead of compressing everything into one embedding. This representation supports searches for details within footage, including small objects.

In 2024, the company reported more than 30,000 developers using its platform, spanning individual experimentation and enterprise integrations. TwelveLabs raised a $100 million Series B in 2026, co-led by NEA and NAVER Ventures, with Amazon among the participants, to support research, development, and geographic expansion.

www.twelvelabs.io

1 talk

Newest first

1 speaker at AIE

Affiliations reflect their AIE appearances, not necessarily current employment.

Messages from the stage

From video representations to searchable context

Le described video as a multimodal spatiotemporal volume, presenting Marengo embeddings, Pegasus video-language reasoning, API-accessible context graphs, and Jockey. The talk also covered knowledge stores, corpus digests, and agentic search.

Affiliations reflect each recorded session, not necessarily current employment.

Company sources · checked 2026-08-28