← All organizations

Voice AI models and agent infrastructure

Cartesia

Cartesia builds voice AI models and infrastructure for developers creating real-time conversational applications. Sonic generates speech from text, Ink transcribes streaming speech, and Managed Agents provides a platform for building and shipping enterprise voice agents. Applications include customer support, sales, and recruiting. Developers can access models through APIs, with deployment options spanning cloud, on-premise, and on-device environments.

Founded in 2023 by Karan Goel, Albert Gu, Arjun Desai, Brandon Yang, and Christopher Ré, Cartesia grew out of research at Stanford AI Lab. Goel is CEO and Gu is Chief Scientist. Its technical foundation includes structured state-space sequence models, which maintain compressed information over long contexts for efficient processing. Its H-Nets research collaboration extends this approach to hierarchical representations: dynamic chunking learns to segment raw data into meaningful units, enabling language modeling directly from bytes rather than relying on fixed tokenization.

Cartesia served more than 50,000 customers by March 2026. In 2025, the company reported that thousands of businesses, including ServiceNow, Cresta, and Decagon, used Sonic to power millions of conversations monthly. Its October 2025 financing milestone brought $100 million from investors including Kleiner Perkins, Index Ventures, Lightspeed, and NVIDIA.

www.cartesia.ai

2 talks

Newest first

2 speakers at AIE

Affiliations reflect their AIE appearances, not necessarily current employment.

Start here

  1. State Space Models for Realtime Multimodal Intelligence

    Start with Goel's talk to understand why instant responses on low-power devices motivate a shift from cloud-based batch processing to streaming systems.

    Karan GoelAI Engineer World's Fair 2024

  2. Serving Voice AI at Scale — Arjun Desai (Cartesia) & Rohit Talluri (AWS)

    Follow with the interview to learn how state-space-model inference fits into enterprise voice serving, with audience questions extending to local models and video-model research.

    Arjun Desai · Rohit TalluriAI Engineer World's Fair 2025

Messages from the stage

Compressed memory for streaming intelligence

Goel describes state space models and Mamba as alternatives to architectures that repeatedly inspect extensive context, emphasizing compressed internal memory and streaming token updates.

Voice serving beyond model inference

The discussion between Desai and Talluri covers Sonic 2, voice customization and cloning, and edge deployment alongside latency tradeoffs across the speech-to-text and language-model pipeline.

Affiliations reflect each recorded session, not necessarily current employment. The voice-serving session is a joint Cartesia–AWS discussion.

Company sources · checked 2026-08-27