← All organizations

Voice AI infrastructure

Deepgram

Deepgram provides speech recognition, speech synthesis, audio intelligence and voice-agent infrastructure for developers, product teams and enterprises. Its Voice Agent API combines speech-to-text, text-to-speech and language-model orchestration, letting teams build conversational applications through one interface. Users can process live or recorded audio, extract summaries and sentiment, and deploy through cloud or self-hosted APIs. Its Flux speech platform handles conversational turn-taking and interruptions. Deepgram also acquired OfOne in 2026, whose restaurant voice technology now anchors Deepgram for Restaurants.

Founded in 2015 by Scott Stephenson, Adam Sypniewski and Noah Shutty, former University of Michigan physicists, Deepgram grew from an effort to search audio captured by wearable recorders using deep learning. Stephenson remains CEO, while Sypniewski leads research and engineering as CTO. Its engineering work extends beyond transcription to deciding when an agent should respond: researchers evaluate Flux on complete conversations, using sequence alignment of transcripts and turn-boundary tokens to measure turn-detection accuracy and latency while accounting for dropped turns.

In August 2026, the company reported more than 200,000 developers and 1,400 organizations using its platform, with cumulative processing exceeding 50,000 years of audio and one trillion transcribed words. Deepgram raised $130 million in Series C funding led by AVP in January 2026 at a $1.3 billion valuation.

deepgram.com

2 talks

Newest first

2 speakers at AIE

Affiliations reflect their AIE appearances, not necessarily current employment.

Start here

  1. Giving a Voice to AI Agents

    Start here to understand the shift from earlier assistants to open-ended voice AI and the components needed to support it.

    Scott StephensonAI Engineer World's Fair 2024

  2. Building & Scaling an AI Agent Swarm of low latency real time voice bots!

    Use this workshop for a practical client/server starting point and an introduction to API access and telephony deployment.

    Damien MurphyAI Engineer World's Fair 2024

Messages from the stage

Latency within a modular architecture

Stephenson highlighted a Daily implementation using Deepgram and Llama that approaches 500-millisecond response times, alongside discussion of multimodal context, enterprise controllability, and operating costs.

Operational choices for conversational systems

Murphy's workshop addressed audio bandwidth, function calling, agent routing, and monitoring, including the tradeoffs involved in speaker identification.

Affiliations reflect each recorded session, not necessarily current employment.

Company sources · checked 2026-08-27