Voice AI models and infrastructure
Gradium AI
Gradium AI builds voice models and infrastructure for developers and enterprises creating conversational applications. Its APIs provide streaming speech recognition and synthesis, voice cloning, and real-time translation, supporting uses from customer support to game characters. Gradium Translate handles speech-to-text and speech-to-speech translation across English, French, German, Spanish and Portuguese. Gradbot, its open-source prototyping framework, coordinates speech recognition, language-model inference and speech synthesis through a Rust engine that manages turn-taking and interruptions.
Founded in 2025 by Neil Zeghidour, Laurent Mazaré, Olivier Teboul and Alexandre Défossez, Gradium grew out of the team’s work at the nonprofit research lab Kyutai. Zeghidour serves as CEO. The founders’ earlier research included neural audio codecs and audio language models; Gradium applies that background to production voice systems. Its Phonon model uses approximately 100 million parameters, continuous audio language modeling and flow matching to synthesize speech entirely on a device, enabling offline applications. Its semantic turn detection estimates whether a person has finished a thought rather than relying only on silence.
In July 2026, Gradium extended its seed financing to $100 million, adding NVIDIA as an investor and announcing a San Francisco Bay Area office. The company reported enterprise customers across customer experience, healthcare, media, AI agents and consumer applications, with revenue beginning within weeks of launch.
1 talk
Newest first1 speaker at AIE
Affiliations reflect their AIE appearances, not necessarily current employment.
Messages from the stage
Speech synthesis on a local CPU
Zeghidour introduced Gradium Phonon, an on-device text-to-speech model that runs locally on a CPU, adding a concrete deployment example to the discussion of conversational voice systems.
Affiliations reflect each recorded session, not necessarily current employment.
Company sources · checked 2026-08-27
- Gradium: Advanced Voice AI for Text to Speech, Speech to Text, and Voice Cloning
- Blog: Voice AI Research, Engineering, and Product Updates
- Gradium: Solving voice
- Gradium Extends Funding to $100 Million and Expands to Silicon Valley
- Phonon update: 1.00% WER on Seed-TTS, smaller than every model we beat
- gradium-ai/gradbot: Open source framework to vibecode and prototype voice agents with Gradium APIs
- Gradium adds $30M from Nvidia, pushing seed round past $100M
