
Speech-to-Speech Model Research at Google DeepMind — Valeria Wu Fon & Tom Ouyang, Google DeepMind
Valeria Wu Fon and Tom Ouyang describe Gemini speech-to-speech research, tracing the transition from specialized speech-recognition pipelines to models trained jointly on audio, video, and text. They frame the central challenge as balancing conversational latency, task intelligence, and multimodal interaction. Embedded demonstrations illustrate streaming…
Valeria Wu Fon · Tom Ouyang