Serving Voice AI at $1/hr: Open-source, LoRAs, Latency, Load Balancing
AI Engineer World's Fair 2025 · 16:09
Real-time multimodal AI infrastructure
Gabber built infrastructure for developers creating AI applications that process live audio, video and text. Its real-time AI engine supported assistants, agents and companions that could receive microphone, camera or screen input and respond through models and tools. Gabber Cloud added hosted orchestration and inference, while the source-available engine supported local deployment and connecting developers’ own models.
Founded by Jack Dwyer, Gabber organized applications as graphs of processing nodes linked through typed connections. Developers could combine transcription, model responses and external API calls, reuse subgraphs, and use state machines to control application behavior. Its technical focus was continuous media processing: multiple models could run in parallel while maintaining state across streams. JavaScript/TypeScript, React and Python SDKs supported integration into other applications.
Gabber’s team subsequently joined LiveKit through a team acquisition. In the joining announcement, Dwyer reported that Gabber had powered millions of voice and vision conversations over the preceding year and a half, using LiveKit throughout. Jack Dwyer, Neil Dwyer and Brian Blumenfeld joined the LiveKit team, with Jack focusing on developer experience.
AI Engineer World's Fair 2025 · 16:09
Affiliations reflect their AIE appearances, not necessarily current employment.
The talk covers latency metrics and leading silence alongside GPU deployment, vLLM-based batched inference with LoRA adapters, and load balancing.
Affiliations reflect each recorded session, not necessarily current employment.