Aleix Conchillo Flaqué is a software engineer at Daily building Pipecat, the open-source framework for real-time voice and multimodal agents. His work connects conversational models with the software needed to use them: streaming media, interchangeable services, interruption handling, and tests that exercise an actual spoken exchange.
From language libraries to real-time systems
Conchillo’s open-source work stretches back to the early 2000s. His Simple C Expat Wrapper (SCEW) gives C programmers a document-oriented interface to XML. Instead of writing parser callbacks for each application, developers can load, inspect, modify, and create documents through a reusable library.
He also developed guile-json, whose implementation dates back to 2013. It translates JSON into native Guile Scheme values and back again, supporting individual documents and sequences of documents. His Guile Homebrew tap makes libraries and applications in that ecosystem easier to install on macOS; the Spritely Institute uses it as an installation route for Hoot, its Scheme-to-WebAssembly toolchain.
In 2019, he and a colleague added WebRTC monitoring to the Rumpus collaboration application. They streamed connection statistics from its Electron client through a gRPC backend into Prometheus and Grafana, making bitrate and packet-loss problems visible during meetings. This work brought the quality of a live media connection into the debugging process—a concern that carries into conversational AI.
Composing and testing voice agents
At Daily, Conchillo builds the infrastructure around conversational models. He co-authored the company’s January 2025 account of its NVIDIA voice-agent collaboration, which combines Pipecat orchestration with NVIDIA NIM services. One important detail is conversation memory: when a person interrupts an agent, its context should reflect what was actually played aloud, rather than treating every generated word as something the person heard.
Two aspects of his work make that infrastructure concrete:
Composable multimedia pipelines: Conchillo’s approach treats a voice agent as a sequence of processors that receive and pass along audio, video, text, or control information. A conventional pipeline connects speech recognition, a language model, and speech synthesis. A native speech-to-speech model can combine those stages, while recording and custom processing remain separate parts of the application. His explanation of Pipecat’s architecture emphasizes interchangeable services and the ability to insert custom processors—for example, to inspect model output before it is spoken. The framework supplies the assembly points; developers decide how their agent should respond and what checks it needs.
Voice-agent release testing: Conchillo built release evaluations in which one bot speaks to another. An evaluation bot asks a simple arithmetic question, listens to the reply, and checks the answer. These bot-to-bot evaluations exercise connected services through a spoken exchange, reducing the repetitive manual work of running examples before a release. The test reaches beyond a language model’s text response: the question and answer must also travel through the audio pipeline.
Taking coding agents away from the desk
Conchillo developed the Pipecat MCP Server to make it possible to talk with a coding agent while away from the desk. Rather than requiring someone to remain at a terminal to interact with tools such as Claude Code, Cursor, or Codex, the project adds voice as another way to stay involved in the work.
The server exposes voice and screen-capture tools to MCP-compatible agents, while a separate browser, WebRTC room, or telephone connection carries the media. Local Whisper transcription and Kokoro speech synthesis provide an option without hosted speech services. This separation lets a coding agent retain its working environment while the person speaking with it uses a different connection. Conchillo extends Pipecat’s composable approach from building conversational applications to making development tools accessible through conversation.
A voice bot needs more than speech generation: this workshop builds the audio pipeline, then explores interruptions, context, tools, testing, and selective routing between agents.