Agents at Scale: Inside MiniMax's Model and the Infrastructure Behind It
AI Engineer World's Fair 2026 · 20:14
AI cloud infrastructure and model inference
Together AI provides cloud infrastructure for developers and enterprises to train, customize and run open AI models. Together Inference offers serverless access, asynchronous batch processing and dedicated model deployments. Its fine-tuning service lets teams adapt models without managing training infrastructure, while Accelerated Compute supplies GPU clusters for larger workloads. The platform also includes managed storage and code sandboxes for AI applications and agents.
Founded in 2022 by Vipul Ved Prakash, Ce Zhang, Percy Liang and Christopher Ré, Together AI is led by Prakash as CEO and Zhang as CTO. Tri Dao joined as chief scientist in 2023 and is now also listed as a founder. Its research includes Dao and collaborators’ FlashAttention, which accelerates exact attention and reduces memory use by reorganizing computation into blocks. FlashAttention-2 improves GPU parallelism and work partitioning to make model training and inference more efficient.
The business serves both application builders and teams developing their own models. In July 2026, the company reported over one million developers and thousands of paying customers, including Cursor, Cognition and Decagon. It also announced an $800 million Series C at an $8.3 billion post-money valuation. Customers can customize models with proprietary datasets, then host the resulting models on Together or download them.
AI Engineer World's Fair 2026 · 20:14
AI Engineer Europe 2026 · 24:35
AI Engineer Europe 2026 · 15:50
AI Engineer World's Fair 2025 · 18:47
Affiliations reflect their AIE appearances, not necessarily current employment.
Learn how UPipe processes attention heads in smaller groups to reduce activation memory, with a five-million-token training example on an eight-H100 node.
Max RyabininAI Engineer Europe 2026
Learn what an approximately 100 ms P90 transcription-completion target demands, and how deployment topology, routing, and function-calling evaluations enter voice-agent engineering.
Rishabh BhargavaAI Engineer Europe 2026
See how Hallmark combines visual references with rules against recurring design problems, and how persistent preferences in AGENTS.md guide subsequent interface generation.
Hassan El MghariAI Engineer World's Fair 2026
In a moderated panel, Together AI's Dan Fu and MiniMax's Olive Song discuss MiniMax-M3, the partnership required to serve an open-weight multimodal model at scale, and the GPU-inference infrastructure supporting agent workloads.
Dan Fu · Olive SongAI Engineer World's Fair 2026
Max Ryabinin presents parallelism, activation offloading, and grouped attention-head processing to stretch training context lengths. The joint session featuring Dan Fu and Olive Song discusses sparse attention and KV-cache challenges when serving agent workloads.
Rishabh Bhargava explains how streaming transcription, reasoning, tool calls, and speech synthesis share a tight latency budget. Turn detection and model colocation address delays beyond the language model itself.
Hassan El Mghari's app-building talk covers prototyping, authentication, cost limits, and launch practices. His design talk adds visual references, explicit design rules, and iterative refinement to improve generated interfaces.
Affiliations reflect each recorded session, not necessarily current employment. The MiniMax model-serving session is a joint discussion featuring Dan Fu and MiniMax's Olive Song.