Unlocking Developer Productivity across CPU and GPU with MAX
AI Engineer World's Fair 2024 · 18:33
AI infrastructure and inference
Modular builds software for developers and enterprises to deploy AI across different hardware architectures. MAX is its model-serving framework, with an OpenAI-compatible endpoint and customizable models and kernels. Mojo provides a systems programming language for CPU and GPU programming; MAX’s kernels are written in Mojo. Modular Cloud offers managed inference through shared endpoints with per-token pricing or dedicated deployments on Modular’s compute or customers’ own infrastructure.
Founded in 2022 by Chris Lattner and Tim Davis, who met at Google, Modular develops a common software foundation for heterogeneous compute. Its platform uses a unified low-level layer in place of vendor-specific runtimes, while Mojo lets developers extend MAX with new algorithms. The Mojo compiler and toolchain became fully open source in 2026 under Apache 2.0 with LLVM exceptions. Lattner now serves as an executive vice president at Qualcomm, and Davis is senior vice president and general manager of Modular.
Qualcomm completed its acquisition of Modular in July 2026, retaining Mojo, MAX and Modular Cloud as products and brands. In August 2026, Modular reported that enterprise customer MiniMax’s dedicated M3 deployment served billions of tokens per minute. Before the acquisition, Modular raised $250 million in September 2025, bringing total funding to $380 million at a $1.6 billion valuation.
AI Engineer World's Fair 2024 · 18:33
Affiliations reflect their AIE appearances, not necessarily current employment.
Lattner argues that fragmented inference frameworks hinder secure, customizable production AI, making framework integration a concern beyond execution speed.
Affiliations reflect each recorded session, not necessarily current employment.