← All organizations

Local AI and distributed inference

EXO Labs

EXO Labs builds exo, open-source software that connects Macs and workstations into a local AI inference cluster. Developers and organizations can pool device memory to run models too large for one machine, load models from Hugging Face, and manage their cluster through a dashboard. Automatic device discovery and model partitioning reduce manual setup, while OpenAI-, Claude- and Ollama-compatible APIs let users connect existing clients to locally hosted models.

Founded in 2024 and based in London, the company was co-founded by Alex Cheema, its CEO, and Mohamed Baioumy. Its engineering work includes separating prompt processing from token generation across different hardware. In its DGX Spark and Mac Studio implementation, EXO assigns these phases to devices with different compute and memory-bandwidth strengths, streaming the model’s KV cache between them while computation continues. This approach coordinates both memory capacity and the work performed during an inference request.

EXO Labs also produces the speed and evaluation data behind local.ai, a reference for choosing local AI setups. It compares combinations of models, hardware, agent harnesses, inference engines and configurations using task quality, completion time and cost. This helps people buying hardware or configuring existing machines assess an entire setup rather than relying on token throughput alone.

exolabs.net

2 talks

Newest first

1 speaker at AIE

Affiliations reflect their AIE appearances, not necessarily current employment.

Start here

  1. Frontier AI at Home (literally)

    Start with the workshop for a practical view of Mac-based infrastructure, cluster dashboards, and large-prompt demonstrations.

    Alex CheemaAI Engineer Europe 2026

  2. State of the Union: Why Local, Why Now

    Follow with the panel to learn how distributed inference on NVIDIA DGX systems and computer-vision applications fit into the discussion of local AI adoption.

    Nader Khalil · Alex Cheema · Matthew Berman · Ahmad Osman · Joseph NelsonAI Engineer World's Fair 2026

Messages from the stage

Hardware and inference choices

Cheema's workshop examines prefill and decode alongside energy-efficient hardware-software co-design, quantization tradeoffs, and test-time compute.

Making local AI easier to adopt

The panel featuring Cheema discusses simpler onboarding, multi-model routing, context management, and specialized models, alongside remaining open-source-access challenges.

Affiliations reflect each recorded session, not necessarily current employment. State of the Union: Why Local, Why Now is a joint panel discussion with speakers from several organizations.

Company sources · checked 2026-08-28