Running Gemma 4 On-Device: 40 Tokens/s on iPhone with MLX
AI Engineer Europe 2026 · 10:51
On-device AI applications
Locally AI lets iPhone, iPad and Mac users run language and vision models on their own devices. Users can download models including Llama, Gemma, Qwen and DeepSeek to ask questions, generate text and analyze images. On-device inference works offline after download, without a login. Locally AI also offers local voice conversations, customizable system prompts and Apple Shortcuts integration for automating tasks.
Created by Adrien Grondin, the app combines a native Apple interface with inference powered by Apple’s MLX framework. Its technical approach uses Apple Silicon and its unified memory architecture to run models locally. Siri integration and access through Control Center, the Lock Screen and the Action Button bring the assistant into everyday device workflows.
LM Studio announced its completed acquisition of Locally AI in April 2026, with Grondin joining to lead native AI experiences across devices; the purchase price was not disclosed. Locally subsequently became the LM Studio mobile app. Its LM Link integration lets iPhone and iPad users access larger models running on their computers, with end-to-end encrypted communication and chats stored locally. This adds remote access to users’ own hardware alongside on-device inference.
AI Engineer Europe 2026 · 10:51
Affiliations reflect their AIE appearances, not necessarily current employment.
Grondin explains integration through mlx-swift-lm and selecting Hugging Face MLX Community models and quantized variants. Audience questions address tool calling and model compatibility.
Affiliations reflect each recorded session, not necessarily current employment.