iPhone / iPad · 2025 · A19

Local AI on the iPhone Air

12 GB in a thin chassis. RAM matches the 17 Pro; thermals do not. 4B is the sweet spot. Long 8B sessions will throttle — Connect is still the 14B path.

Advertised RAM
12 GB
Usable for a model
~6.2 GB
Compute
Metal / ANE
LM Mini path
GGUF / Metal (MLX with Pro)

Start in LM Mini

Install the app, pick one of these, chat offline. The free plan includes four small starter models (1–1.7B); models marked Pro need LM Mini Pro.

Get LM Mini and load a model that actually fits — on-device, or via Connect to a PC.

What actually fits

Q4 weights plus KV cache. “Tight” means close Chrome first.

ModelParamsQ4HereWhere
Qwen 3 0.6B 0.6B ~400–500 MB Runs well In LM Mini · Pro
Gemma 3 1B Instruct 1B ~700–750 MB Runs well In LM Mini · Free
Llama 3.2 1B Instruct 1B ~700–808 MB Runs well In LM Mini · Free
DeepSeek R1 Distill 1.5B 1.5B ~1.0–1.1 GB Runs well In LM Mini · Free
Qwen 3 1.7B 1.7B ~1.1–1.2 GB Runs well In LM Mini · Free
Llama 3.2 3B Instruct 3B ~1.9–2.0 GB Runs well In LM Mini · Pro
Phi-4 Mini Instruct 3.8B ~2.3–2.4 GB Runs well In LM Mini · Pro
Qwen 3 4B Instruct 2507 4B ~2.4–2.5 GB Runs well In LM Mini · Pro
Gemma 3 4B Instruct 4B ~2.7–3.7 GB Runs well In LM Mini · Pro
Qwen 3.5 0.8B 0.8B ~500–600 MB Runs well Studio / Ollama
SmolLM2 1.7B Instruct 1.7B ~1.0 GB Runs well Studio / Ollama
Qwen 2.5 1.5B Instruct 1.5B ~1.0 GB Runs well Studio / Ollama
Qwen 3.5 2B 2B ~1.3–1.5 GB Runs well Studio / Ollama
Gemma 3n E2B 2B ~1.5 GB Runs well Studio / Ollama
Gemma 4 E2B Instruct 2.3B ~1.4–1.8 GB (Q4/QAT) Runs well Studio / Ollama
Qwen 2.5 3B Instruct 3B ~1.9 GB Runs well Studio / Ollama
SmolLM3 3B 3B ~1.9 GB (Q4) Runs well Studio / Ollama
Nemotron 3 Nano 4B 4B ~2.5 GB (Q4) Runs well Studio / Ollama
Qwen 3.5 4B 4B ~2.5–2.8 GB Runs well Studio / Ollama
Gemma 3n E4B 4B ~2.6 GB (Q4) Runs well Studio / Ollama
Gemma 4 E4B Instruct 4.5B ~2.6–3.2 GB (Q4/QAT) Runs well Studio / Ollama
Mistral 7B Instruct 7B ~4.1 GB (Q4) Tight Studio / Ollama
OLMo 3 7B Instruct 7B ~4.3 GB (Q4) Tight Studio / Ollama
Qwen 2.5 7B Instruct 7B ~4.4 GB (Q4) Tight Studio / Ollama

Too big for this phone

Host these on LM Studio or Ollama, then chat from iPhone Air.

How we score this

iPhone Air ships with 12 GB. After iPhone / iPad we budget ~6.2 GB for inference. A 7B Q4 is ~4.4 GB on disk and wants ~6.5 GB live — that is why it fails on most 8 GB phones. 12–16 GB Android can try 8B. Apple 8 GB phones should live in 1B–4B.