iPhone / iPad · 2021 · A15

Local AI on the iPhone 13

4 GB is tight. Qwen 3 0.6B or Llama 3.2 1B only. Bigger models belong on a PC via Connect.

Advertised RAM
4 GB
Usable for a model
~1.7 GB
Compute
Metal / ANE
LM Mini path
GGUF / Metal (MLX with Pro)

Start in LM Mini

Install the app, pick one of these, chat offline. The free plan includes four small starter models (1–1.7B); models marked Pro need LM Mini Pro.

Get LM Mini and load a model that actually fits — on-device, or via Connect to a PC.

What actually fits

Q4 weights plus KV cache. “Tight” means close Chrome first.

ModelParamsQ4HereWhere
Qwen 3 0.6B 0.6B ~400–500 MB Fits In LM Mini · Pro
Qwen 3.5 0.8B 0.8B ~500–600 MB Fits Studio / Ollama
Gemma 3 1B Instruct 1B ~700–750 MB Tight In LM Mini · Free
Llama 3.2 1B Instruct 1B ~700–808 MB Tight In LM Mini · Free

Too big for this phone

Host these on LM Studio or Ollama, then chat from iPhone 13.

How we score this

iPhone 13 ships with 4 GB. After iPhone / iPad we budget ~1.7 GB for inference. A 7B Q4 is ~4.4 GB on disk and wants ~6.5 GB live — that is why it fails on most 8 GB phones. 12–16 GB Android can try 8B. Apple 8 GB phones should live in 1B–4B.