Model guide
Which model should I run?
A plain-language list of 51 local models. Read the first sentence if you are new; the grey line is for specs. Download smaller ones in LM Mini, or load bigger ones on a home PC with LM Studio / Ollama and connect remotely. Not sure about RAM? Find your phone.
In LM Mini · Free models download on the free plan (GGUF). In LM Mini · Pro models — and every MLX build — need LM Mini Pro. Anything can also run in LM Studio or Ollama on your computer for free.
Everyday phones
Small models that fit most iPhones and Androids. Start here if you are new to local AI.
-
Qwen 3 0.6B
In LM Mini · ProTiny and quick — good for short replies when your phone is low on memory.
Qwen3-0.6B; Q4 GGUF & 4-bit MLX. Hugging Face downloads in the tens of millions.
-
Qwen 3.5 0.8B
LM Studio / Ollama2026 Qwen tiny — hybrid long-context brain in a phone-sized file. Best when you want the newest Qwen on 4–6 GB devices.
Qwen3.5-0.8B (Feb 2026). Gated DeltaNet hybrid; 262K native context. Community GGUF/MLX as engines catch up — import in LM Mini or host via Connect.
-
Gemma 3 1B Instruct
In LM Mini · FreeGoogle’s smallest Gemma 3 chat model. Great starter for everyday questions.
google/gemma-3-1b-it. Instruct-tuned; GGUF + MLX.
-
Gemma 4 E2B Instruct
LM Studio / OllamaGoogle’s 2026 edge Gemma — ~2.3B effective, text + image + audio. The new tiny phone default if your engine supports Gemma 4.
google/gemma-4-E2B-it (Apr 2026). Official QAT GGUF. Apache 2.0. ~5.1B with embeddings; plan ~2–3 GB live.
-
Llama 3.2 1B Instruct
In LM Mini · FreeMeta’s pocket Llama — fast, tool-friendly, and solid for light chat.
Llama-3.2-1B-Instruct; tool calling. GGUF + MLX.
-
SmolLM2 1.7B Instruct
LM Studio / OllamaHugging Face’s small all-rounder. Surprisingly capable for its size.
SmolLM2-1.7B-Instruct Q4_K_M via LM Studio / custom HF import.
-
DeepSeek R1 Distill 1.5B
In LM Mini · FreeA mini “thinking” model — stronger at math and step-by-step reasoning.
R1 distill of Qwen 2.5 1.5B; reasoning traces. GGUF + MLX.
-
Qwen 3 1.7B
In LM Mini · FreeBest everyday phone pick for most people — chat, tools, and general help.
Qwen3-1.7B; primary free-slot family in LM Mini. GGUF + MLX.
-
Qwen 3.5 2B
LM Studio / OllamaNewest small Qwen that still fits everyday phones — sharper than 1.7B-class, still downloadable on cellular.
Qwen3.5-2B. Hybrid attention + native multimodal in the 3.5 family. Prefer Q4 GGUF once your llama.cpp/MLX build lists Qwen3.5.
-
Qwen 2.5 1.5B Instruct
LM Studio / OllamaPrevious-gen Qwen that’s still a reliable tiny assistant.
Qwen2.5-1.5B-Instruct; still huge Hugging Face download volume.
-
Gemma 3n E2B
LM Studio / OllamaGoogle’s on-device Gemma 3n (effective 2B). Built for phones, including audio-aware builds.
google/gemma-3n-E2B-it. LiteRT / GGUF community quants. Newer than Gemma 3 1B.
High-RAM phones & tablets
More capable answers — needs roughly 6 GB+ free RAM on the device.
-
Llama 3.2 3B Instruct
In LM Mini · ProNoticeably smarter than 1B — good for writing help and longer chats on Pro phones.
Llama-3.2-3B-Instruct; tool calling. GGUF + MLX.
-
SmolLM3 3B
LM Studio / OllamaHugging Face’s 2025 small model — a step up from SmolLM2, still phone-sized.
HuggingFaceTB/SmolLM3-3B. GGUF via bartowski / community.
-
Gemma 3 4B Instruct
In LM Mini · ProGoogle’s mid-size Gemma 3 — sharper answers; some builds can look at images.
google/gemma-3-4b-it; multimodal on supported builds. GGUF + MLX.
-
Qwen 3.5 4B
LM Studio / Ollama2026’s 4B daily driver — long context and stronger tools than Qwen 3 4B, if your phone has ~6 GB free.
Qwen3.5-4B; millions of HF downloads. Hybrid Gated DeltaNet. Import a Q4 when the engine supports it; Qwen 3 4B Instruct 2507 remains the in-app pick today.
-
Gemma 4 E4B Instruct
LM Studio / OllamaGoogle’s on-device Gemma 4 (~4.5B effective) with image and audio. The 2026 upgrade from Gemma 3 4B / 3n E4B.
google/gemma-4-E4B-it. Official QAT GGUF. ~8B with embeddings; budget ~4 GB live. Apache 2.0.
-
Phi-4 Mini Instruct
In LM Mini · ProStrong reasoning in a phone-friendly package — great for explanations and code-ish help.
microsoft/Phi-4-mini-instruct (~3.8B). GGUF + MLX.
-
Qwen 3 4B Instruct 2507
In LM Mini · ProBalanced daily driver for Pro phones and Macs — quality without a desktop GPU.
July 2025 Instruct refresh; millions of HF downloads. GGUF + MLX.
-
Qwen 2.5 3B Instruct
LM Studio / OllamaCompact Qwen for chat and light coding when 4B feels heavy.
Qwen2.5-3B-Instruct Q4_K_M; LM Studio / Ollama.
-
Gemma 3n E4B
LM Studio / OllamaGemma 3n effective-4B — Google’s newer on-device stack, including multimodal 3n builds.
google/gemma-3n-E4B-it. Community GGUF (bartowski). Heavier than E2B.
-
Nemotron 3 Nano 4B
LM Studio / OllamaNVIDIA’s 2026 tiny Nemotron — dense 4B-class, aimed at local agents.
nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16. Use a GGUF quant on phone; BF16 is a desktop download.
Mac (M-series) & strong phones
8B-class models. Comfortable on 16 GB+ Macs; some high-RAM phones can run them too.
-
Mistral 7B Instruct
LM Studio / OllamaClassic open model — clear writing and solid general knowledge on Mac or GPU.
Mistral-7B-Instruct v0.3; ubiquitous GGUF quants.
-
Llama 3.1 8B Instruct
LM Studio / OllamaMeta’s workhorse 8B — excellent general assistant on Mac or a home GPU.
Llama-3.1-8B-Instruct; tool calling; LM Studio / Ollama staple.
-
Qwen 3 8B
In LM Mini · ProFrontier feel on-device for high-RAM phones and M-series Macs.
Qwen3-8B; tool calling. GGUF + MLX in LM Mini.
-
Qwen 3.5 9B
LM Studio / OllamaFrontier feel on a 16 GB phone or any M-series Mac — the 2026 8B-class upgrade.
Qwen3.5-9B (most-downloaded 3.5 size on HF). Not in the LM Mini in-app list yet; LM Studio / Ollama / custom import + Connect.
-
Qwen 2.5 7B Instruct
LM Studio / OllamaPopular all-rounder for coding help, summaries, and chat on desktop or Mac.
Qwen2.5-7B-Instruct — still one of the most downloaded instruct models on HF.
-
DeepSeek R1 Distill 8B
LM Studio / OllamaReasoning-focused — walks through hard problems before answering.
R1 distill on Llama 8B; longer CoT traces.
-
OLMo 3 7B Instruct
LM Studio / OllamaAllen AI’s fully open 7B instruct model — a transparent alternative to Llama/Qwen.
allenai/Olmo-3-7B-Instruct (2025). Community GGUF.
-
Nemotron Nano 9B v2
LM Studio / OllamaNVIDIA 9B local model — stronger than 8B class if you have the RAM.
nvidia/NVIDIA-Nemotron-Nano-9B-v2. Prefer Q4 GGUF on Mac/GPU.
-
Gemma 2 9B Instruct
LM Studio / OllamaGoogle’s 9B chat model — polished answers when you have Mac RAM or a GPU.
Gemma-2-9B-IT; still excellent instruction following.
-
Hermes 3 8B
LM Studio / OllamaCommunity favorite for creative writing and agent-style tool use.
NousResearch Hermes-3-Llama-3.1-8B; function calling.
High-memory Mac
12B–14B class. Best on 24–32 GB+ unified memory, or remote from a GPU PC.
-
Qwen 3 30B-A3B Instruct
LM Studio / OllamaMixture-of-experts: 30B total, ~3B active. Big-model taste on a strong Mac or GPU.
Qwen3-30B-A3B-Instruct-2507 MoE. Q4 is lighter than dense 30B; still not a phone model.
-
Qwen 3.5 27B
LM Studio / OllamaDense 2026 Qwen for a high-RAM Mac or a 24 GB GPU — then use the phone as the remote keyboard.
Qwen3.5-27B. Q4 ~16 GB. Also Qwen3.6-27B as a later dense refresh — same RAM class.
-
Gemma 3 27B Instruct
In LM Mini · ProLarge Gemma 3 with vision — Mac-first in LM Mini (4-bit MLX), or a home GPU via Connect.
mlx-community/gemma-3-27b-it-4bit in LM Mini (~15.5 GB). Needs ~24 GB unified memory.
-
Gemma 3 12B Instruct
In LM Mini · ProMac-first Gemma 3 — premium quality when you have unified memory to spare.
google/gemma-3-12b-it. MLX 4-bit in LM Mini on Mac.
-
Gemma 4 12B Instruct
LM Studio / Ollama2026 Gemma 4 12B — text, image, and audio. Mac or GPU; too big for phones.
google/gemma-4-12B-it. Official QAT GGUF. 256K context. Host in LM Studio then chat from LM Mini.
-
Qwen 3 14B
In LM Mini · ProDesktop-quality Qwen 3 in LM Mini on Mac — 4-bit MLX on 16 GB+ unified memory.
mlx-community/Qwen3-14B-4bit in LM Mini (~8.2 GB). GGUF Q4 similar. Phone: Connect only.
-
Qwen 2.5 14B Instruct
LM Studio / OllamaSerious desktop quality — best on a home GPU or a high-RAM Mac.
Qwen2.5-14B-Instruct Q4/Q5; LM Studio / Ollama.
-
DeepSeek R1 Distill 14B
LM Studio / OllamaHeavy reasoning model for research-style answers on GPU or big Macs.
R1 14B distill; long context CoT; prefer desktop.
Home GPU / remote via Connect
Desktop-class models. Run in LM Studio or Ollama on a PC, then chat from your phone with LM Mini Connect.
-
Mistral Small 3.2 24B
LM Studio / Ollama2025 Mistral Small — rich writing and analysis. GPU or 48 GB+ Mac.
Mistral-Small-3.2-24B-Instruct-2506. Q4 ~13 GB.
-
Qwen 2.5 32B Instruct
LM Studio / OllamaNear-frontier open weights for a well-equipped GPU PC.
Qwen2.5-32B-Instruct; Q4 ~18 GB; LM Studio remote.
-
Qwen 3 32B
In LM Mini · ProTop in-app Qwen 3 on 32 GB+ Macs (4-bit MLX). Phones should Connect, not download.
mlx-community/Qwen3-32B-4bit in LM Mini (~18 GB). minRam 32 GB.
-
Llama 4 Scout
LM Studio / OllamaMeta’s Llama 4 Scout MoE (17B×16E). Home GPU / big Mac — then chat from the phone via Connect.
Llama-4-Scout-17B-16E-Instruct. Mixture-of-experts; Q4 still tens of GB.
-
Llama 4 Maverick
LM Studio / OllamaMeta’s larger Llama 4 MoE (17B×128E). Multi-GPU / 80 GB class — Connect from the phone; do not download to the handset.
Llama-4-Maverick-17B-128E-Instruct. Far heavier than Scout. FP8 and GGUF community quants.
-
Gemma 4 26B-A4B Instruct
LM Studio / OllamaMoE Gemma 4: 26B total, ~4B active. Big-model taste at 8B-class compute — GPU or 32 GB+ Mac, then Connect.
google/gemma-4-26B-A4B-it. Official QAT GGUF. Efficiency king of the Gemma 4 stack.
-
Gemma 4 31B Instruct
LM Studio / OllamaDense Gemma 4 flagship — reasoning, vision, coding. Home GPU / 48 GB Mac, chat from the phone via Connect.
google/gemma-4-31B-it. Official QAT GGUF. 256K context. Apache 2.0.
-
Qwen 3.5 35B-A3B
LM Studio / Ollama2026 Qwen MoE: 35B total, ~3B active. Strong Mac/GPU pick — not a phone download.
Qwen3.5-35B-A3B. Lighter than dense 27B at similar quality. Qwen3.6-35B-A3B is the later FP8/MoE refresh in the same RAM band.
-
Llama 3.3 70B Instruct
LM Studio / OllamaTop-tier Meta 70B — run at home on a strong GPU, chat from your phone via Connect.
Llama-3.3-70B-Instruct; Q4 ~40 GB; multi-GPU friendly.
-
Qwen 2.5 72B Instruct
LM Studio / OllamaHuge Qwen for maximum quality on a serious home or office GPU rig.
Qwen2.5-72B-Instruct; Q4 ~40+ GB; Ollama / LM Studio.
-
DeepSeek R1 Distill 70B
LM Studio / OllamaResearch-grade reasoning — pair with Connect so your phone stays light.
70B R1 distill; expects large VRAM or multi-GPU.
-
Mixtral 8x7B Instruct
LM Studio / OllamaMixture-of-experts — high quality without always loading a dense 70B.
Sparse MoE; ~12–26 GB depending on quant; LM Studio classic.