Providers & models

Switch between on-device GGUF/MLX, LM Studio, Ollama, JAN AI, oMLX, Unsloth, and cloud — plus the in-app catalog, what's free vs Pro, downloads, and reasoning.

Answer first

LM Mini talks to whichever backend you select. Chats stay on that backend — this device, a computer on Wi‑Fi (LM Studio, Ollama, JAN AI, oMLX, Unsloth), LM Mini Home, or a cloud key you own.

Provider kinds

ProviderWhere it runsBest for
On-device GGUFPhone / MacOffline, privacy, travel. Free: four starter models
On-device MLX (Pro)Apple SiliconFaster Apple-native inference
LM StudioYour PC (LAN or Home)Bigger GPU models, MCP
OllamaYour PC (LAN or Home)Simple local API
JAN AIThe JAN app on this computerLocal models from JAN, usually port 1337
oMLXA networked Apple boxShared MLX server, usually port 8000
Unsloth DesktopYour computerLocal models from Unsloth, http://IP:8888, needs an sk-unsloth-… key
Cloud (Pro)Vendor / OpenAI-compatible APIsHosted models you already pay for

Change provider in Settings → Server / Providers, or pin one per chat, Persona, or group participant.

On-device

Plain-language overview: What is on-device AI?

FreeLM Mini Pro
Catalog downloadsFour starter GGUF models (1–1.7B): Qwen 3 1.7B, Llama 3.2 1B, Gemma 3 1B, DeepSeek R1 Distill 1.5BThe full catalog, including bigger models and every MLX build
Your own filesOne imported GGUF fileUnlimited Hugging Face / GGUF / MLX imports
  1. Open Settings → Models → Browse (or finish the welcome / Get Started sheet).
  2. Pick Faster, Balanced, or Best for this device — or browse the catalog (Qwen 3, Gemma 3, Llama 3.2, Phi-4 Mini, DeepSeek distill, and more). Picks above 2B parameters open the Pro screen on the free plan.
  3. Download. iPhone shows Lock Screen / Dynamic Island progress; Android uses a progress notification.
  4. Start a new chat and select that model.

With Pro, MLX is preferred on Apple Silicon when the catalog has an MLX build. Otherwise LM Mini uses GGUF.

Samplers that apply on-device: temperature, top-p, top-k, max tokens, repeat penalty. LM Studio-only knobs (for example Min-P, some load options) stay hidden in the per-chat sheet when on-device is active.

Tip

After an Android download, open a new chat if the old thread still thinks no model is selected.

LM Studio

  1. Start the Developer server with Serve on Local Network.
  2. In LM Mini, set the server URL (http://IP:1234).
  3. Paste an API token if LM Studio requires auth.
  4. Pick a loaded model. LM Mini can also ask the server to load one when that path is available.

Settings → Models can set load options the server supports: context length, flash attention, KV cache offload.

Away from Wi‑Fi on a Mac, use LM Mini Home instead of typing a LAN IP.

Ollama

Point LM Mini at http://IP:11434 (unless you changed the port). Pick a pulled model. Connection help is Ollama-specific — you will not get LM Studio steps here.

JAN AI

JAN is a free local provider (not a cloud key). Add it under Settings → Server → Add server and choose JAN AI.

  1. Start the JAN app and turn on its local API.
  2. In LM Mini, use http://YOUR_LAN_IP:1337 (default). On the same Mac as JAN, http://localhost:1337 is fine.
  3. Leave the API key empty unless you set one in JAN.
  4. Pick a model JAN has loaded.

Do not type localhost on a phone — that is the phone, not the computer running JAN.

Unsloth

Unsloth Desktop is a free local provider too (since 1.9.10). Add it under Settings → Server → Add server and choose Unsloth.

  1. Start Unsloth Desktop on your computer and load a model.
  2. In LM Mini, use http://YOUR_LAN_IP:8888 (default). On the same Mac, http://localhost:8888 is fine.
  3. Paste the sk-unsloth-… API key from Unsloth. Requests without it are rejected.
  4. Pick the model. Settings → Models → Context Length loads it at that size and asks before reloading if it is already loaded at another size.

Cloud

Add API keys in Settings → Server / Add server (cloud types). Keys stay on the device. A chat or Persona can stay on OpenRouter / Mistral / DeepSeek / Gemini / Z.AI / Vercel AI Gateway / any OpenAI-compatible URL while the rest of the app uses local models. OpenAI direct is offered on Android; on iPhone, iPad, and Mac use OpenRouter or an OpenAI-compatible URL. Claude models are reachable through OpenRouter or Vercel AI Gateway.

Cloud types and Pro Search need LM Mini Pro. LM Studio, Ollama, oMLX, JAN AI, and Unsloth do not.

Reasoning / thinking

When a model thinks aloud (including Gemma-style thought), LM Mini shows that stream, then the answer. Toggle reasoning per chat when the model supports it. If the toggle is missing, the server did not advertise the capability — try another model or update.

Custom headers

For reverse proxies and Zero Trust:

See also USB / advanced networking notes under Connect.

Models UI tips