Personal LLM
Free on iOS & Android

Private AI that
never phones home.

Personal LLM runs open models — Qwen 3.5, Gemma 4, GLM, Ministral — on your phone's own chip. Download one once, then chat, ask about photos and watch it reason with no connection at all. There is no server. Your conversations have nowhere to go.

no account works offline understands photos open-weight models no subscription
The one idea

Everything interesting happens on the phone.

Cloud chatbots send every word to a data centre and hope you trust them. Personal LLM doesn't ask you to: the model is a file on your phone, the answer is computed by your phone, and the only things that ever touch the network are the two below.

Your phone
  • Your chats stay
  • Photos you ask about stay
  • The models themselves stay
  • Settings & system prompts stay
  • Usage stats stay
  • Backups you export yours

Computed here, stored here, deleted when you say. Uninstall and it's all gone — nothing was ever anywhere else.

Model file · once, from Hugging Face
Ads in the free app · only while online
Chats, photos, prompts · never
The internet
  • Hugging Face model downloads
  • Google AdMob ads
  • Our servers there aren't any

After the download, switch on airplane mode if you like. The app won't notice. Ads can't load without a connection — everything else carries on.

How it works

Download once. Talk forever.

Three screens you'll actually see, in the order you'll see them.

1

Pick a model that fits

Seven open models from 0.8 to 9 billion parameters, each with a “Fits your device” badge read off your phone's RAM. Download once; pause and resume if the Wi-Fi drops.

The Models tab: Qwen 3.5 0.8B, 4B and 9B cards with Ultra Fast, Vision and Fits-your-device badges, an Install Image Support button and a Download 6.54 GB button
2

Chat anywhere

On a plane, in a basement, in a tunnel. Answers stream in with rendered Markdown and a live tokens-per-second readout, and nothing is waiting on a signal.

A chat titled Spanish for the trip on Gemma 4 E2B: how to politely ask for the check, answered with La cuenta, por favor and two variations, at 35.4 tokens per second
3

Ask about a photo

Every catalog model can take an image. Add the image-support file once (195 MB–1 GB) and point it at a menu, a form, a whiteboard, a sign in a language you don't read.

A vision chat: a photo of a glowing purple brain attached to the question What do you see in this image, and Qwen 3.5 4B describing it as a stylized digital illustration of a human brain
The catalog

Open weights, phone sizes.

Quantized GGUF builds that run through llama.cpp — Metal on iPhone, OpenCL on Snapdragon Adreno, CPU everywhere else. Tell us your phone's RAM and we'll badge them exactly the way the app does.

How much RAM does your phone have?

Roughly: iPhone 13 → 4 GB · iPhone 14–15 → 6 GB · iPhone 15 Pro and 16 onward → 8 GB or more · recent Android flagships → 8–12 GB. The app reads the real number, so you never have to guess.

Qwen 3.5 0.8B
Ultra FastVision
0.81 GB · 0.8B params · needs 2+ GB RAM

For very low-end phones. Thinking mode on; quality is limited, speed is not.

Gemma 4 E2B
Mobile-FirstVision
2.04 GB · 2.3B effective · needs 3+ GB RAM

Google DeepMind's phone-first model, quantization-aware trained. 140+ languages.

Ministral 3 3B
Fast VisionVision
2.15 GB · 3B params · needs 3+ GB RAM

Mistral's compact model with a 256k context window — quick on documents and screenshots.

Qwen 3.5 4B
FastVision
2.74 GB · 4B params · needs 3+ GB RAM

The sweet spot: speed, quality and thinking mode on an everyday phone. Our recommended first download.

Gemma 4 E4B
Mobile-FirstVision
3.00 GB · 4.5B effective · needs 5+ GB RAM

The higher-quality Gemma, with hybrid thinking for harder questions.

Qwen 3.5 9B
RecommendedVision
5.68 GB · 9B params · needs 6+ GB RAM

The best answers in the catalog. Thinking mode for math, code and anything you'd like to check.

GLM 4.6V Flash
High-endVision
6.17 GB · 9B params · needs 8+ GB RAM

Zhipu AI's flagship-class vision model: visual Q&A, documents, UI screenshots.

Bring your own
Any GGUF+ projector
Add by URL · Hugging Face links work

Paste a link to any GGUF and it joins the catalog, with an optional vision projector for image input.

Same rule as the app: “Fits your device” means at least 2 GB of headroom over the model's minimum; “Should run” means it meets it. Image support is a separate 195 MB–1 GB download per model.

Thinking mode

Watch it reason, then fold it away.

On reasoning models, switch on thinking mode and the step-by-step working shows up in a collapsible Reasoning panel above the answer — handy for math, code and anything worth checking. It's off by default, because most questions don't need it and it's faster that way.

  • One-tap presets — Creative, Balanced, Precise, Thinking
  • Temperature, top-p, top-k, context size, max tokens — all yours to tune
  • A system prompt per chat — the recipe assistant doesn't have to sound like the code reviewer
  • Edit, regenerate, stop — a stopped reply keeps what it had
Creative t 1.2 Balanced t 0.7 Precise t 0.3 Thinking reasoning on
The bat-and-ball puzzle in chat: a collapsible Reasoning panel works through the algebra, then the answer: the ball costs 5 cents
Built for phones

A real chat app, not a tech demo.

Running a language model on a phone is the clever part. The rest of Personal LLM is the unglamorous part — the stuff that makes it the app you actually open.

GPU acceleration

Metal on iPhone, OpenCL on Snapdragon Adreno 700+, via llama.cpp. Falls back to CPU quietly when there isn't one.

Honest numbers

Tokens per second on every reply, context usage in the header, and a Fits-your-device badge before you download.

Markdown that renders

Code blocks, tables and lists come out formatted. Copy any reply in one tap.

Switch models mid-chat

Start on the fast one, hand a hard question to the 9B from the chat header. Same conversation.

Read aloud

On-device text-to-speech for any answer — no voice service, no upload.

Backup & restore

Every chat and setting as one JSON file, yours to keep. Restore from the file or the clipboard.

Downloads that survive

Pause, resume, cancel. Interrupted downloads pick up where they stopped, and storage is checked before they start.

Light and dark, search and pin

Follows your system theme or pick one. Search every chat, pin the ones you come back to.

The Chats tab, subtitled Private, on-device AI: a search field, a pinned Weeknight dinners chat and four more, each labelled with the model that answered
Questions

The short answers.

Is it free?

Yes. There's no subscription and no paywall on any feature — the models are free open weights and the app costs nothing. It's supported by ads: a small banner in chat and an occasional full-screen ad when you open a new chat or while a model downloads, never in the middle of an answer. Offline, none load.

Does anything leave my phone?

Your chats, the photos you attach and the models themselves never do — inference runs on the phone's own chip and we don't operate a server. Two things use the network: downloading a model file (once, straight from Hugging Face) and ads. Uninstall the app and everything is gone, because nothing was ever anywhere else.

Which phone do I need?

iOS 15.1 or later, or Android 7 or later. Small models run in 3 GB of RAM; the 9B models want 6–8 GB. GPU acceleration uses Metal on iPhone and OpenCL on Snapdragon Adreno 700+ phones, with a CPU fallback everywhere else. Every model card shows a “Fits your device” badge read off your actual RAM before you download.

How much storage do the models take?

From 0.81 GB (Qwen 3.5 0.8B) to 6.17 GB (GLM 4.6V Flash), plus 195 MB–1 GB if you add image support to a model. Downloads can be paused and resumed, and the app checks free space before it starts.

How fast is it?

That depends on the phone and the model — a 4B model on a recent phone streams at reading pace; a 9B is slower and smarter. Every reply shows a live tokens-per-second readout so you can compare models on your own hardware, and you can switch models mid-conversation from the chat header.

Can I bring my own model?

Yes. Add any GGUF by URL — Hugging Face links work — with an optional vision projector for image input. Temperature, top-p, top-k, context size, max tokens, thinking mode and per-chat system prompts are all adjustable, or use the Creative, Balanced, Precise and Thinking presets.

What about ads and tracking?

Ads come from Google AdMob. On iOS the app asks for App Tracking Transparency permission; decline it and ads are simply non-personalized. AdMob may collect device information to serve ads — that's spelled out in the privacy policy. The app itself has no account, no cloud sync and no analytics on your conversations.

The smartest thing on your phone can also be the most private.

Free. No account. Works on a plane.