Skip to content

Now in early access · Android

Your private AI that knows your phone.

Runs open-weight models fully on-device. Works offline. Nothing leaves your phone unless you ask it to.

Download APK — coming soon Android 9+, 64-bit

on a mid-range phone
~14 tok/s
on a mid-range phone
accounts or trackers
0
accounts or trackers
open models to choose from
20
open models to choose from

Animated preview of the PocketMind chat: an offline question, a nearby places search using OpenStreetMap, and a calculator tool call, each with live speed, CPU, memory and temperature readings.

Powered by open models via llama.cpp

  • Qwen
  • MiniCPM
  • Gemma
  • Llama
  • LFM
  • SmolLM
  • Phi
  • DeepSeek-R1
  • Granite

Features

An assistant, not a model playground.

Everything you expect from a modern AI chat, plus the one thing cloud assistants can't offer: it works with your phone's data without that data ever leaving it.

Knows your phone

"Who called me today?" "Am I free at 4?" "Pharmacy near me." "Add Rahul's new number." PocketMind uses typed tools for contacts, calendar, location, notifications, and (as your default assistant) call history and SMS.

  • Contacts
  • Calendar
  • Nearby places
  • Notifications
  • Call history
  • Messages
  • Web search
  • Calculator

Private by default

No account, no server, no analytics. What you ask stays on your phone.

  • Chats & history on device
  • Memories on device
  • Model files on device
  • Voice audio on device
  • Contacts, calendar, SMS on device

Real-time metrics & a benchmark lab

Honest performance: live tokens/s, CPU, RAM and temperature while you chat, plus a benchmark lab to find the model that suits your phone. Export results as CSV or JSON.

Works offline

On a flight, in the mountains, or on a patchy connection. Once a model is downloaded, chat, local voice input and on-phone tools like contacts and calendar work in airplane mode. No per-query cost, ever.

Voice input

Speak instead of typing. Use Google's on-device recognizer, or a fully local Moonshine or Whisper model.

4 s of audio → text in 83 ms

Thinking that doesn't overthink

Auto thinking reasons only when it helps (math, code, planning). A budget and loop detection stop small models from going in circles.

Memory you control

Ask it to remember something and it saves a short, editable note on your phone. View, edit or delete any memory at any time.

Your models, your choice

Pick from 20 open-weight models sized for your phone, from 292 MB to 3.1 GB, with a device-fit badge on each. Paste any GGUF link from Hugging Face. Or connect an OpenAI-compatible server you trust, such as Ollama on your laptop; API keys are encrypted on the device.

How it works

Up and running in a minute.

  1. Step 01

    Download a model

    One tap. PocketMind reads your phone's RAM and recommends models that will run well, with a fit badge on every option.

    from 292 MB · resumable download

  2. Step 02

    Ask anything

    Type or speak. Answers stream token by token with Markdown, code, tables and LaTeX, and your history stays on the phone.

    loads in 1.5–3 s · works offline

  3. Step 03

    Let it use your phone, with permission

    Switch on only the sources you want: contacts, calendar, location, notifications. Every change it proposes waits for your OK.

    opt-in per source · OK / Cancel

Privacy

Private by architecture, not by promise.

The model runs on your phone's CPU, so your questions, chats and personal data are processed in your pocket. Data only crosses the line for things that need the internet, and you control each one.

  • No account, no server

    There is nothing to sign up for and no PocketMind backend that sees your chats.

  • No ads, analytics or trackers

    No tracking SDKs in the app. Your data is never the business model.

  • Only the minimum, under your control

    Web search sends just the query and can be switched off. Nearby places sends an area rounded to about 100 m.

  • Delete anything

    Clear chats, memories and models in the app. Uninstalling removes everything.

Read the privacy policy

on-device boundary

  • Your prompts
  • Contacts · Calendar
  • SMS · Calls
  • Voice audio

llama.cpp
on your CPU

  • Answers
  • Chat history
  • Memories
  • Model files

Web search

Query only → DuckDuckGo, when online. Switch it off any time.

Remote model

Opt-in. Only to a server you add, labelled in chat.

Nearby places (OpenStreetMap) and model downloads (Hugging Face) are also user-initiated.

Models & performance

Real numbers from a real mid-range phone.

No cherry-picked flagship. Every figure below was measured in PocketMind's benchmark lab on an iQOO Z10 (Snapdragon 7s Gen 3, 8 GB RAM).

Choose a metric

Decode speed, higher is better Prompt processing (pp512), higher is better Time to first token, lower is better

  • MiniCPM5 1B Q4_K_M 14.1 115 552
  • LFM2.5 1.2B Instruct Q4_K_M 14 113 580
  • Qwen3.5 0.8B Q8_0 10.5 131 422
  • MiniCPM5 2B Q4_K_M 6.7 45 1128
  • Qwen3.5 2B Q4_K_M 3.2 n/a n/a

4 threads, Q4_K_M / Q8_0 GGUF, measured 4 Oct 2026. *Qwen3.5 2B measured in chat with a 692-token prompt. Decode speed is limited by memory bandwidth, so smaller files are faster.

1.5–3 s

to load a model

83 ms

to transcribe 4 s of speech (Moonshine)

Need more power?

Point PocketMind at your own GPU. Qwen3.5 4B via Ollama on a GTX 1650 laptop streams at ~38–44 tok/s over Wi-Fi.

20 models in the catalog 292 MB – 3.1 GB

  • Qwen3 / Qwen3.5 0.6B – 4B
  • MiniCPM5 1B, 2B
  • Gemma 3 / Gemma 4 270M – E2B
  • LFM2.5 350M – 2.6B
  • Llama 3.2 1B, 3B
  • SmolLM3 3B
  • Phi-4 mini 3.8B
  • DeepSeek-R1 Distill 1.5B
  • Granite 4.0 1B
Full benchmark table
Benchmark results on iQOO Z10, Snapdragon 7s Gen 3, 8 GB RAM
Model Quant File Generation tok/s Prompt tok/s First token Peak RAM
MiniCPM5 1B Q4_K_M 688 MB 13.8–14.4 112–118 552 ms 1.40 GB
LFM2.5 1.2B Instruct Q4_K_M 731 MB 13.5–14.4 101–125 541–622 ms 1.72 GB
Qwen3.5 0.8B Q8_0 812 MB 10.5 131 422 ms —
MiniCPM5 2B Q4_K_M 1.6 GB 6.4–7.0 39–50 1128 ms —
Qwen3.5 2B Q4_K_M 1.3 GB 3.2* — — 1.2 GB

The app

Built for everyday use.

Real screenshots from PocketMind running on an iQOO Z10.

FAQ

Questions, answered.

Something else? Email shivanggupta9696@gmail.com.

Is it really offline?
Yes. After you download a model, chat runs entirely on your phone's processor, and so do local voice input and on-phone tools like contacts and calendar. Try it in airplane mode. Only the features that need the internet (web search, nearby places, remote models and downloading models) need a connection.
Which phones does it support?
Android 9 or later on a 64-bit (arm64) phone. We recommend 8 GB of RAM; 6 GB works well with small models (about 1.5 GB or less). The app shows a fit badge on every model so you know what will run well. Our reference device is an iQOO Z10 (Snapdragon 7s Gen 3, 8 GB).
What data does it access?
Only what you switch on in Settings → Permissions: contacts, calendar, location, notifications, and (only if you make PocketMind your default assistant) call history and SMS. Data is read on your phone only when you ask a question that needs it, and it is never uploaded. Any edit, delete or new event needs your OK first. See the privacy policy for details.
Is it free?
Yes, early access is free, and there is no account or subscription to sign up for. We may add optional paid extras later, but there will be no ads and your data will never be the business model.
Why isn't it on the Play Store yet?
PocketMind is in early access while we test it on more phones and finish Google Play's review requirements for assistant apps (for example the SMS and call-log declarations). Want to try it now? Request early access and we'll send you the APK.
Can I use my own server or models?
Yes. Paste any GGUF model link from Hugging Face to run it on-device, or add any OpenAI-compatible endpoint (for example Ollama or LM Studio on your computer, or a hosted API). API keys are encrypted with the Android Keystore, and chats that use a remote model are clearly labelled because those messages go to that server.

An AI that lives on your phone, not in someone's cloud.

PocketMind is in early access on Android. Tell us your phone model and we'll send you a build.

Get early access Download APK — coming soon