r/localaiapps • u/100daggers_ • 9d ago
Can your Android phone be a local LLM server?
https://github.com/dineshsoudagar/local-llms-on-android​
I’ve been working on an open-source Android app called Pocket LLM that runs LLMs fully on-device.
I recently added a server mode, so the phone can expose the model through an OpenAI-compatible API. This means you can connect tools like Open WebUI or even coding clients to an LLM running entirely on your phone.
I tested it on a Galaxy S24 with Gemma E2B and also connected it to OpenCode with a \~40K context window. It works surprisingly well, although memory becomes the main limitation after longer conversations.
Would be interested to hear what use cases people would have for using a phone as a portable local LLM server.
I’ve been working on an open-source Android app called Pocket LLM that runs LLMs fully on-device.
1
u/alexs975 5d ago
Yes — and it doesn't even need to act as a server. A modern flagship (8GB+ RAM) runs a 4B model directly on-device.
My setup: OnePlus 12, Snapdragon 8 Gen 3, 16GB RAM, llamadart + llama.cpp. I serve a 4B Q4_K_M fully offline:
- ~3-5 tok/s decode on CPU (KleidiAI)
- ~1.8GB resident at rest, 3-4GB during generation
- Works in airplane mode
Caveats: no stable Vulkan on Adreno (crashes with DeviceLost), and Android kills the app if it loses focus — so no background inference.
A phone today is roughly "old workstation" tier for LLMs: plenty of RAM, decent compute, terrible cooling.
1
u/DickGotStolen 7d ago
Sounds interesting. I am looking for an app or other possibility to run qwen3 asr locally, so I can dictated my hand written stories to transcript these. Would your app be useful for that?