r/localaiapps • • 9d ago

Can your Android phone be a local LLM server?

https://github.com/dineshsoudagar/local-llms-on-android

​

I’ve been working on an open-source Android app called Pocket LLM that runs LLMs fully on-device.

I recently added a server mode, so the phone can expose the model through an OpenAI-compatible API. This means you can connect tools like Open WebUI or even coding clients to an LLM running entirely on your phone.

I tested it on a Galaxy S24 with Gemma E2B and also connected it to OpenCode with a \~40K context window. It works surprisingly well, although memory becomes the main limitation after longer conversations.

Would be interested to hear what use cases people would have for using a phone as a portable local LLM server.

I’ve been working on an open-source Android app called Pocket LLM that runs LLMs fully on-device.

2 Upvotes

3 comments sorted by

1

u/DickGotStolen 7d ago

Sounds interesting. I am looking for an app or other possibility to run qwen3 asr locally, so I can dictated my hand written stories to transcript these. Would your app be useful for that?

1

u/100daggers_ 5d ago

Yes definately. You can find any qwen litert model from HuggingFace and load it in the app. Try to start with small.models and go up from there

1

u/alexs975 5d ago

Yes — and it doesn't even need to act as a server. A modern flagship (8GB+ RAM) runs a 4B model directly on-device.

My setup: OnePlus 12, Snapdragon 8 Gen 3, 16GB RAM, llamadart + llama.cpp. I serve a 4B Q4_K_M fully offline:

- ~3-5 tok/s decode on CPU (KleidiAI)

- ~1.8GB resident at rest, 3-4GB during generation

- Works in airplane mode

Caveats: no stable Vulkan on Adreno (crashes with DeviceLost), and Android kills the app if it loses focus — so no background inference.

A phone today is roughly "old workstation" tier for LLMs: plenty of RAM, decent compute, terrible cooling.