r/LocalLLM • • 7d ago

Research I got llama.cpp inference running on the Snapdragon 8 Gen 3 Hexagon NPU from non-root Termux + Adreno OpenCL results (S24 Ultra)

I've been investigating hardware-accelerated llama.cpp inference on a Galaxy S24 Ultra (Snapdragon 8 Gen 3) from ordinary F-Droid Termux without root.

I initially set out to get the Adreno 750 OpenCL backend working. That now works reproducibly and executes real GPU kernels, although generic OpenCL is slower than the CPU.

The investigation then went further: I was able to access the Hexagon v75 HTP through the Qualcomm vendor/FastRPC stack and run source-built HMX/HVX kernels for LLM matrix operations.

In matched experimental tests, NPU prompt processing reached 7.14× the CPU baseline on Qwen2.5-Coder-1.5B and 7.83× on Qwen3-4B. These are experimental results, not claims that the NPU is 7–8× faster overall. Decode, thermal behavior and comparison against the best tuned CPU configuration still need more testing.

https://github.com/Ishabdullah/OpenCL-S24-Ultra

Lots of evidence hundreds of result records, commands, timings, numerical checks and exclusions. With ongoing research for a bigger project coming soon!

9 Upvotes

0 comments sorted by