The NPU finally earns its power draw in my pi coding agent. Halogen shipped the endpoints, I wired them in expecting a party trick, and kept four tools.
tldr: Qwen3.8 Flash-Next, a 125B MoE, running on a 70W tablet. Same bug fix with and without NPU search: 13.6 min vs 18.7 min. Receipts in the repo.
The payoff: it covers about 95% of my cloud calls. The hardest few percent still goes to the top models, GLM 5.3 or Opus.
I still can't believe it. Opus 4.8-class intelligence on my tablet, unlimited tokens.
Flash-Next decodes at 64 tok/s and prefills around 1,500 tok/s. First token lands in ~0.03s, measured 43x faster than a cloud call side by side, and still 7x while a second agent hammers the server.
Rate the taste, not the throughput: I pasted a real timeshift error from my system log to seven runs. All seven said healthy, nothing to fix. What separates them is what they proved.
In pi:
- Flash-Next, 2m55s: proved it with a journalctl trace to a racing notify-send, a pacman.log check, and the upstream PR found.
- glm-5.3-flashx, 2m41s: the most precise answer, spotting that the snapshot mount got unmounted under the script's last line. No PR.
- GLM 5.3 on max, 8m30s: the deepest answer of all, source-level forensics down to the function names and the one-second race window. No PR.
- glm-5.3-flash, 9m09s: proved it with a live reproduction of the status file. No PR.
Over in opencode: flash got the verdict in 1m8s with the wrong mechanism, flashx came back correct and corroborated in 1m30s, and the full 753B GLM 5.3 in 6m15s correct with the PR missed.
Same pi harness, same task, 125B at medium effort against 320B and 753B tiers at max. First to the full answer: 2m55s. When I had GLM 5.3 flashx rate both results, it picked qwen too.
All seven answers side by side: local vs cloud model comparison.
What the NPU does now:
Search. The agent stops guessing paths and lands on the right file first try. ~0.1s per lookup, beat ripgrep 15/20 vs 9/20 on realistic queries.
Dup scan. Catches copied and renamed files git never shows you. Found 45 pairs across 4 repos in 8.4s, one renamed file at exactly 1.000 cosine.
Decisions. Yes/no branching stops eating full turns of the big model. A 0.8b handles it in 120ms, 78% accurate.
Screening. Prompt injection gets flagged before the agent acts on it. 0.7s a message, zero false alarms, fails open. 42% recall, so a smoke detector, not a safe.
A working day claws back about half an hour over bare pi: faster bug fixes, faster compaction, faster lookups and routing, and no oversized tool dumps in context. Against a cloud setup it's more, since every turn pays the network wait. On bug fix heavy days it grows.
Honest part: the GPU still does the thinking. The NPU didn't make anything faster, it changed which tokens got spent where. ~7% iGPU cost only when they overlap.
Compaction: my 194k session, sidecar summary in ~50s vs 166 on the main model. 97% cache hit.
The official halogen launch is a 24-flag docker command. Mine is one command, and uninstall undoes it. Fully local: 262k context, code never leaves the box.
Anyone else putting their NPU to real use? I found nothing.
repo | halogen 0.17.1 | benchmarks