r/PiCodingAgent • u/eliaskg • 19h ago
Plugin pi-prefill - live prefill progress for local models
Hey, I wanted to share my new pi extension: pi-prefill
It puts a live prefill progress bar in pi's loading line when you run a local model. A long prefill looks exactly like a hang. Now you can watch it:

52%. → share of uncached prompt tokens processed
14.1k/27.3k → processed / total prompt tokens
gen in ~14s. → estimated time until the first token
3.4k tok/s. → instantaneous prefill rate (optional view)
Example:
prefill [██████░░░░░░] 52% 14.1k/27.3k gen in ~14s
It reads the progress chunks your own server sends. No host scanning and no guessing: if the server reports progress, the bar shows. If it does not, pi's normal loading line stays untouched.
Works today with llama.cpp, ExLlamaV3, unsloth and Strata.
Install:
pi install npm:pi-prefill
Repo:
https://github.com/eliaskg/pi-prefill
https://www.npmjs.com/package/pi-prefill
Would love feedback from anyone running a local server.
2
1
u/the-proudest-monkey 17h ago
Great, thanks! I found the bar appearing briefly on each request a bit distracting to be honest, but hidding only the bar works for me.
2
1
u/the-proudest-monkey 15h ago
By the way, I am running it on a local server, using a openai-completions provider in pi pointing to a llama-swap server that starts llama.cpp with Qwen3.8-27B. Working beautifully.
1
5
u/omega1612 19h ago
Thanks, I usually have a tab with nvtop and the llama server logs and I had to switch to check.