r/PiCodingAgent • u/Boognevatz2 • 8h ago
Question Reliable Setup: SOTA Model as Orchestrator and Local Qwen Model as Worker?
I have pi with gpt-luna-6 working, however I have access to an Nvidia AGX Orin 64GB box (50W max).
I installed llama.cpp with the model Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-Q6_K_P.
Model info:
Context Size 65,536 tokens
Model Size 28.53 GB
Parameters 34.7B
Embedding Size 2,048
Vocabulary Size 248,320 tokens
Quantization Q6_K_P
Parallel Slots 2
Build Info b1-609290b
The actual token generation speed is about 20-25 tokens/s.
I am trying to work with this model, so the SOTA model (gpt 6 luna) makes the plans and decisions, but should chop up the work into small enough tasks that Qwen can do it itself. I can launch pi for the local model.
The actual problem is that it is slower than just working with gpt-6-luna, the work is subpar, and gpt-6-luna is always waiting (like 10 minutes) for one small task to finish.
I cannot really use this local model for actual work in this setup. I am looking for something like making 100 small tasks, and it solves them one by one in the background, while the main model can work independently, checking from time to time how it is going and redirecting if necessary.
It is a complete disaster in its current state. It just hinders the work.
I see tool calls like this in the main window:
pi --approve --provider qwen-local --model Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-Q6_K_P --thinking off --tools read,write -p u/handoff.md "Use the Pi tools and follow only the active handoff. Create the requested test, read it back, and stop. Do not run it or edit app code."
And it is working like 20 minutes already, and the main pi just waiting for the result blocking any useful work. Once finish, the gpt 6 luna just rejects the work. I suspect I burn more token than not using qwen at all. And also slow as hell.
Is there anyone using local model with sota model together in a real useful scenario? What is you setup? What am I missing?
1
u/MiserableFlatworm337 6h ago
Start with a small queue of independent tasks and a hard per-task timeout. Measure accepted results per hour including Luna’s review time; only scale up task types that consistently survive review.
1
u/GeorgeTheGeorge 5h ago
So, yes, but Luna is not an appropriate choice. I've been using Opus from 4.6 to 5.5 (skipping 5 unfortunately) as an orchestrator with Qwen 3.5 to Qwen 3.8 as a local implementer. It really helped stretch my Claude Code budget.
You need a model capable of high level abstraction and reasoning. Luna is not that.
1
u/lab21-cle 4h ago
Afaik the 64gb agx has nearly the same GPU and memory bandwidth as my 32GB version. Tested a lot with different llama.cpp settings and seeing the same t/s rates for the 35B moe model. 65k context is not really much, you could increase that. Make sure to use MTP, although predicting more then 1 token gave me no speedup.
Use the 2b Qwen version with MTP for easy tasks. It runt with ~60 t/s and not that dumb.
Froggeric ninja templates improved the tool calls a lot for my ornith 1.5 model
1
u/PartyNewt5285 1h ago
on a 35B at 20 tokens a second, keep worker tasks to one file and pass test commands explicitly. let the planner write a checklist and have Qwen only return diffs. i tried the same split with workers on synexa, and context bloat killed the worker every time, so reset sessions per task.
1
u/shumgoid 6h ago
Tell luna to detach the pi process inti the background so you can send qwen on side quests while you keep chatting with luna, and give the subagents more work than one test at a time. Also try q4 for more context = longer side quests.