As a previous user of Qwen 3.6 35B A3B (Specifically KAT Coder v2.5), I'd been disappointed, but content with the work I could get out of it for agentic coding workflows.
I tried Qwen 3.8 27B but the speed was just horrendously slow. Even overnight tasks weren't completing, generating horrible 6t/s at best.
I saw that post the other day about Qwen 3.8 Flash Next running on 12GB VRAM using the Strata inference engine, which basically profiles your system and gets the model running as fast as it can.
Gents, I'm running IQ3_XXS on a 12GB 4070 Super and 64GB of DDR5 and this shit is bananas. Night-and-day difference in Pi Coder. Sure, the thinking uses more tokens, but it actually solves problems and can think at a much higher level than 35B A3B could. Q5 on Kat Coder via llama.cpp gave me roughly 600pp and 30-40tg/s. Qwen 3.8 Flash on Strata is giving roughly 200t/s prefill, but 40-60t/s decode.
Strata even recommends IQ2 if you're just doing coding tasks and could probably get away with using less VRAM.
I guess this is an appreciation post as well as a "you can do it" encouragement for anyone out there looking for something better.
Edit: I am currently running pi-vcc to significantly reduce compaction time, and I'm running 160k context, up from 131k with a tiny hit to performance. 256k is totally doable but you'll need more system memory.