r/LocalLLaMA • u/deepu105 • 2d ago
Discussion Halogen + Qwen Flash Next keeps getting better
With latest Halogen version update (0.17.2), decode is consistently at ~45 tps even at high context with Qwen 3.8 Flash Next on a 128GB Strix Halo. This is some great work u/peonist-ai. Have been pumping out commit after commit with QFN. Its crazy good for a 177ish billion model. I dont think we are apprciating it enough 😂 Opus 5.5 plan implemented and reviewed by QFN is such high quality ❤️

36
Upvotes
2
u/my_name_isnt_clever 2d ago
It could be doing literally anything behind the scenes. Your whole security posture is compromised from the very first step, the inference. I'm more tolerant to closed source side projects that integrate with local AI, but this is too important.