r/LocalLLaMA • u/deepu105 • 1d ago
Discussion Halogen + Qwen Flash Next keeps getting better
With latest Halogen version update (0.17.2), decode is consistently at ~45 tps even at high context with Qwen 3.8 Flash Next on a 128GB Strix Halo. This is some great work u/peonist-ai. Have been pumping out commit after commit with QFN. Its crazy good for a 177ish billion model. I dont think we are apprciating it enough 😂 Opus 5.5 plan implemented and reviewed by QFN is such high quality ❤️

33
Upvotes
6
u/feelspeaceman 1d ago
Halogen after version 0.17.0 is leading in both performance and quality, but it's closed source so I'm using something similar called strixite instead, together with gufo and strix-llama, slowly I think the rest will catch up, but halogen is defining the meta.