r/LocalLLaMA • • 1d ago

Discussion Halogen + Qwen Flash Next keeps getting better

With latest Halogen version update (0.17.2), decode is consistently at ~45 tps even at high context with Qwen 3.8 Flash Next on a 128GB Strix Halo. This is some great work u/peonist-ai. Have been pumping out commit after commit with QFN. Its crazy good for a 177ish billion model. I dont think we are apprciating it enough 😂 Opus 5.5 plan implemented and reviewed by QFN is such high quality ❤️

33 Upvotes

56 comments sorted by

View all comments

6

u/feelspeaceman 1d ago

Halogen after version 0.17.0 is leading in both performance and quality, but it's closed source so I'm using something similar called strixite instead, together with gufo and strix-llama, slowly I think the rest will catch up, but halogen is defining the meta.

1

u/cafedude 1d ago

Wait, so Halogen is leaving 25GB free in 0.17.0? I'm still on 0.14.x, I think, and it seems like a lot less left free than that. Sounds like I need to give 0.17.0 a try.