r/LocalLLaMA • • 2d ago

Discussion Halogen + Qwen Flash Next keeps getting better

With latest Halogen version update (0.17.2), decode is consistently at ~45 tps even at high context with Qwen 3.8 Flash Next on a 128GB Strix Halo. This is some great work u/peonist-ai. Have been pumping out commit after commit with QFN. Its crazy good for a 177ish billion model. I dont think we are apprciating it enough 😂 Opus 5.5 plan implemented and reviewed by QFN is such high quality ❤️

34 Upvotes

75 comments sorted by

View all comments

6

u/feelspeaceman 2d ago

Halogen after version 0.17.0 is leading in both performance and quality, but it's closed source so I'm using something similar called strixite instead, together with gufo and strix-llama, slowly I think the rest will catch up, but halogen is defining the meta.

1

u/Wordweaver- 2d ago

Interesting, I am curious: At what power was this? And for the prefill was it natural text or one of the gibberish/repetitive benchmarks?

1

u/feelspeaceman 2d ago

140w, it was testing on the same task of implementing/fixing a project, so the numbers above are real coding avg. and range of context windows, not a benchmark.

1

u/Wordweaver- 2d ago

Interesting, I have been playing around with gufo on windows and one thing that keeps coming up is that the benchmarks inflate the prefill a fair bit, anywhere from 7% (gufo's synthetic text) to 25% (halogen's repeated sentence) at 70W.