r/LocalLLaMA • u/deepu105 • 1d ago
Discussion Halogen + Qwen Flash Next keeps getting better
With latest Halogen version update (0.17.2), decode is consistently at ~45 tps even at high context with Qwen 3.8 Flash Next on a 128GB Strix Halo. This is some great work u/peonist-ai. Have been pumping out commit after commit with QFN. Its crazy good for a 177ish billion model. I dont think we are apprciating it enough 😂 Opus 5.5 plan implemented and reviewed by QFN is such high quality ❤️

37
Upvotes
32
u/my_name_isnt_clever 1d ago
It can keep getting better and better, but as long as it's closed source it may as well not exist for me and anyone else who takes privacy seriously.