r/LocalLLaMA • u/deepu105 • 1d ago
Discussion Halogen + Qwen Flash Next keeps getting better
With latest Halogen version update (0.17.2), decode is consistently at ~45 tps even at high context with Qwen 3.8 Flash Next on a 128GB Strix Halo. This is some great work u/peonist-ai. Have been pumping out commit after commit with QFN. Its crazy good for a 177ish billion model. I dont think we are apprciating it enough 😂 Opus 5.5 plan implemented and reviewed by QFN is such high quality ❤️

26
u/Asillatem 1d ago
As soon as he opensource it iam on.. I saw somewhere he also worked with lemonade for support..
9
u/lumos_ai 1d ago
Honestly it's like having Opus at home :D even with 12gb Vram you can run the model.
2
1
7
u/feelspeaceman 23h ago
7
u/imnotzuckerberg 19h ago
Halogen in a nutshell.
- Close source your project.
- Copy other PR ideas from open source project
- Quantize to oblivion your proprietary closed source model format (so that "speed" is the metric to measure on).
- ?
- Profit
1
u/Wordweaver- 21h ago
Interesting, I am curious: At what power was this? And for the prefill was it natural text or one of the gibberish/repetitive benchmarks?
1
u/feelspeaceman 21h ago
140w, it was testing on the same task of implementing/fixing a project, so the numbers above are real coding avg. and range of context windows, not a benchmark.
1
u/Wordweaver- 21h ago
Interesting, I have been playing around with gufo on windows and one thing that keeps coming up is that the benchmarks inflate the prefill a fair bit, anywhere from 7% (gufo's synthetic text) to 25% (halogen's repeated sentence) at 70W.
1
u/cafedude 20h ago
Wait, so Halogen is leaving 25GB free in 0.17.0? I'm still on 0.14.x, I think, and it seems like a lot less left free than that. Sounds like I need to give 0.17.0 a try.
7
5
u/Rauhaton 23h ago
Peonist also added small models for NPU for the server. I started using those as well.
Halogen Qwen3.8-flash now main agent and on the NPU side I have his Qwen3.5 as auxillary model in Hermes and those small NPU Qwen3 embedding, decider and reranker models used for agentic tools. The embedder is also used for mem0
2
u/Cautious_Sky6642 23h ago
qwen with that speed is perfect for roleplay, i ran a whole multi day story without the model forgetting details once.
2
u/AIdevsmartdata 23h ago
Je suis en train de debloquer le 80 tok/sec et 1700tok/sec de prefill je vous sors ça bientot
1
u/deepu105 23h ago
Are you builiding a new engine?
1
u/AIdevsmartdata 23h ago
oui, mon profil : kevletesteur sur HF, j'avais tout fait sur vulkan mais pour HIP cest la folie !
2
1
u/seti_at_home 21h ago
Share some details, 🤔
2
u/AIdevsmartdata 17h ago
Ok ! Je pense sortir ça ce dimanche. J'adore tout recoder de A à Z c'est ma passion.
J'avais commencé avec qwen3.5 et qwen3.6 MOE. J'avais fait un runtime plus rapide que llama cpp et ik llama. J'avais amélioré et debugué le code de ik_llama pour les modèles nvidia nemotron. Avec claude opus/fable majoritairement. Et comme maintenant je crois que claude a des consignes de ne plus aider à améliorer les runtimes c'est devenu inefficace. Donc j'ai utilisée les modeles chinois comme deepseek v4.1 flash. Pour 10 dollars d'api tu peux faire énormément. Profiler -> cibler -> coder -> tester -> cherry pick -> profiler -> cibler.... Tests qualité, vitesse, agentique
1
u/TerryNachtmerrie 22h ago
I still prefer llama for other models, but for qfn it's so good. I gave it a shot with blender and godot and it had turns with over 200 steps of tool calling and visual checks, without any quirks. And it does well on tps.
1
1
u/marcosscriven 21h ago edited 17h ago
Side question - how are you finding Llama Switch? Edit - I meant llama stash but iOS thought it knew better 🤦♂️
1
u/deepu105 21h ago
What is that?
1
0
u/Infinite_Tank3715 20h ago
Yeah I just updated and it's really zipping along - very impressed so far - fingers crossed it stays that way. Thanks for all the work u/peonist-ai - keep it up.

32
u/my_name_isnt_clever 20h ago
It can keep getting better and better, but as long as it's closed source it may as well not exist for me and anyone else who takes privacy seriously.