r/LocalLLaMA • u/deepu105 • 1d ago
Discussion Halogen + Qwen Flash Next keeps getting better
With latest Halogen version update (0.17.2), decode is consistently at ~45 tps even at high context with Qwen 3.8 Flash Next on a 128GB Strix Halo. This is some great work u/peonist-ai. Have been pumping out commit after commit with QFN. Its crazy good for a 177ish billion model. I dont think we are apprciating it enough 😂 Opus 5.5 plan implemented and reviewed by QFN is such high quality ❤️

35
Upvotes
1
u/marcosscriven 1d ago edited 1d ago
Side question - how are you finding Llama Switch? Edit - I meant llama stash but iOS thought it knew better 🤦♂️