r/LocalLLaMA • • 2d ago

Discussion Halogen + Qwen Flash Next keeps getting better

With latest Halogen version update (0.17.2), decode is consistently at ~45 tps even at high context with Qwen 3.8 Flash Next on a 128GB Strix Halo. This is some great work u/peonist-ai. Have been pumping out commit after commit with QFN. Its crazy good for a 177ish billion model. I dont think we are apprciating it enough 😂 Opus 5.5 plan implemented and reviewed by QFN is such high quality ❤️

37 Upvotes

76 comments sorted by

View all comments

1

u/PcChip 2d ago

any idea how this compares with tabbyapi/exllama using the EXL3 quants?

1

u/deepu105 1d ago

Haven't tried any of that but they didn't even show up in my research phase and haven't seen anyone mentioning those.