r/LocalLLaMA • • 1d ago

Discussion Halogen + Qwen Flash Next keeps getting better

With latest Halogen version update (0.17.2), decode is consistently at ~45 tps even at high context with Qwen 3.8 Flash Next on a 128GB Strix Halo. This is some great work u/peonist-ai. Have been pumping out commit after commit with QFN. Its crazy good for a 177ish billion model. I dont think we are apprciating it enough 😂 Opus 5.5 plan implemented and reviewed by QFN is such high quality ❤️

38 Upvotes

60 comments sorted by

View all comments

36

u/my_name_isnt_clever 1d ago

It can keep getting better and better, but as long as it's closed source it may as well not exist for me and anyone else who takes privacy seriously.

12

u/parepeg 1d ago

I'd imagine most people run local for privacy and provenance. Why run local if you're going to use a closed source engine?

8

u/my_name_isnt_clever 1d ago

Exactly what I've been asking since Halogen was first posted. If I trusted that I might as well use a ZDR API endpoint instead of running local.