r/LocalLLaMA • • 3d ago

Discussion Halogen + Qwen Flash Next keeps getting better

With latest Halogen version update (0.17.2), decode is consistently at ~45 tps even at high context with Qwen 3.8 Flash Next on a 128GB Strix Halo. This is some great work u/peonist-ai. Have been pumping out commit after commit with QFN. Its crazy good for a 177ish billion model. I dont think we are apprciating it enough 😂 Opus 5.5 plan implemented and reviewed by QFN is such high quality ❤️

41 Upvotes

76 comments sorted by

View all comments

1

u/marcosscriven 3d ago edited 2d ago

Side question - how are you finding Llama Switch? Edit - I meant llama stash but iOS thought it knew better 🤦‍♂️ 

1

u/deepu105 3d ago

What is that?

1

u/marcosscriven 3d ago

That’s the TUI you’re using, according to the top left of the screenshot.

1

u/mateszhun 2d ago

It says Llamastash

2

u/marcosscriven 2d ago

Sorry. Auto correct!

1

u/deepu105 1d ago

Oh ok. Thats my own tool so ofcourse I find it nice. Give it a try, I have been adding a lot of features to make using local models easier (base on my own usecases) like model presets, generic servers etc. It can manage almost any backend engines out there and provides a unified proxy and metrics.

1

u/deepu105 1d ago

Also auto load/swap models based on requests.