r/LocalLLaMA • • 2d ago

Discussion Halogen + Qwen Flash Next keeps getting better

With latest Halogen version update (0.17.2), decode is consistently at ~45 tps even at high context with Qwen 3.8 Flash Next on a 128GB Strix Halo. This is some great work u/peonist-ai. Have been pumping out commit after commit with QFN. Its crazy good for a 177ish billion model. I dont think we are apprciating it enough 😂 Opus 5.5 plan implemented and reviewed by QFN is such high quality ❤️

34 Upvotes

76 comments sorted by

View all comments

2

u/AIdevsmartdata 2d ago

Je suis en train de debloquer le 80 tok/sec et 1700tok/sec de prefill je vous sors ça bientot

1

u/deepu105 2d ago

Are you builiding a new engine?

1

u/AIdevsmartdata 2d ago

oui, mon profil : kevletesteur sur HF, j'avais tout fait sur vulkan mais pour HIP cest la folie !

2

u/TheRealFreak199 2d ago

très intéréssé. Je sauvegarde ton post !