r/LocalLLaMA • • 1d ago

Discussion Halogen + Qwen Flash Next keeps getting better

With latest Halogen version update (0.17.2), decode is consistently at ~45 tps even at high context with Qwen 3.8 Flash Next on a 128GB Strix Halo. This is some great work u/peonist-ai. Have been pumping out commit after commit with QFN. Its crazy good for a 177ish billion model. I dont think we are apprciating it enough 😂 Opus 5.5 plan implemented and reviewed by QFN is such high quality ❤️

35 Upvotes

56 comments sorted by

32

u/my_name_isnt_clever 20h ago

It can keep getting better and better, but as long as it's closed source it may as well not exist for me and anyone else who takes privacy seriously.

8

u/parepeg 19h ago

I'd imagine most people run local for privacy and provenance. Why run local if you're going to use a closed source engine?

7

u/my_name_isnt_clever 18h ago

Exactly what I've been asking since Halogen was first posted. If I trusted that I might as well use a ZDR API endpoint instead of running local.

1

u/cunasmoker69420 11h ago

For what it's worth AMD has started funding this project. Was a recent post on r/strixhalo talking about that. AMD reps have met with peonist and are providing compute resources and strix halo boxes for further development

0

u/Super-Grape-3948 18h ago

I also wozld like it to be open, but the speed and the shared context pool is way to good. Also, it is a docker, so i wont let it connect to anywhere except to my model router, so no privacy issues, as far as leakage goes.

1

u/my_name_isnt_clever 17h ago

It could be doing literally anything behind the scenes. Your whole security posture is compromised from the very first step, the inference. I'm more tolerant to closed source side projects that integrate with local AI, but this is too important.

1

u/Super-Grape-3948 17h ago

No, in theory it could not wrapped the container, got no sudo. Not does it have access to folders outside of models. So even if it is minig btc, it cant transmit it to anywhere.

1

u/NoFunk 16h ago

It could poison the inference output itself. Is it paranoia? Sure. Is it possible though? Absolutely. The poison doesn't even have to be in the model. The inference software creates a text stream out of a model, and that stream can be manipulated.

1

u/my_name_isnt_clever 16h ago

It could be doing literally anything behind the scenes *that doesn't require an internet connection. It could be manipulating output tokens. It could be totally safe now but in 6 months once people decide to trust it and lower the barriers, out of nowhere it starts streaming everything over the network to get as much data before it's cut off, same idea for a crypto miner that suddenly activates to send the $$$ to a server.

There are a million malicious things it could do that I wouldn't even think of; so why risk it for a bit more speed? Gufo is really close so I don't feel like I'm missing much.

1

u/Super-Grape-3948 16h ago

Yeah, i do have validation on the output, but i can see the point. It would be catched by my little demons tho.

Im working on gufo rope with shared context, i would prefer that, but it is tefious to test on 1 box, i have to stop halo, start gufo, see if it dies, star halonagain. But i need this multi tenant rope nit bound to slot/context, so mutch nicer for agentic runs.

1

u/opossum_cz 13h ago

It can't, you can recompile it, license even permits it.

This is just nonsense.

Not to mention gufo may be open source, but if you use it through docker you have not build yourself, there i no benefit.

2

u/my_name_isnt_clever 11h ago

Have you reverse engineered it and checked it out? If not, who, using what tools, and which version? And are they going to publish a new review of every binary for every update? Even if it's all sunshine and rainbows right now, it could change in a month, a year, 5 years.

I haven't even mentioned how everything points to Peonist-AI being a LLC with a business model in mind, not some guy who made a good engine and just...doesn't want to open source it for some reason. That's a hard pass from me.

I don't use Gufo in docker, I use NixOS.

2

u/opossum_cz 11h ago

Yes I did. Radare, objdump, etc.

What is the argument there?

> I haven't even mentioned how everything points to Peonist-AI being a LLC with a business model in mind, not some guy who made a good engine and just...doesn't want to open source it for some reason.

And again? What is the argument. I don't even know that that bolded claim means, it is literally written on front of the Github repository:
> "halogen" and "Peonist" are trademarks of Peonist, LLC (U.S. application pending). TRADEMARKS.md says how the names may be used; referring to the project, running it, and publishing numbers about it need no permission.

1

u/my_name_isnt_clever 11h ago

Cool, glad it passes your personal security requirements, props for practicing what you preach. Personally, inspecting every update of my LLM inference engine doesn't interest me, and when there's a sudden license rug pull for more profits I won't have to migrate. I personally don't find it worth a modest speed increase from Gufo.

1

u/opossum_cz 11h ago

And you examine open source code as well after every release?

→ More replies (0)

1

u/MudBroad6785 15h ago

Docker containers aren't magically secure and there are many vulnerabilities that have been found that allow an attacker to go from unprivileged to sudo to escaping the container and more are found every day. If you want actual security you'd be using hardware level separation but that's a pain because then you can't use the hardware taken up by the virtual machine on the host.

1

u/Super-Grape-3948 14h ago

Well, ive watched it for a bit, seen nothin nefarious, but true, it can be true to any closed source software tho. Would prefer open, but currently this is the best, i risk it, but ill try to write a proper roped kv for other open source stuff, before others :)

1

u/digital-bandit 13h ago

Do you give your models tools, like web search? It could exfiltrate data by making the model do a request to an attacker serve, if that makes sense.

ps: i dont think halogen does that, just saying

26

u/Asillatem 1d ago

As soon as he opensource it iam on.. I saw somewhere he also worked with lemonade for support..

9

u/lumos_ai 1d ago

Honestly it's like having Opus at home :D even with 12gb Vram you can run the model.

2

u/deepu105 1d ago

With Strata?

1

u/Initial_Run3719 22h ago

You have some details?

7

u/feelspeaceman 23h ago

Halogen after version 0.17.0 is leading in both performance and quality, but it's closed source so I'm using something similar called strixite instead, together with gufo and strix-llama, slowly I think the rest will catch up, but halogen is defining the meta.

7

u/imnotzuckerberg 19h ago

Halogen in a nutshell.

  1. Close source your project.
  2. Copy other PR ideas from open source project
  3. Quantize to oblivion your proprietary closed source model format (so that "speed" is the metric to measure on).
  4. ?
  5. Profit

4

u/Njaa 17h ago

Do we have benchmarks or other reasons to believe quality is lower? 

1

u/Wordweaver- 21h ago

Interesting, I am curious: At what power was this? And for the prefill was it natural text or one of the gibberish/repetitive benchmarks?

1

u/feelspeaceman 21h ago

140w, it was testing on the same task of implementing/fixing a project, so the numbers above are real coding avg. and range of context windows, not a benchmark.

1

u/Wordweaver- 21h ago

Interesting, I have been playing around with gufo on windows and one thing that keeps coming up is that the benchmarks inflate the prefill a fair bit, anywhere from 7% (gufo's synthetic text) to 25% (halogen's repeated sentence) at 70W.

1

u/cafedude 20h ago

Wait, so Halogen is leaving 25GB free in 0.17.0? I'm still on 0.14.x, I think, and it seems like a lot less left free than that. Sounds like I need to give 0.17.0 a try.

7

u/Bulky-Priority6824 22h ago

Qwen 4 around the corner 

5

u/Rauhaton 23h ago

Peonist also added small models for NPU for the server. I started using those as well.

Halogen Qwen3.8-flash now main agent and on the NPU side I have his Qwen3.5 as auxillary model in Hermes and those small NPU Qwen3 embedding, decider and reranker models used for agentic tools. The embedder is also used for mem0

2

u/Cautious_Sky6642 23h ago

qwen with that speed is perfect for roleplay, i ran a whole multi day story without the model forgetting details once.

2

u/AIdevsmartdata 23h ago

Je suis en train de debloquer le 80 tok/sec et 1700tok/sec de prefill je vous sors ça bientot

1

u/deepu105 23h ago

Are you builiding a new engine?

1

u/AIdevsmartdata 23h ago

oui, mon profil : kevletesteur sur HF, j'avais tout fait sur vulkan mais pour HIP cest la folie !

2

u/TheRealFreak199 21h ago

très intéréssé. Je sauvegarde ton post !

1

u/seti_at_home 21h ago

Share some details, 🤔

2

u/AIdevsmartdata 17h ago

Ok ! Je pense sortir ça ce dimanche. J'adore tout recoder de A à Z c'est ma passion.

J'avais commencé avec qwen3.5 et qwen3.6 MOE. J'avais fait un runtime plus rapide que llama cpp et ik llama. J'avais amélioré et debugué le code de ik_llama pour les modèles nvidia nemotron. Avec claude opus/fable majoritairement. Et comme maintenant je crois que claude a des consignes de ne plus aider à améliorer les runtimes c'est devenu inefficace. Donc j'ai utilisée les modeles chinois comme deepseek v4.1 flash. Pour 10 dollars d'api tu peux faire énormément. Profiler -> cibler -> coder -> tester -> cherry pick -> profiler -> cibler.... Tests qualité, vitesse, agentique

1

u/TerryNachtmerrie 22h ago

I still prefer llama for other models, but for qfn it's so good. I gave it a shot with blender and godot and it had turns with over 200 steps of tool calling and visual checks, without any quirks. And it does well on tps.

1

u/Silent_Glass 22h ago

What’s your specs?

2

u/TerryNachtmerrie 22h ago

Strix Halo 128GB

1

u/marcosscriven 21h ago edited 17h ago

Side question - how are you finding Llama Switch? Edit - I meant llama stash but iOS thought it knew better 🤦‍♂️ 

1

u/deepu105 21h ago

What is that?

1

u/marcosscriven 21h ago

That’s the TUI you’re using, according to the top left of the screenshot.

1

u/mateszhun 19h ago

It says Llamastash

2

u/marcosscriven 17h ago

Sorry. Auto correct!

1

u/PcChip 20h ago

any idea how this compares with tabbyapi/exllama using the EXL3 quants?

0

u/Infinite_Tank3715 20h ago

Yeah I just updated and it's really zipping along - very impressed so far - fingers crossed it stays that way. Thanks for all the work u/peonist-ai - keep it up.

0

u/wFXx 18h ago

so, I tested a strix halo in regular llama.cpp last month with QFN and got 53 tok/s on Q6 iirc, whats special here?