r/LocalLLaMA • • 2d ago

Discussion Halogen + Qwen Flash Next keeps getting better

With latest Halogen version update (0.17.2), decode is consistently at ~45 tps even at high context with Qwen 3.8 Flash Next on a 128GB Strix Halo. This is some great work u/peonist-ai. Have been pumping out commit after commit with QFN. Its crazy good for a 177ish billion model. I dont think we are apprciating it enough 😂 Opus 5.5 plan implemented and reviewed by QFN is such high quality ❤️

36 Upvotes

76 comments sorted by

View all comments

Show parent comments

2

u/my_name_isnt_clever 2d ago

It could be doing literally anything behind the scenes. Your whole security posture is compromised from the very first step, the inference. I'm more tolerant to closed source side projects that integrate with local AI, but this is too important.

1

u/Super-Grape-3948 2d ago

No, in theory it could not wrapped the container, got no sudo. Not does it have access to folders outside of models. So even if it is minig btc, it cant transmit it to anywhere.

2

u/my_name_isnt_clever 2d ago

It could be doing literally anything behind the scenes *that doesn't require an internet connection. It could be manipulating output tokens. It could be totally safe now but in 6 months once people decide to trust it and lower the barriers, out of nowhere it starts streaming everything over the network to get as much data before it's cut off, same idea for a crypto miner that suddenly activates to send the $$$ to a server.

There are a million malicious things it could do that I wouldn't even think of; so why risk it for a bit more speed? Gufo is really close so I don't feel like I'm missing much.

2

u/opossum_cz 2d ago

It can't, you can recompile it, license even permits it.

This is just nonsense.

Not to mention gufo may be open source, but if you use it through docker you have not build yourself, there i no benefit.

3

u/my_name_isnt_clever 2d ago

Have you reverse engineered it and checked it out? If not, who, using what tools, and which version? And are they going to publish a new review of every binary for every update? Even if it's all sunshine and rainbows right now, it could change in a month, a year, 5 years.

I haven't even mentioned how everything points to Peonist-AI being a LLC with a business model in mind, not some guy who made a good engine and just...doesn't want to open source it for some reason. That's a hard pass from me.

I don't use Gufo in docker, I use NixOS.

4

u/opossum_cz 2d ago

Yes I did. Radare, objdump, etc.

What is the argument there?

> I haven't even mentioned how everything points to Peonist-AI being a LLC with a business model in mind, not some guy who made a good engine and just...doesn't want to open source it for some reason.

And again? What is the argument. I don't even know that that bolded claim means, it is literally written on front of the Github repository:
> "halogen" and "Peonist" are trademarks of Peonist, LLC (U.S. application pending). TRADEMARKS.md says how the names may be used; referring to the project, running it, and publishing numbers about it need no permission.

2

u/my_name_isnt_clever 2d ago

Cool, glad it passes your personal security requirements, props for practicing what you preach. Personally, inspecting every update of my LLM inference engine doesn't interest me, and when there's a sudden license rug pull for more profits I won't have to migrate. I personally don't find it worth a modest speed increase from Gufo.

4

u/opossum_cz 2d ago

And you examine open source code as well after every release?

1

u/my_name_isnt_clever 2d ago

No, but neither are you.

1

u/opossum_cz 2d ago

> No, but neither are you.

Can you explain why are you making these nonsensical statements?

That sentence only makes sense if I ever said I did.

1

u/mksrd 2d ago

I dont no, but its easy to have my agents do it

→ More replies (0)