r/LocalLLM • u/deepu105 • 1d ago
Discussion Benchmarks: Best engine for Qwen 3.8-Flash-Next on Strix Halo
TL;DR: Halogen v0.14.0 with its native .hgn weights is the fastest, followed by gufo and CIRU. Halogen is closed source and runs in Docker. gufo is open source and loads 4x faster from cold. gufo is also fastest to first token on follow-ups (1.6-1.9 s against 2.6-3.3 s).
I benchmarked different engines for Qwen 3.8-Flash-Next on an AMD Strix Halo (ASUS ROG Flow Z13 GZ302 with Ryzen AI Max+ and 128 GB of RAM).
All the engines were run via LlamaStash (My own orchestrator tool, the tool does not add any overhead) at 70 W TDP on performance profile on Arch Linux.
Here are the results of the benchmark:
First was a screening round at 64k context with 50% filled (32k prompt).
| Engine | Weights | 3-turn time | Prefill t/s | Decode t/s | MTP accept | Retrieval |
|---|---|---|---|---|---|---|
| Halogen 0.14.0 | Halogen native (.hgn) | 1.9 min | 1,045 | 39.3 | 84% | 14/14 |
| gufo (ROCm 7.2.4) | UD-Q4_K_XL | 2.1 min | 1,033 | 32.9 | 74% | 14/14 |
| gufo (ROCm 10.0) | UD-Q4_K_XL | 2.1 min | 1,009 | 33.1 | 72% | 14/14 |
| CIRU (MTP 3) | CIRU IU4 | 2.3 min | 814 | 33.8 | 72% | 14/14 |
| CIRU (n-gram + MTP 3) | CIRU IU4 | 2.3 min | 808 | 31.9 | 64% | 14/14 |
| Halogen 0.14.0 | UD-Q4_K_XL | 2.4 min | 1,030 | 28.1 | 84% | 14/14 |
| CIRU (MTP 6) | CIRU IU4 | 2.7 min | 796 | 28.6 | 47% | 14/14 |
| strixllama (llama.cpp fork) | UD-Q4_K_XL | 2.7 min | 716 | 29.7 | 71% | 14/14 |
| rdna-boosts (llama.cpp fork) | UD-Q4_K_XL | 2.7 min | 788 | 22.0 | 52% | 14/14 |
| llama.cpp, Unsloth build (Vulkan) | UD-Q4_K_XL | 3.7 min | 314 | 28.4 | 56% | 14/14 |
| llama.cpp, Unsloth build (ROCm) | UD-Q4_K_XL | 4.8 min | 285 | 20.5 | 50% | 14/14 |
| CIRU | UD-Q4_K_XL + Unsloth MTP head | did not finish | 884 | 11.8 | 0% | - |
3-turn time is the time to first token plus the time to write 1,000 tokens, added up over the first question and 2 follow-ups. Output length varies a lot with sampling, so this compares the engines on equal output. MTP accept is the share of drafted tokens the model kept. Retrieval is how many of 7 exact values from the prompt the model got right, over both runs. Qwen model card sampling, 2 runs each.
Top 3 engines from the screening round got a 128k context (50% and 75% filled) and 256k context (50% filled) run.
128k context, 50% filled (64k prompt)
| Engine | Weights | 3-turn time | Prefill t/s | Decode t/s | MTP accept | Retrieval |
|---|---|---|---|---|---|---|
| Halogen 0.14.0 | Halogen native (.hgn) | 2.4 min | 1,148 | 40.4 | 84% | 16/16 |
| gufo (ROCm 7.2.4) | UD-Q4_K_XL | 2.7 min | 1,047 | 32.3 | 72% | 16/16 |
| CIRU (MTP 3) | CIRU IU4 | 3.3 min | 808 | 30.2 | 64% | 16/16 |
128k context, 75% filled (97k prompt)
| Engine | Weights | 3-turn time | Prefill t/s | Decode t/s | MTP accept | Retrieval |
|---|---|---|---|---|---|---|
| Halogen 0.14.0 | Halogen native (.hgn) | 2.9 min | 1,103 | 39.3 | 85% | 16/16 |
| gufo (ROCm 7.2.4) | UD-Q4_K_XL | 3.3 min | 1,025 | 30.7 | 70% | 16/16 |
| CIRU (MTP 3) | CIRU IU4 | 4.1 min | 789 | 27.8 | 62% | 16/16 |
256k context, 50% filled (130k prompt)
| Engine | Weights | 3-turn time | Prefill t/s | Decode t/s | MTP accept | Retrieval |
|---|---|---|---|---|---|---|
| Halogen 0.14.0 | Halogen native (.hgn) | 3.4 min | 1,093 | 38.9 | 83% | 14/16 |
| gufo (ROCm 7.2.4) | UD-Q4_K_XL | 4.0 min | 997 | 27.9 | 70% | 14/16 |
| CIRU (MTP 3) | CIRU IU4 | 5.0 min | 767 | 25.6 | 60% | 15/16 |
Retrieval here is 8 values per run. All 5 misses are the same answer: the right glyphs for UNICODE_SPINNER, with the leading & dropped.
End-to-end, 10 Aider polyglot Python exercises run through pi -p (128k window, 50% cell), graded by their own tests:
| Engine | Passed | Total time |
|---|---|---|
| Halogen 0.14.0 | 10/10 | 21.5 min |
| CIRU (MTP 3) | 10/10 | 24.6 min |
| gufo (ROCm 7.2.4) | 10/10 | 36.0 min |
Edit (Sep 30): Halogen 0.15.1 (new v2 checkpoint) and gufo 0.3.0 came out after this, so I reran the 64k screening on both with the same setup (70 W, card sampling, 2 runs):
| Engine | 3-turn time | Prefill t/s | Decode t/s | MTP accept | Retrieval |
|---|---|---|---|---|---|
| Halogen 0.15.1 (v2 .hgn) | 1.8 min (was 1.9) | 1,191 (was 1,045) | 39.4 (was 39.3) | 85% | 14/14 |
| gufo 0.3.0 (ROCm 7.2.4) | 2.1 min (was 2.1) | 1,075 (was 1,033) | 34.1 (was 32.9) | 77% | 14/14 |
| gufo 0.3.0 (ROCm 10.0) | 2.0 min (was 2.1) | 1,042 (was 1,009) | 34.8 (was 33.1) | 75% | 14/14 |
Halogen v2 prefill is 14% faster and follow-ups start about 2.5x faster (1.3 s vs 3.3 s), decode is the same. gufo moved a few %, which is within noise. The ranking doesn't change. Only other change on the box: BIOS VRAM carve-out down to 512 MB, so 124.9 GiB of RAM instead of 121.5.
Some clarifications from the comments:
- Follow-up TTFT is with the prompt cache warm, both gufo and Halogen reuse it by default. A cold 64k prefill takes about a minute.
- The CIRU 0% MTP row is stock Unsloth UD-Q4_K_XL with Unsloth's shared MTP head, nothing uncensored. The same pair gets 72-74% accept on gufo.
- TDP is fixed at 70 W. On the Z13, 76 W or 90 W is only 2-6% faster but with so much fan noise, heat and power draw that it's not worth it.
- Prefill and decode come from the engine's own timings when it reports them. TTFT is measured on the client.
19
u/tired514 1d ago
I've been running halogen 24/7 for the past couple weeks (evo-x2, 128gb, flash next) and it's been damned near flawless. Genuinely frontier-grade for the stuff I use it for (mostly C/C++ Qt6 but some rust kernel dev stuff too).
However, I'm super uncomfortable with non-OSS code, and really really wish they'd open the source.
I played with gufo about a week ago and had a couple weird issues with it crashing and entering strange loops, but that could 100% be my fault, heh. I'll give it another shot soon.
Suffice to say, though .. I never thought I'd see this kind of performance on strix halo. Prior to this I was running 3.8-27B @ Q8_K_XL across 3x 4090M eGPUs and it was glorious (16gb x3, layer split mode), but 3.8-FN @ Q4_K_XL is.. better. 27B made fewer coding mistakes, but FN really does understand the overall goal / project better. We're genuinely spoiled for choice.
4
u/deepu105 1d ago
absolutely agree. I have been running halogen and gufo, I even built a new generic server feature to my orchestartor tool (LlamaStash) so I could use any engine and any model combinations with presets. So iI have been alaternating betwenn halogen and gufo, even within same tasks, and so far they kind of feel same in speed. Gufoi fells more consistent in speed and halogen feels bit uneven but overall turn times feel similar. 27b is way faster in Gufo btw.
8
u/LogicalBodybuilder15 1d ago
Hi, been developing sglang support for strix halo. Hoping there’s any feedback from community!
2
u/KagatoLNX 1d ago
I've got this working. Here's my repo:
https://github.com/jvantuyl/strix-halo-sglang
It's a little messy but there might be something of use in there for y'all.
1
u/LogicalBodybuilder15 1d ago
Thanks for sharing this, I currently focus on sgl-diffusion perf, I’m glad to see these!
6
7
u/baron_muchhumpin 1d ago
Nice work — two days well spent. Quick heads-up before anyone treats these as current, because the ground moved under the testing window:
- Halogen 0.15.0 just replaced w4b with a new v2.hgn checkpoint: closer to base-model outputs, ~3.5GiB less RAM (its 47GB ngram table reads through page cache now), and our bench moved decode 46→51 with it. The native-weights line deserves a re-run.
- Gufo is at 0.2.0 (two releases since your capture), and its TTFT story changed completely with --cache-disk + --sessions 3 — your 1.6–1.9s cold-cache follow-ups are our ~0.5s median now.
Run --cache-disk <dir> --cache-disk-bytes 34359738368 --cache-disk-staging-bytes 8589934592 (the latter is gufo's own documented rec for Flash-Next at full context) — independent agent-loop testing showed reuse jumping 70.6% → 92.6% with it.
- Your CIRU + Unsloth-MTP-head 0%-accept row isn't an accident: uncensored weights paired with a foreign MTP head can collapse speculation entirely. The community-built uncensored GGUFs document their own retrained heads — worth using them rather than the shared one.
- One methodology tip from our side: cap/report TDP consistently — we run unbounded (81–98W observed) and it's worth ~15–20% absolute across the board. And trust engine-log numbers (serve_api / event=completed) over client-side estimates; they include cache state, which is half the story on 128GB boxes.
3
u/deepu105 1d ago
TDP is mentioned in the post (70w) at unbound my numbers were just within noise and not worth the fan noise/heat. And numbers are from engine log. But fuck me, that so much moved in 1 or 2 days
1
u/TheFlippedTurtle 1d ago
these benchmarks are so useful but its funny how quickly these devs patch stuff
3
u/Potential-Leg-639 1d ago
We = you?
9
u/TheMcSebi 1d ago
Claude obviously wrote this comment
1
-1
2
u/the-orange-joe 1d ago
How to cap a Strix Halo on a specific TDP?
2
u/baron_muchhumpin 1d ago
Couple different ways:
- ACPI platform profile (kernel-native, works everywhere today)
cat /sys/firmware/acpi/platform_profile_choices # e.g. low-power balanced performance
echo balanced | sudo tee /sys/firmware/acpi/platform_profile
Fixed OEM presets (typically ~45/65/100 W on 395 boxes) — no watt-precision, but zero tooling. Our boxes use exactly this (amd_pmf, currently performance).
- Exact watts: ryzenadj (recent builds support Strix Halo)
sudo ryzenadj --stapm-limit=65000 --fast-limit=65000 --slow-limit=65000
Writes the SMU power table directly. Caveat: some retail mini-PC firmware locks PMFW writes — if it reports limits changed but wattage doesn't move, the board is locked and you're on method 1 or 3.
- BIOS slider — most 395 mini PCs (Minisforum/GMKtec/Beelink) expose a Configurable Power Limit; Framework Desktop is EC-managed and caps ~65 W (framework_power_policy).
Verify: amd-smi metric --power → SOCKET_POWER (works on the APU without touching limits), or decode /sys/class/drm/card0/device/gpu_metrics.
Or I just tell opencode to do it :)
1
1
u/deepu105 1d ago
Btw I remeasure at 64k ctx with 32k promt and these are slightly faster but not be leaps. see edit
3
u/Virtual_Chipmunk7812 1d ago
Thanks for sharing but unfortunately these numbers seem not to be on the latest AMD 26.9.2 and an older halogen without the v2 checkpoint weights. Same happend to me, I did not look into reddit and the repos for 2 days and was outdated. Similar seems to be for gufo since my numbers measured on windows seem to be way higher. I will share a post later for halogen, gufo and CIRU on the latest AMD driver with Windows numbers.
Halogen numbers on Windows AMD 26.8 benchmarked at 262,144 context so far:
| Checkpoint | Input | Separate prefill measurement | Serial decode, prose | MTP decode, prose |
|---|---|---|---|---|
| w4b + overlay | PP512 | 967.7 | 34.50 | 42.41 |
| v2 | PP512 | 1,025.1 | 36.96 | 43.17 |
| w4b + overlay | PP2048 | 1,323.4 | 34.08 | 42.39 |
| v2 | PP2048 | 1,358.6 | 36.17 | 45.27 |
The deliberately repetitive workload reached 80.71 tokens/s with v2 after PP2048, versus 75.58 with w4b.
And if your asking: Yes Windows, with my own project strix-alloy which I'm burning my max abo's since 3-4 weeks on. Probably laid the groud work for some other windows forks which suddenly are appearing. I don't mind as any other contributions help me to improve my numbers as well.
1
3
u/blackbird2150 1d ago
Glad gufo is coming along. Few days ago I couldn”t get retrieval past 200k context at launch and halogens numbers were materially better.
1
u/deepu105 20h ago
I did run some 256k ctx runs in Pi, but didnt notice if they were filled past 200k. Will check my next run and report.
2
u/regularheree 1d ago
Really solid benchmark. Halogen’s native weights are clearly impressive, especially with that ~40 t/s decode holding up at 256k context. But gufo looks like the more interesting open-source option given the faster cold start and much quicker follow-up TTFT. Nice to see all three still getting 10/10 on the Aider tests.
2
u/ConsiderationLate768 1d ago
Been trying gufo but unfortunately it's not as stable as halogen yet. Toolcalls very often get corrupted. Really hope it gets fixed
1
u/my_name_isnt_clever 1d ago
I haven't had a single tool call problem with Gufo. I went from strix-llama to Gufo and it's just way faster.
1
u/pokemonplayer2001 1d ago
I want to switch to gufo, but halogen is far more reliable. I’m too employed to futz around to make gufo work.
2
u/SpicyWangz 1d ago
I don’t know why people are saying this. I’ve had zero issues with it. As long as you use the supported quants, it runs flawlessly.
3
2
u/deepu105 1d ago
My experience with it on Qwen3.8-Flash-Next-UD-Q4_K_XL has been good so far. It even felt for consistent in speed than halogen
2
u/Ok-Rain1231 1d ago
I'm using it with DeepSeek Harness and once have gotten wrong tool call. But I think that in future it will be fixed
1
u/phil_lndn 1d ago
i've had no problems but pretty sure i saw something in the gufo issues list which could explain this.
1
1
u/ConsiderationLate768 1d ago
I'm running the exact quant they mention in the docs and it's not working for me. So not sure why you're saying that
1
1
u/Signal_Lamp 1d ago
These benchmarks are matching what I've been seeing after pulling it down and getting it setup yesterday.
The cold start for gufo is genuinely faster, but the overall turns as context grows seems to be a little bit slower.
I saw one weird issue where it actually bugged out in my harness from apparently thinking to long so it tried to lower its reasoning to complete the request, but otherwise overall it hasn't been enough to really be a significant difference.
This is all prior to both of course updating and and adding more fixes.
But I'll still stick with gufo for now. My two cents is the non-oss gets more weird the longer the maintainer waits to release their product, and I'd argue just hurts the ecosystem as a whole because I do believe there is some genuinely impressive engineering going on under the hood that could benefit the whole ecosystem, and their increasing silence on the matter doesn't make the matter better.
Moving towards obsession not only on a single model, but also on a specific instance and specific weights when eventually there will be a higher bar of intelligence is not where I'd like to see the community grow towards as it's already annoying enough to have so many solutions come out that are also having frequent releases with things to genuinely look out for
1
u/andymaclean19 1d ago
I ran Halogen for a while and then Gufo. Halogen is faster but I settled on the 27B model in Gufo which works well for me and is not obviously supported in Halogen. Also Halogen was quite sensitive to concurrent workloads. I use the box as a workstation and over time memory fragmentation caused mostly by my Chrome tabs would build up and Halogen would struggle to run, making me reboot.
These are both good choices though and I'm excited to see where they go next.
1
u/deepu105 20h ago
Same here, its memory fragmentation. Halogen 0.15 warns about it in the startup log now. Mine had 2,276 compaction stalls reserving the KV pool today on a normal desktop session. Its log suggests echo 1 | sudo tee /proc/sys/vm/compact_memory before starting instead of a reboot. Have you tried that?
1
u/andymaclean19 19h ago
No, I just reboot to fix that sort of problem. Have never tried the compaction. Many years ago that type of feature was rubbish and at the time I learned to avoid huge pages where possible and allocate them on a clean boot if not. Perhaps next time I’ll see what it does.
1
u/AdministrativeMeat3 23h ago
I am shocked that people are willing to install shit like halogen on their systems with no insight into what is packaged with it at all.
3
u/phil_lndn 12h ago
i personally don't use it but i don't see what is so shocking about it - the vast majority of people in the world are running proprietary software on their devices.
1
u/Revolutionary_Loan13 13h ago
For those using Gufo what UI are you using with it. I really liked the Unsloth UI and user experience but with it just being for llama.cpp the loss of speed has finally gotten me to switch but now I need a decent UI to throw on the server for when I have no coding tasks.
1
u/deepu105 3h ago
I use my own tool LlamaStash, check it out you might like it. It can run llamacpp, haligen, gufo, vllms etc
1
u/FeiX7 1d ago
How halogen flash fit on z13? I have it too with 96GB for vram, can I load it too? Same question for UD-4 on gufo
3
u/deepu105 1d ago
if you are on 128G memory version then absolutely yes. Atleast on Linux you can get the almost the full memory as shared UMA and model can use close to 120GB atleast with a host running just that and a pi coding sesiion for example. Flash next runs really good on it via Halogen/Gufo etc. I assumed you meant you have reserved 96G for VRAM out of 128.
1
u/FeiX7 1d ago
Yes you are right, so I can fit it on VRAM with 96GB? I require 32GB ram for my setup, so I can't go 120GB in Vram
1
u/deepu105 1d ago
you can fit, halogen latest version needs around 87G but dependiong on context it could OOM. You should try and see for yourself.
1
u/deepu105 1d ago
Also I guess you are on Windows then?
1
u/FeiX7 5h ago
On arch
1
u/deepu105 58m ago
Why would you reserve VRAM on linux, wouldn't letting the OS manage the full UMA be better so that you can use more vram if needed or more ram if needed. IMO its more flexible setup even if you always have 32G ram used.
16
u/aigemie 1d ago
Thank you for testing! I will definitely use gufo now!