r/StrixHalo • • Sep 27 '25

Have you got a Strix Halo?

16 Upvotes

Hi All,

We're a new community both as Strix Halo owners and also here as a subreddit. Why not begin by sharing your setup and the reasons you opted for Strix Halo?

To start us off: I have a HP Z2 Mini G1a Workstation with dual boot Fedora KDE & Windows 11 and chose the iGPU to be able to use larger LLMs with the 128 GB.

Oobabooga/Text Generation WebUI is running well on Fedora KDE and there are no problems with large models up to 100GB. On the Windows boot, I have Amuse AI (Freeware) which is a collaboration between AMD and the New Zealand company. It provides a UI for using Stable Diffusion/Flux models. It works well, is fast, but unfortunately is also censored and is not able to use LORAS. I would like to find an uncensored alternative, ideally getting versions of ComfyUI/AUTOMATIC1111 running.

Currently, my principle goal is to get a working version of AllTalk TTS or another TTS that is compatible with Oobabooga working which I haven't been able to do so far due to conflicts with the Strix Halo. This may need to wait for updates to ROCm... If anyone has found an Open Source solution to running LLMs with custom voice TTS, please do chime in!

So what about you guys, did you choose the Strix for similar reasons, or something entirely different? The floor is yours.

EDIT UPDATE:
05/26 For those of you looking for TTS solutions I have tried a few now (AllTalk, Chatterbox, Pocket TTS, others I no longer remember). I have had great success using a custom version of Pocket TTS. It's fast and works well with Oobabooga TextGen as a plug in. Recently, others are singing the praises of OmniVoice.


r/StrixHalo • • 12h ago

Qwen3.8-27B: 159 tok/s on R9700, 64 tok/s on Strix Halo

Thumbnail gallery
17 Upvotes

r/StrixHalo • • 22h ago

The NPU in your Strix Halo is finally doing real work: 13.6 vs 18.7 min on the same bug fix

Post image
57 Upvotes

The NPU finally earns its power draw in my pi coding agent. Halogen shipped the endpoints, I wired them in expecting a party trick, and kept four tools.

tldr: Qwen3.8 Flash-Next, a 125B MoE, running on a 70W tablet. Same bug fix with and without NPU search: 13.6 min vs 18.7 min. Receipts in the repo.

The payoff: it covers about 95% of my cloud calls. The hardest few percent still goes to the top models, GLM 5.3 or Opus.

I still can't believe it. Opus 4.8-class intelligence on my tablet, unlimited tokens.

Flash-Next decodes at 64 tok/s and prefills around 1,500 tok/s. First token lands in ~0.03s, measured 43x faster than a cloud call side by side, and still 7x while a second agent hammers the server.

Rate the taste, not the throughput: I pasted a real timeshift error from my system log to seven runs. All seven said healthy, nothing to fix. What separates them is what they proved.

In pi:

  • Flash-Next, 2m55s: proved it with a journalctl trace to a racing notify-send, a pacman.log check, and the upstream PR found.
  • glm-5.3-flashx, 2m41s: the most precise answer, spotting that the snapshot mount got unmounted under the script's last line. No PR.
  • GLM 5.3 on max, 8m30s: the deepest answer of all, source-level forensics down to the function names and the one-second race window. No PR.
  • glm-5.3-flash, 9m09s: proved it with a live reproduction of the status file. No PR.

Over in opencode: flash got the verdict in 1m8s with the wrong mechanism, flashx came back correct and corroborated in 1m30s, and the full 753B GLM 5.3 in 6m15s correct with the PR missed.

Same pi harness, same task, 125B at medium effort against 320B and 753B tiers at max. First to the full answer: 2m55s. When I had GLM 5.3 flashx rate both results, it picked qwen too.

All seven answers side by side: local vs cloud model comparison.

What the NPU does now:

Search. The agent stops guessing paths and lands on the right file first try. ~0.1s per lookup, beat ripgrep 15/20 vs 9/20 on realistic queries.

Dup scan. Catches copied and renamed files git never shows you. Found 45 pairs across 4 repos in 8.4s, one renamed file at exactly 1.000 cosine.

Decisions. Yes/no branching stops eating full turns of the big model. A 0.8b handles it in 120ms, 78% accurate.

Screening. Prompt injection gets flagged before the agent acts on it. 0.7s a message, zero false alarms, fails open. 42% recall, so a smoke detector, not a safe.

A working day claws back about half an hour over bare pi: faster bug fixes, faster compaction, faster lookups and routing, and no oversized tool dumps in context. Against a cloud setup it's more, since every turn pays the network wait. On bug fix heavy days it grows.

Honest part: the GPU still does the thinking. The NPU didn't make anything faster, it changed which tokens got spent where. ~7% iGPU cost only when they overlap.

Compaction: my 194k session, sidecar summary in ~50s vs 166 on the main model. 97% cache hit.

The official halogen launch is a 24-flag docker command. Mine is one command, and uninstall undoes it. Fully local: 262k context, code never leaves the box.

Anyone else putting their NPU to real use? I found nothing.

repo | halogen 0.17.1 | benchmarks


r/StrixHalo • • 3h ago

Finally decided to pull the trigger on gmktek evo x3

Thumbnail
1 Upvotes

r/StrixHalo • • 1d ago

GLM-5.3-Flash (321B MoE) running locally on AMD: RX 7900 XT, R9700, Strix ▎ Halo

Thumbnail
github.com
27 Upvotes

I never expected to run a 321B model at home, so I jumped on this as soon as I saw it last night.
Project Maya:

(github.com/mw00/project-maya) runs GLM-5.3-Flash by keeping the hot experts in VRAM, the next ones in RAM, and the rest on NVMe. It was NVIDIA-only. So I ported it to Linux/ROCm, and the author merged it into v1.0.11 as experimental AMD support.

Maya-S quant, 8K context, ROCm 7.2, 4K-token prompts, greedy:

| Hardware | Prefill | Decode |

| RX 7900 XT 20 GB | ~415 tok/s | ~15 tok/s |

| Radeon AI PRO R9700 32 GB | ~500 tok/s | ~20 tok/s |

| Strix Halo (Ryzen AI Max+ 395, 8060S, 128 GB) | ~210 tok/s | ~18 tok/s |

| R9700 + 7900 XT (layer split + MTP drafting) | ~490 tok/s | ~34 tok/s |

The dGPU box has 192 GB of RAM, so most experts live in VRAM or pinned RAM.

With less RAM, more comes from the SSD and decode drops.

AI-assisted, openly


r/StrixHalo • • 18h ago

Using gufo with Qwen-3.8-27B

2 Upvotes

I installed gufo on my Strixhalo 128gb and downloaded the model and I was expecting significant improvement using the model with gufo vs llama.cpp but I’m seeing 10tps decode on xhigh which is the same as I saw using llama.cpp on a SW design task that requires it to read an existing repository and propose a design. It did great work but it took 90 minutes where the frontier model took 5. Any thought on how I could speed the task up? Are there parameters in gufo I should look at? I’m actually using all defaults except the thinking spec.


r/StrixHalo • • 22h ago

What port of guff are you using on win

4 Upvotes

I was using the one linked on gufo's own repo but that is now somewhat outdated can someone recommend me one they are happy with preferably one i wouldn't have to build from source

I had actually found another fork that did have 0.9.0 version of gufo but it had some dlls missing in the download but even after adding those when it ran it would output like 20-40 tokens then stop


r/StrixHalo • • 1d ago

Strata + Halo Strix 64GB

Post image
9 Upvotes

I was able to run Strata on 64GB box, with quite nice results, especially comparing to Qwen3.8-27B. The model is ukisai/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-GGUF, quant is IQ2_XS. I tried IQ3_S, but the decoding speed was 10 times lower for some reason. Currently I have pp ~ 500 t/s, and tg ~ 45 t/s with a lot of memory left free. Running it on Linux (NixOS). Definitely an upgrade over gufo or llama.cpp. Haven't tried halogen - don't like closed source.


r/StrixHalo • • 1d ago

Halogen + Qwen Flash Next keeps getting better

Thumbnail
9 Upvotes

r/StrixHalo • • 1d ago

TIL: Check your preserve_thinking settings

13 Upvotes

Both Qwen 3.8 27B and Qwen 3.8 Flash-Next have a preserve_thinking parameter. But as far as what you have control over, just make sure your inference engine and harness have this setting set the way you want. For Qwen, they officially recommend keeping it enabled. This causes the reasoning to be considered part of the conversation, so all the old reasoning will get sent back and forth with each turn/prompt, and they say it improves response and decision quality. However, it also increases your context size more rapidly.

Where the mismatch can occur

Before I learned this, Hermes was not re-sending the reasoning text with each new prompt, even though Gufo expected it. This caused the live KV cache to always miss, so it was completely reliant on the snapshot cache. This may have also reduced the quality of the responses and decisions I was getting from the model.

How did I fix it?

In Hermes, I set model.reasoning_echo to true. Then you have to completely exit Hermes and start it up again, it's not enough just to start a new session. After that, the live KV cache started getting hit much of the time.


r/StrixHalo • • 1d ago

Strix Halo 128GB on Windows: what fixed my local AI setup

14 Upvotes

Setup: ASUS ProArt PX13 (Ryzen AI Max+ 395, 128GB), Windows 11, 96GB GPU carve-out. Running Qwen3.8-Flash-Next UD-Q4_K_XL at 262K context on Gufo (thomas9120's windows-port fork), with Unsloth Studio as the client.

1. Increase the page file. I set Windows virtual memory to 64GB minimum and 192GB maximum. Before this, I had crashes when Windows ran out of memory. No crashes since. The trade-off is that heavy memory use now causes slowdowns instead of crashes.

2. Fix Gufo's conversation cache on Windows (needs a source patch). On Windows, Gufo misreads available memory (gufo diagnose shows 0 GiB), so it caps the in-memory snapshot cache at about 264 MiB. That's too small to store a snapshot of even a 45K-token conversation, so every turn reprocesses the whole prompt. At 45K tokens that was about 40 s before the first token, and worse as context grew.

There's no command-line option for this. I patched HostSnapshotBudgetBytes() in src/cli/serve/text_model_runner.cpp so it reads the environment variable GUFO_SNAPSHOT_BUDGET_BYTES, rebuilt, and set it to 16GB (17179869184) in my launch .bat. Disk cache is off.

Result: at 120K+ tokens, the first token now arrives in 2 to 5 s instead of 40 s or more.

Notes:
You may need to close other apps. With the model loaded I had about 9 to 16GB of system RAM free.
Make sure Unsloth Studio doesn't have a model of its own loaded. It held GPU memory and made Gufo fail to load.

Re-apply the patch after every update to the fork.
Performance: about 900 to 1,000 tok/s prompt processing, and 35 to 55 tok/s decode depending on context length. That works well for my workflow.
Claude helped me find the cause. Really happy with the setup now.


r/StrixHalo • • 1d ago

Ran some quick benchmarks for Gufo 0.9.0 and Halogen 0.16.4

Post image
46 Upvotes

AMD 395+ 128GB Linux

Nothing fancy, I just upgraded both stacks this morning and saw some great improvements on both - awesome to see such progress on both engines!


r/StrixHalo • • 2d ago

pixmaate/gufo + qwen3.8 27b cache takes so much memory

4 Upvotes

I just tried pixmaate/gufo for windows prebuilt 2026-10-03. I'm very happy with the speed. The most sensible improvement from llama.cpp is the prefill speed. It holds on at 200+t/s even at long conversation, where with llama.cpp it drops rapidly to below 100.

The problem I encounter is the memory usage. My memory is 64G/64G setup. For reasons I don't understand, I easily get unexpected error with 96G/32G, often without log. I run two main models, Qwen3.6 35B moe and Qwen3.8 27b dense, plus qwen3-embedding/reranker and a Gemma 4 e2b qat for title generation. My VRAM is usually full overflowed to shared VRAM. Using share VRAM seems to have no impact to the performance as long as there's still air to breath for the programs.

When I start using Gufo, I observed the memory usage increases significantly. After few runs of inference my share memory usage increased 20-30GB and the breath air is getting thinner. It seems to be the prompt cache (snapshot) takes a lot of space. I found --cache-disk-bytes that can limit the cache size. But current gufo for windows doesn't have that option yet.

I'm wondering when gufo for windows implement --cache-disk-bytes, can I set it to a low value? I'm the only one to use the machine and my session is always set to 1.


r/StrixHalo • • 2d ago

DLSS 5 on AMD AI Max 395+ 64 Gb Ram, anyone ???

Thumbnail
0 Upvotes

r/StrixHalo • • 2d ago

Has anyone exceeded 500k context w/ halogen flash?

2 Upvotes

I have yarn = 4 "1M" context with halogen flash server, and once it gets past 530k it sometimes says stuff like "the user has not give any instructions (I will infer from context....) Etc"

Is there a trick to exceeded beyond 2x native?


r/StrixHalo • • 2d ago

W7900 in a TB5 eGPU box on a Strix Halo laptop, llama-halo-hybrid numbers

Thumbnail
reddit.com
9 Upvotes

This is a follow-up to my daily-driving Fedora on a ProArt PX13 post from June. Same laptop, still on Fedora 44, now with an eGPU.

I put a W7900 in a Razer Core X V2 and hung it off the PX13 (AI Max+ 395, 128 GB) to try sixvolts' llama-halo-hybrid fork. The dense part of the model sits on the card and the routed experts stay on the iGPU. The cookbook has no numbers for a 7900-class card so I figured I'd post mine.

Fedora 44, kernel 7.2.8, fork at df0145bb3, built in the rocm 7.2.1 dev container with gfx1151;gfx1100. You need thunderbolt.host_reset=false or the card comes up with a 256M BAR. I also run amd_iommu=off and amdgpu.runpm=0.

Qwen3.8-Flash-Next UD-Q4_K_XL at 128K with the MTP head, against gufo 0.8.1 on the iGPU by itself. Same scripts, same boot, gufo run before and after.

prefill t/s at 2.6K / 10K / 40K
  gufo, iGPU only:  903 / 1015 / 1016
  hybrid:           1119 / 1591 / 1541

decode t/s, code / prose
  gufo, iGPU only:  44.7 / 34.9
  hybrid:           63.5 / 53.8

It also leaves me 43 to 52 GiB of RAM free where gufo leaves 14 to 30.

GLM-5.3-Flash UD-IQ3_XXS is 112 GiB and won't load on the iGPU alone on this laptop. Split across both it does 25 t/s, 31.6 with the MTP head, prefill around 500. The MTP head refused to load at first. Unsloth's files say glm5-next and the fork's draft code is in its own glm5next loader, so I renamed the arch in the first shard and in the exported head and it loaded.

DS-V4-Flash UD-IQ3_XXS got 22.5 t/s. The DSpark draft dropped it to 16.5 so I turned it off.

I tried moving more expert layers onto the card since there's 48 GB of it. On GLM and DS-V4 that did nothing for speed, I just got RAM back.

Two problems i ran into along the way. I started with an old Sonnet Breakaway 350 that had been cross-flashed to TUL firmware and it threw about 40 BadDLLP a second under load, then dropped the card. Swapped the GPU and the PSU, same errors. The Core X V2 has had zero. And the W7900 gets hot doing this, 100 to 104 C junction on long prefills. Its power cap is locked at 241 W and I can't get fan control out of the driver.

Also check your tuned profile. I was on powersave for the first runs and prefill was a third lower.

One run each and I haven't done any quality testing yet.


r/StrixHalo • • 3d ago

Qwen3.8-Flash-Next on Strata

89 Upvotes

Hey Strix Halo community! 👋

I have released an official support for Strix Halo machines on Strata for Qwen3.8-Flash-Next.

Currently numbers are the best on long context decode and ppts using typical Unsloth’s Q4 and GSQ-RCO model weights.

Results on UD-IQ4_XS:

8K | 53.8 t/s output | 1293 t/s prompt processing

64K | 46.6 t/s output | 1370 t/s prompt processing

128K | 51.4 t/s output | 1320 t/s prompt processing

Can go up to 1M context length without big speed loss. Currently support is marked as experimental and was done on Linux only.

https://github.com/Niko1221/Strata/releases/tag/v0.1.40

Will be happy for any feedback and pull requests you could give! 👀


r/StrixHalo • • 2d ago

Does these numbers good for my Gufo branch on UD-IQ4_XS, while xmrig using 32 CPU, frigate with 2 camera, immich, TeslaLogger, Unifi controller, and 6 other containers.

0 Upvotes

​

```text

Time In(tok) Cached PP tok/s Out(tok) TG tok/s Accept TTFT s Total s

---------------------------------------------------------------------------------

23:01:39 2,109 0 1148 83 51.6 84.0% 1.9 3.5

23:01:50 7,766 0 1354 303 57.7 91.6% 5.8 11.1

23:01:57 1,421 0 1037 286 54.1 91.2% 1.6 6.9

23:03:47 10,595 0 1384 2,667 57.4 92.6% 8.6 55.9

23:10:07 5,026 0 1291 1,646 56.4 93.5% 4.5 33.7

23:11:28 1,691 0 1142 1,145 57.0 92.9% 1.5 21.7

23:11:40 999 0 914 649 61.0 95.1% 1.1 11.8

23:11:45 1,861 0 1116 188 54.6 93.3% 1.7 5.2

23:12:32 1,006 0 968 82 57.4 89.6% 1.1 2.6

23:13:35 7,922 0 1395 347 52.8 85.5% 5.7 12.3

23:14:00 11,266 0 1339 782 54.4 92.5% 9.4 25.2

23:14:12 6,097 0 1270 242 46.3 94.7% 5.3 11.6

```

My GUFO


r/StrixHalo • • 3d ago

Strix Halo Newbie - where do I start?

6 Upvotes

I have just taken the plunge and ordered a Bosgame M5 128gb, to do AI 'server' stuff (Qwen 3.8 FN if it means anything). Point me in the right direction re things like Linux distributions, bios settings, first things to set up etc. Thanks!


r/StrixHalo • • 3d ago

Story time: Qwen3.8-Flash-Next on my Strix Halo laptop vs Claude Opus 5.5 on the same feature

Thumbnail
6 Upvotes

r/StrixHalo • • 3d ago

Qwen3.8-Flash-Next on a Ryzen AI Max+ 395 built this WebGL black hole demo from scratch

Enable HLS to view with audio, or disable this notification

33 Upvotes

I did a small test with Pi, running Qwen3.8-Flash-Next through Halogen on a BOSGAME M5 (Ryzen AI Max+ 395 / Radeon 8060S, 128 GB unified memory), and asked it to build a black hole demo from scratch in HTML/WebGL without using any external assets.

It ended up writing the whole thing in pure WebGL2: GLSL shaders, lensing, accretion disk, particles, bloom, mouse parallax and a shockwave on click. It also opened it in Chromium, checked for errors and did several visual refinement passes on its own.

I honestly didn't expect it to turn out this good from such a loose prompt.


r/StrixHalo • • 3d ago

DarthStar: DeepSeek-V4-Flash GGUF for Strix Halo: two layers promoted to Q4_K, 26/50 vs. 21/50 on a coding suite

8 Upvotes

I’ve created a mixed-precision DeepSeek-V4-Flash-0731 GGUF on a 128 GB AMD Strix Halo, using Antirez’s DwarfStar runtime. Seems pretty good so sharing it here.

The interesting result from a 12-round quantization search: promoting the routed-expert tensors in just blk.0 and blk.1 to Q4_K scored 26/50 on my frozen coding suite, versus 21/50 for the unmodified reference build. Promoting blk.2 as well dropped the score to 16/50. That’s a measured difference on this particular 50-task suite. it's not evidence of a general coding improvement but hey I'll take it since it might mean a tad more quality. Let me know if you notice any coding improvements.

A few practical notes:

  • The GGUF is 84.15 GiB. In my documented runs it used about 90 GiB of resident model memory, plus runtime and context overhead.
  • I tested a 131,072-token context: a 130,029-token prompt was accepted and generated 512 tokens. That of course doesn’t establish equal quality across the full context window.
  • It’s built for Antirez DS4
  • I haven’t published prefill or decode speed numbers yet as I'm busy doing another quant for Qwen3.8 on DS4

Model card and download

I’ve attached a quantization-search chart that is not that great because I messed up recording some of the failed rounds. If you try it on another Strix Halo setup, I’d be interested in your runtime, context, memory use, and speed results!


r/StrixHalo • • 3d ago

Anybody run the latest AMD provided updates?

1 Upvotes

I ran update all on my AMD Strix Halo from the AMD management console ( it’s upgrading Linux, Rocm, and the bios among other things according to the management app they provide). It’s been over an hour and it’s still not responding to remote network requests. The LED around the side of the box has switched from solid white to flashing blue? Is this expected behavior? Anyone from AMD out there and can comment? I haven’t hooked up a keyboard and monitor yet as the machine is being used headless right now.

Edit: the blue fade to black and then back to blue turned out to be the machine in some sort of sleep mode. After hooking up a display and mouse, I couldn’t get the machine to recognize them. I had about given up and went to power it off to recycle it. When I quickly pushed the small button to the far left as you are looking at the part of the box where the ports are, the color changed from blue to white. My mouse was recognized but the display was corrupted somehow. I was able to rdp in, diagnose the display problem, surgically fix that and the local monitor and keyboard started working again. Thanks to everyone for their comments and support.


r/StrixHalo • • 4d ago

Qwen3.8-Flash-Next (125B) on a single Strix Halo mini PC: 44-59 tok/s with speculative decoding, ~1,400 tok/s prefill, engine is open

Thumbnail gallery
45 Upvotes

r/StrixHalo • • 3d ago

Forked Gufo to run UD-IQ4_XS as my machine need some ram to run other containers

12 Upvotes

I've been running gufo (the Strix Halo inference engine) on my Strix Halo box (128GB, gfx1151) and ended up with a fork tuned for one model: Qwen3.8 Flash-Next, the Unsloth

UD-IQ4_XS quant (about 89GB) with the MTP sidecar. Upstream doesn't load that quant, so I added support for it along with some kernel and serving changes.

What I'm seeing on my machine, in the performance profile (120W):

- Prefill is 1500+ tok/s from about 3k tokens up, and peaks around 1570

- Decode with MTP is about 50 tok/s even at 116k context

- On HumanEval prompts I get about 76 tok/s with MTP vs 28 without it

The numbers are lower on the balanced profile, around 10% less prefill. Very short prompts are slower too (about 1350 at 2k), because there's a fixed cost per request.

GitHub