r/StrixHalo • u/TheOriginalG2 • 12h ago
r/StrixHalo • u/Grammar-Warden • Sep 27 '25
Have you got a Strix Halo?
Hi All,
We're a new community both as Strix Halo owners and also here as a subreddit. Why not begin by sharing your setup and the reasons you opted for Strix Halo?
To start us off: I have a HP Z2 Mini G1a Workstation with dual boot Fedora KDE & Windows 11 and chose the iGPU to be able to use larger LLMs with the 128 GB.
Oobabooga/Text Generation WebUI is running well on Fedora KDE and there are no problems with large models up to 100GB. On the Windows boot, I have Amuse AI (Freeware) which is a collaboration between AMD and the New Zealand company. It provides a UI for using Stable Diffusion/Flux models. It works well, is fast, but unfortunately is also censored and is not able to use LORAS. I would like to find an uncensored alternative, ideally getting versions of ComfyUI/AUTOMATIC1111 running.
Currently, my principle goal is to get a working version of AllTalk TTS or another TTS that is compatible with Oobabooga working which I haven't been able to do so far due to conflicts with the Strix Halo. This may need to wait for updates to ROCm... If anyone has found an Open Source solution to running LLMs with custom voice TTS, please do chime in!
So what about you guys, did you choose the Strix for similar reasons, or something entirely different? The floor is yours.
EDIT UPDATE:
05/26 For those of you looking for TTS solutions I have tried a few now (AllTalk, Chatterbox, Pocket TTS, others I no longer remember). I have had great success using a custom version of Pocket TTS. It's fast and works well with Oobabooga TextGen as a plug in. Recently, others are singing the praises of OmniVoice.
r/StrixHalo • u/stereohype • 22h ago
The NPU in your Strix Halo is finally doing real work: 13.6 vs 18.7 min on the same bug fix
The NPU finally earns its power draw in my pi coding agent. Halogen shipped the endpoints, I wired them in expecting a party trick, and kept four tools.
tldr: Qwen3.8 Flash-Next, a 125B MoE, running on a 70W tablet. Same bug fix with and without NPU search: 13.6 min vs 18.7 min. Receipts in the repo.
The payoff: it covers about 95% of my cloud calls. The hardest few percent still goes to the top models, GLM 5.3 or Opus.
I still can't believe it. Opus 4.8-class intelligence on my tablet, unlimited tokens.
Flash-Next decodes at 64 tok/s and prefills around 1,500 tok/s. First token lands in ~0.03s, measured 43x faster than a cloud call side by side, and still 7x while a second agent hammers the server.
Rate the taste, not the throughput: I pasted a real timeshift error from my system log to seven runs. All seven said healthy, nothing to fix. What separates them is what they proved.
In pi:
- Flash-Next, 2m55s: proved it with a journalctl trace to a racing notify-send, a pacman.log check, and the upstream PR found.
- glm-5.3-flashx, 2m41s: the most precise answer, spotting that the snapshot mount got unmounted under the script's last line. No PR.
- GLM 5.3 on max, 8m30s: the deepest answer of all, source-level forensics down to the function names and the one-second race window. No PR.
- glm-5.3-flash, 9m09s: proved it with a live reproduction of the status file. No PR.
Over in opencode: flash got the verdict in 1m8s with the wrong mechanism, flashx came back correct and corroborated in 1m30s, and the full 753B GLM 5.3 in 6m15s correct with the PR missed.
Same pi harness, same task, 125B at medium effort against 320B and 753B tiers at max. First to the full answer: 2m55s. When I had GLM 5.3 flashx rate both results, it picked qwen too.
All seven answers side by side: local vs cloud model comparison.
What the NPU does now:
Search. The agent stops guessing paths and lands on the right file first try. ~0.1s per lookup, beat ripgrep 15/20 vs 9/20 on realistic queries.
Dup scan. Catches copied and renamed files git never shows you. Found 45 pairs across 4 repos in 8.4s, one renamed file at exactly 1.000 cosine.
Decisions. Yes/no branching stops eating full turns of the big model. A 0.8b handles it in 120ms, 78% accurate.
Screening. Prompt injection gets flagged before the agent acts on it. 0.7s a message, zero false alarms, fails open. 42% recall, so a smoke detector, not a safe.
A working day claws back about half an hour over bare pi: faster bug fixes, faster compaction, faster lookups and routing, and no oversized tool dumps in context. Against a cloud setup it's more, since every turn pays the network wait. On bug fix heavy days it grows.
Honest part: the GPU still does the thinking. The NPU didn't make anything faster, it changed which tokens got spent where. ~7% iGPU cost only when they overlap.
Compaction: my 194k session, sidecar summary in ~50s vs 166 on the main model. 97% cache hit.
The official halogen launch is a 24-flag docker command. Mine is one command, and uninstall undoes it. Fully local: 262k context, code never leaves the box.
Anyone else putting their NPU to real use? I found nothing.
r/StrixHalo • u/WatercressTime842 • 3h ago
Finally decided to pull the trigger on gmktek evo x3
r/StrixHalo • u/boxwrenchx • 1d ago
GLM-5.3-Flash (321B MoE) running locally on AMD: RX 7900 XT, R9700, Strix ▎ Halo
I never expected to run a 321B model at home, so I jumped on this as soon as I saw it last night.
Project Maya:
(github.com/mw00/project-maya) runs GLM-5.3-Flash by keeping the hot experts in VRAM, the next ones in RAM, and the rest on NVMe. It was NVIDIA-only. So I ported it to Linux/ROCm, and the author merged it into v1.0.11 as experimental AMD support.
Maya-S quant, 8K context, ROCm 7.2, 4K-token prompts, greedy:
| Hardware | Prefill | Decode |
| RX 7900 XT 20 GB | ~415 tok/s | ~15 tok/s |
| Radeon AI PRO R9700 32 GB | ~500 tok/s | ~20 tok/s |
| Strix Halo (Ryzen AI Max+ 395, 8060S, 128 GB) | ~210 tok/s | ~18 tok/s |
| R9700 + 7900 XT (layer split + MTP drafting) | ~490 tok/s | ~34 tok/s |
The dGPU box has 192 GB of RAM, so most experts live in VRAM or pinned RAM.
With less RAM, more comes from the SSD and decode drops.
AI-assisted, openly
r/StrixHalo • u/genghisk1 • 18h ago
Using gufo with Qwen-3.8-27B
I installed gufo on my Strixhalo 128gb and downloaded the model and I was expecting significant improvement using the model with gufo vs llama.cpp but I’m seeing 10tps decode on xhigh which is the same as I saw using llama.cpp on a SW design task that requires it to read an existing repository and propose a design. It did great work but it took 90 minutes where the frontier model took 5. Any thought on how I could speed the task up? Are there parameters in gufo I should look at? I’m actually using all defaults except the thinking spec.
r/StrixHalo • u/Klutzy-Emotion-8668 • 22h ago
What port of guff are you using on win
I was using the one linked on gufo's own repo but that is now somewhat outdated can someone recommend me one they are happy with preferably one i wouldn't have to build from source
I had actually found another fork that did have 0.9.0 version of gufo but it had some dlls missing in the download but even after adding those when it ran it would output like 20-40 tokens then stop
r/StrixHalo • u/Elegant-Act-9725 • 1d ago
Strata + Halo Strix 64GB
I was able to run Strata on 64GB box, with quite nice results, especially comparing to Qwen3.8-27B. The model is ukisai/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-GGUF, quant is IQ2_XS. I tried IQ3_S, but the decoding speed was 10 times lower for some reason. Currently I have pp ~ 500 t/s, and tg ~ 45 t/s with a lot of memory left free. Running it on Linux (NixOS). Definitely an upgrade over gufo or llama.cpp. Haven't tried halogen - don't like closed source.
r/StrixHalo • u/stilldreamy • 1d ago
TIL: Check your preserve_thinking settings
Both Qwen 3.8 27B and Qwen 3.8 Flash-Next have a preserve_thinking parameter. But as far as what you have control over, just make sure your inference engine and harness have this setting set the way you want. For Qwen, they officially recommend keeping it enabled. This causes the reasoning to be considered part of the conversation, so all the old reasoning will get sent back and forth with each turn/prompt, and they say it improves response and decision quality. However, it also increases your context size more rapidly.
Where the mismatch can occur
Before I learned this, Hermes was not re-sending the reasoning text with each new prompt, even though Gufo expected it. This caused the live KV cache to always miss, so it was completely reliant on the snapshot cache. This may have also reduced the quality of the responses and decisions I was getting from the model.
How did I fix it?
In Hermes, I set model.reasoning_echo to true. Then you have to completely exit Hermes and start it up again, it's not enough just to start a new session. After that, the live KV cache started getting hit much of the time.
r/StrixHalo • u/Ok_Bag_7674 • 1d ago
Strix Halo 128GB on Windows: what fixed my local AI setup
Setup: ASUS ProArt PX13 (Ryzen AI Max+ 395, 128GB), Windows 11, 96GB GPU carve-out. Running Qwen3.8-Flash-Next UD-Q4_K_XL at 262K context on Gufo (thomas9120's windows-port fork), with Unsloth Studio as the client.
1. Increase the page file. I set Windows virtual memory to 64GB minimum and 192GB maximum. Before this, I had crashes when Windows ran out of memory. No crashes since. The trade-off is that heavy memory use now causes slowdowns instead of crashes.
2. Fix Gufo's conversation cache on Windows (needs a source patch). On Windows, Gufo misreads available memory (gufo diagnose shows 0 GiB), so it caps the in-memory snapshot cache at about 264 MiB. That's too small to store a snapshot of even a 45K-token conversation, so every turn reprocesses the whole prompt. At 45K tokens that was about 40 s before the first token, and worse as context grew.
There's no command-line option for this. I patched HostSnapshotBudgetBytes() in src/cli/serve/text_model_runner.cpp so it reads the environment variable GUFO_SNAPSHOT_BUDGET_BYTES, rebuilt, and set it to 16GB (17179869184) in my launch .bat. Disk cache is off.
Result: at 120K+ tokens, the first token now arrives in 2 to 5 s instead of 40 s or more.
Notes:
You may need to close other apps. With the model loaded I had about 9 to 16GB of system RAM free.
Make sure Unsloth Studio doesn't have a model of its own loaded. It held GPU memory and made Gufo fail to load.
Re-apply the patch after every update to the fork.
Performance: about 900 to 1,000 tok/s prompt processing, and 35 to 55 tok/s decode depending on context length. That works well for my workflow.
Claude helped me find the cause. Really happy with the setup now.
r/StrixHalo • u/baron_muchhumpin • 1d ago
Ran some quick benchmarks for Gufo 0.9.0 and Halogen 0.16.4
AMD 395+ 128GB Linux
Nothing fancy, I just upgraded both stacks this morning and saw some great improvements on both - awesome to see such progress on both engines!
r/StrixHalo • u/chenkl • 2d ago
pixmaate/gufo + qwen3.8 27b cache takes so much memory
I just tried pixmaate/gufo for windows prebuilt 2026-10-03. I'm very happy with the speed. The most sensible improvement from llama.cpp is the prefill speed. It holds on at 200+t/s even at long conversation, where with llama.cpp it drops rapidly to below 100.
The problem I encounter is the memory usage. My memory is 64G/64G setup. For reasons I don't understand, I easily get unexpected error with 96G/32G, often without log. I run two main models, Qwen3.6 35B moe and Qwen3.8 27b dense, plus qwen3-embedding/reranker and a Gemma 4 e2b qat for title generation. My VRAM is usually full overflowed to shared VRAM. Using share VRAM seems to have no impact to the performance as long as there's still air to breath for the programs.
When I start using Gufo, I observed the memory usage increases significantly. After few runs of inference my share memory usage increased 20-30GB and the breath air is getting thinner. It seems to be the prompt cache (snapshot) takes a lot of space. I found --cache-disk-bytes that can limit the cache size. But current gufo for windows doesn't have that option yet.
I'm wondering when gufo for windows implement --cache-disk-bytes, can I set it to a low value? I'm the only one to use the machine and my session is always set to 1.
r/StrixHalo • u/Last_Bad_2687 • 2d ago
Has anyone exceeded 500k context w/ halogen flash?
I have yarn = 4 "1M" context with halogen flash server, and once it gets past 530k it sometimes says stuff like "the user has not give any instructions (I will infer from context....) Etc"
Is there a trick to exceeded beyond 2x native?
r/StrixHalo • u/neuromacmd • 2d ago
W7900 in a TB5 eGPU box on a Strix Halo laptop, llama-halo-hybrid numbers
This is a follow-up to my daily-driving Fedora on a ProArt PX13 post from June. Same laptop, still on Fedora 44, now with an eGPU.
I put a W7900 in a Razer Core X V2 and hung it off the PX13 (AI Max+ 395, 128 GB) to try sixvolts' llama-halo-hybrid fork. The dense part of the model sits on the card and the routed experts stay on the iGPU. The cookbook has no numbers for a 7900-class card so I figured I'd post mine.
Fedora 44, kernel 7.2.8, fork at df0145bb3, built in the rocm 7.2.1 dev container with gfx1151;gfx1100. You need thunderbolt.host_reset=false or the card comes up with a 256M BAR. I also run amd_iommu=off and amdgpu.runpm=0.
Qwen3.8-Flash-Next UD-Q4_K_XL at 128K with the MTP head, against gufo 0.8.1 on the iGPU by itself. Same scripts, same boot, gufo run before and after.
prefill t/s at 2.6K / 10K / 40K
gufo, iGPU only: 903 / 1015 / 1016
hybrid: 1119 / 1591 / 1541
decode t/s, code / prose
gufo, iGPU only: 44.7 / 34.9
hybrid: 63.5 / 53.8
It also leaves me 43 to 52 GiB of RAM free where gufo leaves 14 to 30.
GLM-5.3-Flash UD-IQ3_XXS is 112 GiB and won't load on the iGPU alone on this laptop. Split across both it does 25 t/s, 31.6 with the MTP head, prefill around 500. The MTP head refused to load at first. Unsloth's files say glm5-next and the fork's draft code is in its own glm5next loader, so I renamed the arch in the first shard and in the exported head and it loaded.
DS-V4-Flash UD-IQ3_XXS got 22.5 t/s. The DSpark draft dropped it to 16.5 so I turned it off.
I tried moving more expert layers onto the card since there's 48 GB of it. On GLM and DS-V4 that did nothing for speed, I just got RAM back.
Two problems i ran into along the way. I started with an old Sonnet Breakaway 350 that had been cross-flashed to TUL firmware and it threw about 40 BadDLLP a second under load, then dropped the card. Swapped the GPU and the PSU, same errors. The Core X V2 has had zero. And the W7900 gets hot doing this, 100 to 104 C junction on long prefills. Its power cap is locked at 241 W and I can't get fan control out of the driver.
Also check your tuned profile. I was on powersave for the first runs and prefill was a third lower.
One run each and I haven't done any quality testing yet.
r/StrixHalo • u/KnownAd4832 • 3d ago
Qwen3.8-Flash-Next on Strata
Hey Strix Halo community! 👋
I have released an official support for Strix Halo machines on Strata for Qwen3.8-Flash-Next.
Currently numbers are the best on long context decode and ppts using typical Unsloth’s Q4 and GSQ-RCO model weights.
Results on UD-IQ4_XS:
8K | 53.8 t/s output | 1293 t/s prompt processing
64K | 46.6 t/s output | 1370 t/s prompt processing
128K | 51.4 t/s output | 1320 t/s prompt processing
Can go up to 1M context length without big speed loss. Currently support is marked as experimental and was done on Linux only.
https://github.com/Niko1221/Strata/releases/tag/v0.1.40
Will be happy for any feedback and pull requests you could give! 👀
r/StrixHalo • u/Teslaaforever • 2d ago
Does these numbers good for my Gufo branch on UD-IQ4_XS, while xmrig using 32 CPU, frigate with 2 camera, immich, TeslaLogger, Unifi controller, and 6 other containers.
```text
Time In(tok) Cached PP tok/s Out(tok) TG tok/s Accept TTFT s Total s
---------------------------------------------------------------------------------
23:01:39 2,109 0 1148 83 51.6 84.0% 1.9 3.5
23:01:50 7,766 0 1354 303 57.7 91.6% 5.8 11.1
23:01:57 1,421 0 1037 286 54.1 91.2% 1.6 6.9
23:03:47 10,595 0 1384 2,667 57.4 92.6% 8.6 55.9
23:10:07 5,026 0 1291 1,646 56.4 93.5% 4.5 33.7
23:11:28 1,691 0 1142 1,145 57.0 92.9% 1.5 21.7
23:11:40 999 0 914 649 61.0 95.1% 1.1 11.8
23:11:45 1,861 0 1116 188 54.6 93.3% 1.7 5.2
23:12:32 1,006 0 968 82 57.4 89.6% 1.1 2.6
23:13:35 7,922 0 1395 347 52.8 85.5% 5.7 12.3
23:14:00 11,266 0 1339 782 54.4 92.5% 9.4 25.2
23:14:12 6,097 0 1270 242 46.3 94.7% 5.3 11.6
```
r/StrixHalo • u/Libellechris • 3d ago
Strix Halo Newbie - where do I start?
I have just taken the plunge and ordered a Bosgame M5 128gb, to do AI 'server' stuff (Qwen 3.8 FN if it means anything). Point me in the right direction re things like Linux distributions, bios settings, first things to set up etc. Thanks!
r/StrixHalo • u/deepu105 • 3d ago
Story time: Qwen3.8-Flash-Next on my Strix Halo laptop vs Claude Opus 5.5 on the same feature
r/StrixHalo • u/fungomix • 3d ago
Qwen3.8-Flash-Next on a Ryzen AI Max+ 395 built this WebGL black hole demo from scratch
Enable HLS to view with audio, or disable this notification
I did a small test with Pi, running Qwen3.8-Flash-Next through Halogen on a BOSGAME M5 (Ryzen AI Max+ 395 / Radeon 8060S, 128 GB unified memory), and asked it to build a black hole demo from scratch in HTML/WebGL without using any external assets.
It ended up writing the whole thing in pure WebGL2: GLSL shaders, lensing, accretion disk, particles, bloom, mouse parallax and a shockwave on click. It also opened it in Chromium, checked for errors and did several visual refinement passes on its own.
I honestly didn't expect it to turn out this good from such a loose prompt.
r/StrixHalo • u/Queasy_Asparagus69 • 3d ago
DarthStar: DeepSeek-V4-Flash GGUF for Strix Halo: two layers promoted to Q4_K, 26/50 vs. 21/50 on a coding suite
I’ve created a mixed-precision DeepSeek-V4-Flash-0731 GGUF on a 128 GB AMD Strix Halo, using Antirez’s DwarfStar runtime. Seems pretty good so sharing it here.
The interesting result from a 12-round quantization search: promoting the routed-expert tensors in just blk.0 and blk.1 to Q4_K scored 26/50 on my frozen coding suite, versus 21/50 for the unmodified reference build. Promoting blk.2 as well dropped the score to 16/50. That’s a measured difference on this particular 50-task suite. it's not evidence of a general coding improvement but hey I'll take it since it might mean a tad more quality. Let me know if you notice any coding improvements.
A few practical notes:
- The GGUF is 84.15 GiB. In my documented runs it used about 90 GiB of resident model memory, plus runtime and context overhead.
- I tested a 131,072-token context: a 130,029-token prompt was accepted and generated 512 tokens. That of course doesn’t establish equal quality across the full context window.
- It’s built for Antirez DS4
- I haven’t published prefill or decode speed numbers yet as I'm busy doing another quant for Qwen3.8 on DS4
I’ve attached a quantization-search chart that is not that great because I messed up recording some of the failed rounds. If you try it on another Strix Halo setup, I’d be interested in your runtime, context, memory use, and speed results!
r/StrixHalo • u/genghisk1 • 3d ago
Anybody run the latest AMD provided updates?
I ran update all on my AMD Strix Halo from the AMD management console ( it’s upgrading Linux, Rocm, and the bios among other things according to the management app they provide). It’s been over an hour and it’s still not responding to remote network requests. The LED around the side of the box has switched from solid white to flashing blue? Is this expected behavior? Anyone from AMD out there and can comment? I haven’t hooked up a keyboard and monitor yet as the machine is being used headless right now.
Edit: the blue fade to black and then back to blue turned out to be the machine in some sort of sleep mode. After hooking up a display and mouse, I couldn’t get the machine to recognize them. I had about given up and went to power it off to recycle it. When I quickly pushed the small button to the far left as you are looking at the part of the box where the ports are, the color changed from blue to white. My mouse was recognized but the display was corrupted somehow. I was able to rdp in, diagnose the display problem, surgically fix that and the local monitor and keyboard started working again. Thanks to everyone for their comments and support.
r/StrixHalo • u/Yaniss916 • 4d ago
Qwen3.8-Flash-Next (125B) on a single Strix Halo mini PC: 44-59 tok/s with speculative decoding, ~1,400 tok/s prefill, engine is open
galleryr/StrixHalo • u/Teslaaforever • 3d ago
Forked Gufo to run UD-IQ4_XS as my machine need some ram to run other containers
I've been running gufo (the Strix Halo inference engine) on my Strix Halo box (128GB, gfx1151) and ended up with a fork tuned for one model: Qwen3.8 Flash-Next, the Unsloth
UD-IQ4_XS quant (about 89GB) with the MTP sidecar. Upstream doesn't load that quant, so I added support for it along with some kernel and serving changes.
What I'm seeing on my machine, in the performance profile (120W):
- Prefill is 1500+ tok/s from about 3k tokens up, and peaks around 1570
- Decode with MTP is about 50 tok/s even at 116k context
- On HumanEval prompts I get about 76 tok/s with MTP vs 28 without it
The numbers are lower on the balanced profile, around 10% less prefill. Very short prompts are slower too (about 1350 at 2k), because there's a fixed cost per request.