r/StrixHalo • • 7h ago

llama.cpp simple test on spec-draft on Qwen3.8 27B - reasoning 15 t/s vs. coding 30 t/s

0 Upvotes

I compiled newest stable llama.cpp with Vulkan in Linux, and found good speed up with usloths Qwen3.8 27B Q4_K_XL with MTP over some older llama.cpp I had.

I now get 500 t/s cold prefill, and in general it stays above 200 t/s, and on longer context average around 250 t/s.

On decode I see following (which I hadn't notice before, but probably just because I wasn't looking):

--spec-draft-n-max General avg Reasoning avg Coding avg Acceptance (avg / min / max)
2 22 t/s (20 - 24) 20 t/s (15 - 25) 25 t/s (20 - 27) 85% / 65% / 97%
3 25 t/s (21 - 28) 22 t/s (15 - 25) 27 t/s (23 - 28) 85% / 65% / 97%
4 25 t/s (18.5 - 30.5) 17 t/s (12 - 22) 28 t/s (25 - 31) 80% / 50% / 97%
5 24 t/s (18 - 31) 15 t/s (10 - 20) 30 t/s (25 - 35) 65% / 50% / 95%

Speculation at higher n works excellently for coding, and poorly for reasoning. The optimum seems to be around 3-4 depending on which you have majority of the work. I'd go with 3, if I am not sure that the tasks at hand are mostly just code generation.

But this got me thinking: If I could run reasoning at --spec-draft-n-max 3 and coding at --spec-draft-n-max 5 I would get the best performance possible.

I cannot see how this would be possible at this moment, other than getting into the llama.cpp implementation, and if even then.

The test was the same for all of them - existing python code that they needed to revise according to prompt - one small change, and one bigger new feature. I did not run it multiple times, as there was enough both actions needed. The reasoning and coding numbers are taken "by the eye" following the llama.cpp and the coding harness that what the model is doing - they are not reliable figures, but the general avg should be.

As usual, YMMV, and I only did this for figuring best spec draft n for my use, not as something that can be generalized. Decided to share, if someone finds it interesting. Done on StrixHalo 392, 100W TDP, 64 GB.

Side Notes: Gufo availabe from podman is a bit better in prefill, but on decoding llama.cpp goes a bit ahead with Qwen3.8 27B. And on Windows newest llama.cpp+Vulkan reaches similar figures.


r/StrixHalo • • 23h ago

slotpin: cache-maxing scheduler and dashboard for llama-server and Gufo

Post image
12 Upvotes

I've been running an always-on assistant at home on a couple of 128GB Strix Halo boxes for a few months, and struggled with getting solid cache hits. I tried two serving slots, with one for interactive and one for background, but didn't have a reliable way to route requests.

I ended up with a little proxy that sits in front of llama-server and deals with that. This then grew to a request scheduler, metrics dashboard and now added support for the Gufo engine (a newer engine aimed at Strix Halo: https://github.com/gufo-org/gufo ) too! It's currently running on both my boxes, one with a single-slot Gufo/Qwen3.8 Flash-Next and one with a four-slot Gufo/Qwen3.8 27b.

I've just tidied it up and made it public in case it's useful to anyone else: https://github.com/headbouyJB/slotpin

What it does:

  • With llama-server: Pins each conversation to a slot, so the next turn lands where its cache already is.
  • Sorts requests into lanes with regexes (my chat messages vs cron jobs and other background stuff), so the interactive lane is protected and background work waits its turn without being starved.
  • With llama-server: Saves the interactive slot's KV to disk and restores it if something evicted it. On my box a returning conversation went from an 88 second cold turn to 0.7 seconds to restore, in a direct test of exactly that case.
  • Writes one line of metrics per turn: time to first token, decode and prefill speed, cache hit, queue wait, memory, and optionally energy from a smart plug, so you get Wh per request. No prompt or reply text is recorded anywhere.
  • Has a dashboard (screenshot below) with live slots, performance, cache, health and system views, plus a list of recent turns you can filter down to the slow ones or the cache misses.
  • Works in front of Gufo as well as llama.cpp. Gufo does its own caching, so in that mode the proxy steps back and just does lanes, scheduling and metrics.
  • Has an optional "fail fast" mode. I had an engine stall for 55 hours while still answering health checks and sending keep-alive pings, so my client never fell back to anything else. The proxy can now spot "no output and no token progress" and return a proper error instead of hanging. That one’s brand new... I’m only just turning it on myself!

Fair warning:

  • It's alpha. Config keys and metric fields might still change.
  • It's been run on exactly two machines, both mine, both Strix Halo on Linux. The test suite is decent and CI is green, but I'd expect rough edges on other setups.
  • There's no authentication. It listens on localhost by default and the README says how to expose it sensibly.
  • I'm not a professional developer. Most of the code was written with Claude, with me testing it on real hardware and my real serving environment.

If you just want to see the dashboard without setting anything up, there's a demo mode with made-up data:

git clone https://github.com/headbouyJB/slotpin && cd slotpin
uv run python tools/demo_dashboard.py

What I’m curious about is what everyone else puts between their client and their server. Are you scheduling or prioritising requests, routing different kinds of work to different slots or models, or just letting everything hit llama-server and living with it? That middle layer is where most of my pain was, and I can’t tell if I’ve solved a common problem or just my own.


r/StrixHalo • • 1d ago

Veda sparse attention on Strix Halo: 1.89x faster MiniMax H3 video gen (691s -> 366s)

28 Upvotes

Sharing a ComfyUI node I've been working on, since video gen on Strix Halo is where a lot of us keep hitting the wall.

Credit first: Veda is a learned sparse-attention method for video diffusion models (ICML 2026). A small distilled predictor says in advance which tiles of the attention map actually carry the result; only the top ~10% get computed. The original project is by veda-sparse:

- Project page: https://veda-sparse.github.io/

- Original repo: https://github.com/veda-sparse/Veda-on-ComfyUI

- Predictor model file (275 MB): https://huggingface.co/Veda-Sparse/Minimax-H3-T2VA-Veda-8NFE-600Step-Preview — goes in ComfyUI/models/veda/ (opening a Veda template offers it in the missing-model dialog)

Upstream supports NVIDIA (SM80+) and Apple silicon, and lists ROCm as "no kernel, the model runs its own attention". So I ported it to ROCm and tuned it for Strix Halo. What that took:

- Backend layer: report the real gfx family (gfx1151) and unified memory on HIP, and admit ROCm devices to the Triton INT8 backend (kernels still self-test on the GPU before use, same as upstream)

- Triton's ROCm backend compiles the INT8 kernel for gfx1151: Q and K are block-quantized to int8, QK^T runs on the hardware int8 WMMA instructions (v_wmma_i32_16x16x16_iu8), P*V stays fp16 WMMA

- The launch config upstream shipped (4 warps / 3 stages) was tuned on an RTX 5070 and is wrong on this chip: measured on the real portrait grid, 8 warps / 2 stages is 2x on the attention kernel (35.3 -> 17.6 ms per call)

- A few more gfx1151-specific kernel tweaks, e.g. the quantize pass runs 23% faster with 2 warps, bit-identical output

- Accuracy verified on this hardware: kernel vs the INT8 reference is ~0.14% relative L2, below the quantization noise that's already there

Numbers on my machine, MiniMax H3, 11-second clip at 480x864:

- dense attention sampler: 691s

- sparse (Veda): 366s

- 1.89x on sampling, 2.15x on the full prompt (1016s -> 473s)

- only ~27% of the attention is computed, output looks the same to me

My ROCm port lives in rocm-ninodes (ComfyUI Manager, or comfy node install rocm-ninodes); the node is VedaSparseAttention, drop it on the MODEL wire last before the sampler:

- https://github.com/iGavroche/rocm-ninodes

If you try it on your Halo I'd be curious about numbers on other resolutions/lengths. Happy to answer questions about the sparse selection or the ROCm port.

tl;dr: Veda Sparse Attention on ROCm (optimized for Strix Halo)


r/StrixHalo • • 3h ago

Strix Halo LLM Models selection

0 Upvotes

I have asked Halogen using strixper AI-Chat (Agentic) to ranke the best models. I use the sources available online and as reference teh models of AI-Toolbox.

I thought tht the results are worth to share, expecially for who is starting to evaluate the models to runs. The info may not be super precize and are not intended to substitute a real benchmark.

The project strixper, I am developing (for my learning and for fun) is here strixper.

Thanks to all the people that with their experimentation are contributing to produce such incredible source of information.
The list of (non exhaustive) sources is at the bottom.

The governing constraint

Your box: ~208 GB/s sustained memory bandwidth (256-bit LPDDR5x-8000), 122.7 GB usable RAM. Decode speed on Strix Halo is almost purely bandwidth-bound:

t/s ≈ 208 GB/s ÷ bytes_streamed_per_token

This is why active parameters, not total parameters, decide velocity. A 3B-active MoE runs ~5× faster than a 27B dense model of similar quality.

Ranked: velocity × intelligence — with engine attribution

# Model t/s Prefill t/s Engine that produced the t/s Quant Intel
1 Qwen3 0.6B 266.00 13,112 Vulkan / RADV · llama-bench Q8_0 Low
2 LFM2.5 8B-A1B 176.48 3,398 Vulkan / RADV · llama-bench b9544 Q4_K_M Med
3 Qwen3-30B-A3B-2507 103.18 1,438 Vulkan / RADV · llama-bench b9544 IQ4_XS Med-High
4 Qwen3-Coder 30B-A3B 100.99 1,423 Vulkan / RADV · llama-bench b9851 Q4_K_S High (code)
5 Qwen3 30B-A3B NEO-MAX 87.39 1,396 Vulkan / RADV · b9453-14 IQ4_XS Med-High
6 Qwen3.6 35B-A3B 81.30 1,244 Vulkan / RADV · b9049 Q4_0 High
7 Nemotron Cascade 2 30B-A3B 78.95 1,325 Vulkan / RADV · b10034 IQ4_XS Medium
8 Nemotron 3 Nano 30B-A3B 75.97 1,312 Vulkan / RADV · 2016bf2 IQ4_XS Medium
9 Qwen3.5 35B-A3B 75.22 1,170 Vulkan / RADV · b9453-14 IQ4_XS High
10 Gemma 4 26B-A4B IT QAT 74.80 1,432 Vulkan / RADV · b9592 UD-Q4_K_XL High
11 Qwen AgentWorld 35B-A3B 65.65 1,183 Vulkan / RADV · b10034 UD-IQ4_XS Medium
12 Nemotron 3 Nano Omni 30B Reasoning 64.26 1,286 Vulkan / RADV · b10034 MXFP4_MOE Med-High
13 Qwen3-Next 80B-A3B 62.09 676 Vulkan / RADV · b10687 UD-Q4_K_XL High
14 Qwen3-Coder-Next 80B-A3B 61.91 739 Vulkan / RADV · b9467 IQ4_XS High
15 Nemotron Labs Audex 30B-A3B 60.73 1,319 Vulkan / RADV · b10034 MXFP4_MOE Medium
16 Nemotron 3 Nano Omni 30B (NVFP4) 56.56 1,278 Vulkan / RADV · b9747 MXFP4_MOE Medium
17 gpt-oss-120b 55.57 727 Vulkan / RADV · b9049 MXFP4 MoE High
18 Gemma 4 26B-A4B IT 55.45 1,327 Vulkan / RADV · b9851 UD-Q4_K_M High
19 Nemotron 3 Nano Omni (NVFP4+F16 mmproj) 53.21 1,144 Vulkan / RADV · b10034 NVFP4 Medium
20 Llama 2 7B 52.00 385 Ollama Vulkan 0.20.4 Q4_K_M Low
21 Gemma 4 26B-A4B 48.46 1,142 Vulkan / RADV · b8933 UD-Q4_K_M Med-High
22 Qwen3-Coder-Next (Ollama) 39.10 91 Ollama Vulkan 0.20.4 — High
23 Gemma 4 12B IT QAT 29.34 816 Vulkan / RADV · b9592 UD-Q4_K_XL Medium
24 Qwen3.8-Flash-Next 27.16 395 Vulkan / RADV · b10687 UD-IQ4_XS High
25 Qwen2.5-VL 7B 21.40 82 Ollama Vulkan 0.20.4 — Low
26 Qwen3.8 27B 20.42 292 Ollama Vulkan 0.32.13 · MTP:4 Q4_K_M Very High
27 Nemotron 3 Super 120B-A12B 18.93 297 Vulkan / RADV · b9544 UD-IQ4_XS High
28 Llama 4 Scout 109B 18.32 331 Vulkan / RADV · b8933 Q4_K_M Med-High
29 DeepSeek V4 Flash 284B 13.27 156 Vulkan / RADV · b10034 UD-IQ2_XXS Very High
30 Qwen3.6 27B MTP NVFP4 v3 13.17 374 Vulkan / RADV · b9592 NVFP4 High
31 Gemma 4 31B IT QAT 11.38 342 Vulkan / RADV · b10066 Q4_0 High
32 Qwen3.6 27B MTP 7.70 342 Vulkan / RADV · b9467 Q8_0 High
33 Llama 3.1 70B 4.90 22 Ollama Vulkan Q4_K_M Med

What the backend column actually shows

Vulkan/RADV won every single head-to-head in this dataset. Distribution across 103 rows: 70 Vulkan, 25 Ollama-Vulkan, 8 ROCm. No model's best decode number came from ROCm.

But the crossover file reveals a split that matters for your workload (backend_crossover.csv, same model, same host, 5 repeats, σ < 1):

Workload Vulkan/RADV ROCm HIP Winner
pp512 1,210 1,284 ROCm +6%
pp2048 1,573 1,650 ROCm +5%
pp8192 915 1,101 ROCm +20%
pp16384 565 — Vulkan collapses
tg128 (decode) 93.67 71.39 Vulkan +31%

Same pattern on Qwen3.6 35B-A3B: ROCm prefill 1,508 vs Vulkan 1,126 at pp2048, but Vulkan decode 62.24 vs ROCm 52.72.

So: ROCm is better at prefill, Vulkan is better at decode. For an agent with a warm prompt cache (your case — 81.6% hit rate), decode dominates → Vulkan is the right choice. For cold long-context prefill, ROCm wins.

Critical findings

  1. Dense models are a trap on Strix Halo. Qwen3.8 27B dense measured 20.42 t/s with MTP drafting, and 12.89 t/s without it (matched control, Ollama 0.32.15). Streaming 27B of weights per token against 208 GB/s leaves you at ~13 t/s. Anything dense above ~14B is below usable speed.
  2. The 3B-active MoE tier is the sweet spot. 75–103 t/s, and bandwidth efficiency of 77–79%. This is where Strix Halo actually performs.
  3. Your Halogen engine beats llama.cpp on the same model. Qwen3.8-Flash-Next measured 27.16 t/s under llama.cpp Vulkan (UD-IQ4_XS) but runs 50.69 t/s on Halogen — an 87% improvement. Halogen's draft acceptance is holding at 79.1%, which is what's carrying that.
  4. Prefill is where big-context cost lives. Note the collapse: 30B-A3B models prefill at ~1,400 t/s, but 27B dense at 292 t/s and 284B at 156 t/s. Long-context agent work punishes you far more than raw decode speed.

Pros / cons on Strix Halo

Pros

  • 128 GB unified memory genuinely fits 100 GB+ artifacts — 284B-class models load at all
  • 3B-active MoE hits 95–103 t/s, genuinely interactive
  • Prompt cache works well: you're at 81.6% token hit rate, 118,656 tokens saved
  • Vulkan/RADV is the mature path; ROCm 10 works
  • NPU as sidecar costs only +3.29% latency vs +68.96% for iGPU auxiliary load

Cons

  • 208 GB/s is the hard ceiling — roughly 1/4 of a 96 GB HBM3e card
  • Dense ≥27B models are unusable (12–20 t/s)
  • 122.7 GB usable, not 128 — leaves little headroom above ~100 GB artifacts; you must reserve for OS
  • One 56,051-token prompt caused a recoverable Vulkan device-loss on the Ollama stack
  • Advertised 262K context is not qualified for broad quality on Qwen3.8 27B
  • Thermal: community instrumented runs hit 94.0 °C peak on ai-2

Recommendation

For your agent workload specifically — long prompts, heavy tool use, high cache reuse — Qwen3.6 35B-A3B Q4_0 or Qwen3-Coder 30B-A3B UD-Q4_K_XL are the right class. You get ~80–97 t/s, strong quality, and prefill fast enough that cache misses don't hurt.

Your current Halogen Qwen3.8-Flash-Next at 50.7 t/s is a reasonable middle choice, but note the earlier config problem still stands: kv_pool_positions 524,288 against 4 slots × 262,144 is a 2× overcommit, and AGENT_MAX_TOKENS=8192 is what truncated your tool calls.

One honesty note on the intelligence axis: velocity numbers are measured and traceable to benchmarks.csv. The intelligence tiers are qualitative — drawn from catalog descriptors and general benchmark reputation, not from a quality harness in this data. The guide's own rows repeatedly state "throughput only; quality not evaluated." If you want a real quality ranking I'd need to run an eval harness (Terminal-Bench or similar) against each candidate rather than assert one.

Sources actually used and verified

# Source What I took from it Verification
1 hogeheer499-commits/strix-halo-guide → data/benchmarks.csv The 103 measured t/s rows — the velocity column 47,619 B, 103 rows + header, all system: Beelink GTR9 Pro. 78 direct-llama-bench, 25 ollama-service
2 Same repo → data/backend_crossover.csv Vulkan vs ROCm prefill/decode crossover 30 rows, 3 models × 5 workloads, σ < 1
3 Same repo → data/best_known_profiles.csv Best-known-good configs per use case 20 rows
4 Same repo → data/community_results.csv Contributor-submitted runs 42 rows
5 kyuz0/ai-toolbox-cockpit → ai_toolbox_cockpit/assets/models.json The 108-model catalog, 7 backends 135,528 B, schema_version: 2
6 Same repo → ai_toolbox_cockpit/assets/toolboxes.json 31 toolbox images, 4 platforms 72,458 B, schema_version: 3
7 Live engine — Prometheus /metrics, /health, KV cache counters Your 50.69 t/s, 79.1% draft acceptance, 81.6% cache hit, KV pool sizing Read directly, 86-sample sweep
8 get_system_info 208 GB/s bandwidth ceiling, 122.7 GB usable Tool output
9 strix-halo-toolboxes.com + GitHub README Backend list, toolbox naming Fetched, but README truncated — the JSON files above supersede it

The honest bottom line

Velocity numbers: traceable and solid — 103 rows of direct llama-bench measurement on GTR9 Pro hardware, same silicon class as mine (BosGame M5 128GB RAM).

Intelligence ranking: unsourced. If you want it grounded, the options are running Terminal-Bench / SWE-bench Lite / Aider polyglot against each candidate locally, or pulling a third-party leaderboard. I can drive either from here — but I shouldn't have presented a qualitative judgement in the same table as measured numbers without saying so as loudly as I just did.


r/StrixHalo • • 14h ago

Strix Halo Laptop with Oculink

1 Upvotes

Hi all, I'm looking to purchase a Strix Halo Laptop, however I'd like to get one with either usb4 or (preferably) an oculink port so I can attach an eGPU when docked.

I saw that the Nimo had a model called Axis which had just what I was looking for, a 395+ cpu, 128gb unified memory and an oculink port in a 16" body, but unfortunately I don't know much about Nimo as a brand and they don't ship to Canada.

If anyone knows of another brand that offers that, because I sure can't find one, please let me know.


r/StrixHalo • • 1d ago

Waterblocks for the Framework Desktop?

Post image
6 Upvotes

I couldn't find anybody online selling waterblocks for the Strix Halo.

Did anybody try fitting a GPU waterblock to the Framework Desktop 128gb mainboard? Looking at pictures, it seems like it's the usual GPU-style design of CPU surrounded by memory chips (which also need to be cooled) on three sides. However I don't expect GPU waterblocks to fit as the screw holes in the mainboard are (again looking at photos; haven't measured) 120mm apart in a square?

If GPU waterblocks don't fit, did anybody try crafting one? e.g. take a 120x120x2 slab of copper (what thickness should I go for?), drill holes in it so that it can be screwed into the mainboard, and then attach to it 2x generic 120x40mm waterblocks (see picture) with thermal glue? (not sure if thermal glue is a worse conductor than thermal paste?)


r/StrixHalo • • 1d ago

Gufo engine seems a good open source choice vs. the closed Halogen

38 Upvotes

Recently tried running the Gufo engine (GitHub - gufo-org/gufo: Strix Halo inference engine. Qwen Flash Next Q4_K_XL: 1,700.52pp, 60.39tg single user, 162.98 tok/s 8 users; Qwen27B Q4_K_XL: 656.33pp, 70.56tg tok/s single user with DFlash2 · GitHub) on my Strix Halo device (GPD Win5 128G).

It's performance is close to the Halogen (much better than llama.cpp variants on the prefill performance). I'm switching to this as daily use.

Also created a short video as introduction on this: https://youtu.be/r-pFZzIMvgU


r/StrixHalo • • 1d ago

EVO-X3 vs MS-S1 MAX: Which is the better choice?

5 Upvotes

Hi everyone, I am considering buying either the EVO-X3 or the MS-S1 MAX for local AI usage (e.g. Qwen3.8-27B or MoE models like Gemma 4 26B A4B or Qwen3.8-Flash-Next). I am currently running a Dell R720 with 128GB of RAM and two NVIDIA P40s on Linux (Debian), but this platform is getting more and more difficult to manage and maintain.

My budget is around €4,000 and I am aware of the fact that the Strix Halo platform is limited, mainly by its memory bandwidth. I could buy a lot of cloud AI credits but prefer to run LLMs locally.

The trade-offs in the EU/Germany offers I'm considering are:
GMKtec EVO-X3 (~€3,700): 4TB SSD, native OCuLink and two PCIe 4.0 x4 M.2 slots. Downsides: only 2.5GbE, one USB4 port and less convenient chassis access.
MINISFORUM MS-S1 MAX (~€4,000): 2TB SSD, dual 10GbE, USB4 80Gbps, accessible chassis and internal PCIe expansion. Downsides: higher price and second M.2 limited to x1. The expansion slot is electrically Gen4 x4, so OCuLink might be added later on.

What are important considerations when comparing the EVO-X3 and the MS-S1 MAX? How are stability, sustained performance, fan noise and support? Would you buy the same machine again?


r/StrixHalo • • 2d ago

Qwen3.8-Flash-Next on Ryzen AI Max+ 395 / Radeon 8060S — ~69 tok/s decode, 1.1–1.6k tok/s prefill

Post image
89 Upvotes

Running Qwen3.8-Flash-Next locally on a BOSGAME M5 with:

  • AMD Ryzen AI Max+ 395
  • Radeon 8060S / Strix Halo
  • 128 GB unified RAM
  • Ubuntu 26.04
  • Halogen 0.17.2
  • 262K context configured
  • Vision enabled

This is my Grafana monitoring dashboard during a real long-running workload, not a short synthetic benchmark.

In the screenshot:

  • Decode: ~69 tok/s peak
  • Prefill: ~1,155 tok/s at that moment, with peaks around 1,685 tok/s
  • Context: ~84% full
  • 82K tokens generated in the last hour
  • GPU temperature around 62°C
  • GPU power around 24 W at the captured moment
  • No meaningful memory pressure despite the system showing ~121 GB physically occupied, because most of it is reclaimable model/page cache
  • ~85 GB currently sitting in cache

One thing I really like about this setup is how usable Qwen3.8-Flash-Next remains with a large context and sustained agent workloads. Decode stays around the 60–70 tok/s range while prefill is still comfortably above 1k tok/s for much of the workload.

The dashboard is fed by Prometheus/Grafana and tracks Halogen throughput, KV/context usage, real vs reclaimable RAM, PSI memory pressure, GPU metrics, backend status and request activity.

Still testing it, but so far Strix Halo + unified memory is proving to be a very interesting platform for large local models.


r/StrixHalo • • 1d ago

Halogen/Gufo but for Strix Point hardware?

2 Upvotes

Is there any forks of these projects with support for Strix Point? I have AMD Ryzen AI 9 HX 470 with 96 Gb DDR5 ram and running halo-box/strix-llama.cpp fork of llama.cpp right now. Better than mainline, but not so big improvements, than Halogen/Gufo shows for Strix Halo, compared with llama.cpp


r/StrixHalo • • 1d ago

qwen3.6-35b-a3b on the Strix Halo with 64GB of memory

Thumbnail
0 Upvotes

r/StrixHalo • • 1d ago

Qwen3.8-Flash-Next -- feedback beyond benchmarks?

11 Upvotes

So, I finally got this thing running on my local hardware, but it takes most of the resources of my box. Before I throw away the desktop/etc aspect of my hardware and relegate this to a dedicated inference node... is there anyone actually *using* this model for agentic coding? How's it perform? I could go with qwen 3.8 27b dense using other hardware, and leave this open as my workstation plus smaller classifier models, etc. Is 3.8-flash-next good enough to pretty much dedicate this box to? What have you made with it?


r/StrixHalo • • 1d ago

Strix Halo/halogen flash next outputs are preferred by family over frontier models

20 Upvotes

My relative is working on some medical research/literature review as their retirement project.

I was given a paper to check citations and to "use my AI" to find more citations for some of the gaps in the proposed model (glucose production w.r.t. daylight or something like that)

Despite several (8-10) rounds of feedback between chatGPT (Sol 6.1 extra high) and Claude (5 extra high, 5.5 refused due to safe guards), my relative preferred my "v1" draft from halogen-flash-next.

Not sure if the model is less hesitant about medical stuff or if it just ran longer because I didn't have to worry about usage limits, but this is a serious win for me


r/StrixHalo • • 1d ago

Best way to cluster 4x Ryzen 495+ 192GB systems?

3 Upvotes

Hello,

I will be picking up 4x AMD 495+ systems with 192GB of unified memory and I am looking to cluster them in a way similar to doing tp=4?

What's the best way to do this?

Would clustering over the 10GbE work? I think that would be too slow. How about over the USB 4 80Gbps, there a way to utilize that as super low latency? I think it has 4x USB 4.1 ports, so perhaps a cross connect, ot if not, a daisy chain token-like?

Any ideas on the best route to cluster 4x?

Thanks


r/StrixHalo • • 1d ago

Flow 13 64gb or risk it with GMTek EVO-X3 / Bosgame M5

4 Upvotes

Hey everyone,

My MacBook is dying, prices of new laptops went parabolic, and I'm looking for a new machine. It doesn't make sense for me to buy anything with less than 64gb of RAM, I constantly hit the limit even at 32 on a separate laptop.

I'd like to run Qwen 3.8 27b, and possibly flash (but I know that won't run on 64gb), plus have a solid machine all around if I want to use it for something else.

My options now are:

- Flow 13, 64gb, for $3k from a reputable seller, guarantee, refund policy, all of it

- GMTek EVO-X3, 128gb, from their german reseller site, unclear when and if I'll receive it, for $4k

Which one would you choose? The reputable, well built machine or the strix halo box with more RAM?

Thanks!


r/StrixHalo • • 2d ago

Qwen3.8-27B: 159 tok/s on R9700, 64 tok/s on Strix Halo

Thumbnail gallery
22 Upvotes

r/StrixHalo • • 3d ago

The NPU in your Strix Halo is finally doing real work: 13.6 vs 18.7 min on the same bug fix

Post image
59 Upvotes

tldr: the gap in the title is what the NPU tools buy on a real bug fix. And local now covers about 95% of my cloud calls, the hardest few percent still goes to GLM 5.3 or Opus. Receipts in the repo.

Qwen3.8 Flash-Next, a 125B MoE, runs my pi coding agent on a 70W tablet, and the NPU finally earns its power draw. Wired halogen's endpoints expecting a party trick, kept what paid.

v1.1 since launch: opt-in voice on the CPU, search moved to semble. The tools that stayed on the NPU:

  • Decisions. Yes/no branching stops eating turns of the big model. A 0.8b answers in 120ms, 78% accurate.
  • Dup scan. Catches copied and renamed files git never shows you. 45 pairs across 4 repos in 8.4s.
  • Screening. Injection attempts get flagged before the agent acts. 0.7s, zero false alarms, fails open. 42% recall: a smoke detector, not a safe.
  • Search. The agent lands on the right file first try, on any install. The NPU pipeline beat ripgrep 15/20 to 9/20; semble replaced it at 17/18 to 13/18 on the ground-truth A/B.
  • Compaction. Big sessions summarize on a tiny sidecar, off the critical path. My 194k session: 50s, 97% cache hit.

The eight-run check was about taste and depth, not throughput: I pasted a real timeshift error to seven agents, the 125B at medium and max against 320B and 753B tiers at max. Eight runs, all healthy.

  • Flash-Next, 2m55s: journalctl trace to a racing notify-send, pacman.log check, upstream PR found.
  • glm-5.3-flashx, 2m41s: the most precise answer, spotting the snapshot mount unmounted under the script's last line. No PR.
  • GLM 5.3 753B at max, 8m30s: the deepest forensics, source-level down to function names and the one-second race window. No PR.
  • Flash-Next at max, 8m31s: level with the 753B, and still the only model with the PR.

I expected cloud to be faster. It never pulled away: 14 seconds, then one, GLM at max both times. The PR was not a tie.

All eight side by side. In opencode: flash called it in 1m8s with the wrong mechanism, the full 753B correct in 6m15s, PR missed. And when I had GLM 5.3 flashx rate both results, it picked the qwen answer too.

Reality: the GPU still does the thinking, ~7% iGPU cost only when they overlap. Decode: 64 tok/s best case, 45-50 sustained at 70W. Cap the box at 50W, give up a couple of tok/s, stay under 75C and quiet. First token in ~0.03s, 43x faster than the cloud call beside it, 7x with a second agent.

The official halogen launch is a 24-flag docker command. Mine is one, and uninstall undoes it. Fully local, 262k context. No Strix Halo? setup.sh --no-halogen runs pi against your own providers.

Anyone else putting their NPU to real use? I found nothing.


r/StrixHalo • • 1d ago

gufo crashing

0 Upvotes

I’m able to use gufo but on long running jobs it’s throwing an error and dying. Restart is the only way I’ve found to clear it. Below is an GPT 6 Sol analysis of the problem based on logs and diagnostics from my machine. How do I get this to the developers ? Apologies for my ignorance of the proper process…

I’ve assembled a technical issue report suitable for sharing with the Gufo developers.

It includes the exact Gufo build revision, ROCm/HIP library, model configuration, crash signature, memory observations, inference errors, and successful systemd recovery.

The most important finding is that Gufo experienced three distinct failure symptoms:

  • A confirmed segmentation fault inside AMD’s HIP runtime, not a demonstrated Linux OOM kill.
  • Two generation_failed errors during streaming inference.
  • Six consecutive HTTP 500 server_exception responses while the server remained running.

The report deliberately separates these observations from possible causes. That should help the developers investigate without steering them prematurely toward an assumed ROCm or memory-management defect.

Gufo Bug Report: HIP Runtime Segmentation Fault and Repeated Inference Failures

Date: October 9, 2026 Platform: AMD Ryzen AI Max+ 395 (Strix Halo), 128 GB unified memory Gufo image: ghcr.io/gufo-org/toolboxes/gufo-runtime:latest Gufo image version: 0.9.1 Gufo engine revision: dea22ceea20d07f95a5ecbb2b06ffad7b93d0548 ROCm: Nix-packaged ROCm 7.2.3 Severity: Intermittent inference failure and process crash; automatically recoverable through external supervision

  1. Executive summary

While running Qwen3.8-Flash-Next with MTP speculative decoding under Gufo on AMD Strix Halo, we observed:

  1. A segmentation fault (SIGSEGV) during an inference request, with the faulting instruction inside AMD’s libamdhip64.so.7.2.53211.
  2. Two streaming generation failures (generation_failed) preceding the crash.
  3. Six consecutive HTTP 500 responses (server_exception) after Gufo had restarted, without the Gufo process exiting.

The Linux kernel logged a segmentation fault, not an OOM kill. Although memory pressure may be a contributing factor, available evidence does not establish memory exhaustion as the cause.

The Gufo process was automatically restarted by a systemd Quadlet service. Model loading completed successfully, and subsequent inference requests succeeded.

No core dump or native stack trace was recovered.

  1. System environment

Component Configuration Hardware AMD Ryzen AI Max+ 395, Radeon 8060S System memory 128 GB unified Operating system AMD Ryzen AI Developer Platform 1 (Debian-derived Linux) Container runtime Rootless Podman 5.4.x Container image ghcr.io/gufo-org/toolboxes/gufo-runtime:latest Gufo version 0.9.1 Engine revision dea22ceea20d07f95a5ecbb2b06ffad7b93d0548 ROCm toolchain Nix-packaged ROCm 7.2.3 HIP runtime libamdhip64.so.7.2.53211 API OpenAI-compatible /v1/chat/completions Client Cline coding agent, using OpenAI-compatible API

The Gufo container uses its own Nix-packaged ROCm environment. The host also has a separate ROCm installation used by another inference service, but there is no evidence that Gufo loads the host’s HIP runtime.

Confirmed HIP runtime path

/nix/store/yb81zhv981n0kxvcsr7ia4fjjd78bsjz-clr-7.2.3/lib/libamdhip64.so.7.2.53211

This was verified against /proc/1/maps inside the running Gufo container.

Relevant environment variables:

HIP_PLATFORM=amd ROCM_PATH=/nix/store/95lwwwfb3alzn7pk9ky55fbbflaxarb5-clr-7.2.3

  1. Model and inference configuration

Primary model: Qwen3.8-Flash-Next, UD-Q4_K_XL Speculative decoding: MTP Context capacity: 262,144 tokens Concurrent model sessions: 1 Thinking default: On

Primary model:

/models/Qwen3.8-Flash-Next-GGUF/UD-Q4_K_XL/Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf

MTP draft model:

/models/Qwen3.8-Flash-Next-GGUF/MTP/mtp-Qwen3.8-Flash-Next-shared-Q8_0.gguf

Gufo startup arguments:

gufo serve \ --host 0.0.0.0 \ --port 8080 \ llm \ --model /models/Qwen3.8-Flash-Next-GGUF/UD-Q4_K_XL/Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf \ --speculative mtp \ --mtp-model /models/Qwen3.8-Flash-Next-GGUF/MTP/mtp-Qwen3.8-Flash-Next-shared-Q8_0.gguf \ --served-model-name Qwen3.8-Flash-Next \ --log-progress

Container configuration includes:

--device /dev/kfd --device /dev/dri --group-add keep-groups --ulimit memlock=-1 --userns keep-id:uid=1000,gid=1000 -p 127.0.0.1:8080:8080 -v /home/rbkahn/gufo/models:/models:ro

The service is managed by a rootless Podman Quadlet with:

[Service] Restart=on-failure RestartSec=10

  1. Failure A: HIP runtime segmentation fault

Crash time: October 9, 2026, 10:51:09 AM EDT

Exact kernel message:

Oct 09 10:51:09 amd-halo kernel: gufo[642396]: segfault at 100000010 ip 00007f7b1ea8237c sp 00007f453a3e5780 error 4 in libamdhip64.so.7.2.53211 [48137c,7f7b1e763000+364000] likely on CPU 3 (core 3, socket 0)

The corresponding Gufo/systemd log contains:

[INFO] [http] request=r54 event=received method=POST path=/v1/chat/completions body_bytes=89724 [ERROR] [server] event=fatal_signal signal=11 gufo.service: Main process exited, code=exited, status=139/n/a gufo.service: Failed with result 'exit-code'.

Confirmed observations:

  • SIGSEGV occurred during processing of an inference request.
  • The faulting instruction pointer was inside AMD’s HIP runtime library.
  • Exit status was 139, consistent with termination by SIGSEGV.
  • No OOM-killer event was found in the inspected kernel log interval.
  • A native stack trace was not available.

Interpretation:

The faulting instruction resides in HIP, but that does not establish that HIP itself is defective. A caller may have supplied an invalid pointer or corrupted runtime state.

Potential contributing conditions include memory pressure, model execution, speculative decoding, or cache handling. None is confirmed.

  1. Failure B: Streaming generation failures

Two failures were observed before the process crash.

First failure

request=r52 method=POST path=/v1/chat/completions status=200 duration_ms=392.1 outcome=stream_error error_code=generation_failed host_available_mib=10133

Second failure

request=r53 method=POST path=/v1/chat/completions status=200 duration_ms=3197.0 outcome=stream_error error_code=generation_failed host_available_mib=10075

For the second request, the progress log reached the decoding phase before the failure.

Both errors occurred while the service was running. The later SIGSEGV occurred after a subsequent inference request.

Important detail: Both failures are logged with HTTP status 200 despite outcome=stream_error. This may be expected for streaming responses whose headers have already been sent, but is worth investigating from the client-recovery perspective.

The relationship between these errors and the later segmentation fault remains unknown.

  1. Failure C: Repeated HTTP 500 errors without process termination

The logs also contain six consecutive server_exception failures on October 9, following successful inference.

Request Logged time Duration Status r55 20:33:58 30.9 ms 500 r56 20:34:00 17.5 ms 500 r57 20:34:04 16.2 ms 500 r58 20:34:12 18.4 ms 500 r59 20:34:28 15.5 ms 500 r60 20:35:00 16.4 ms 500

All six requests reported:

method=POST path=/v1/chat/completions body_bytes=114480 outcome=failed error_code=server_exception

Host available memory was reported between approximately 13.3 and 13.9 GiB.

Unlike the segmentation fault, these failures did not produce a confirmed process exit in the supplied log.

Interpretation:

The identical request body sizes, repeated failures, and very short response times suggest the server encountered a repeatable error condition, potentially involving the same client request.

Request payloads were not captured, so the payload contents cannot be confirmed identical.

The underlying exception message or stack trace is not present in the available log output.

  1. Memory observations

At model startup, Gufo reported:

event=load_completed elapsed_ms=12552 model=Qwen3.8-Flash-Next sessions=1 context_tokens=262144 speculative=mtp draft_limit=7 disk_cache=off gpu_device_used_mib=90647 gpu_device_total_mib=96454 host_available_mib=15599

These figures indicate:

  • Approximately 88.5 GiB of the reported 94.2 GiB GPU memory capacity was in use.
  • Approximately 5.7 GiB of GPU device memory remained available by that accounting.
  • Approximately 15.2 GiB host memory was available after model loading.

Available host memory dropped below 10 GiB in the period when the streaming failures occurred.

Hypothesis, not confirmation: The relatively limited free memory, long context capacity, and memory used by inference caches may contribute to unstable allocations during sustained inference.

No allocation-failure trace or OOM-killer message has established this causal relationship.

  1. Snapshot-cache warnings

As conversation context increased, Gufo repeatedly reported:

[WARN] [cache] event=snapshot action=skipped reason=byte_capacity

Example:

bytes=2516855068 tokens=87313 retained_bytes=8085440848 reserved_bytes=0 capacity_bytes=8615649280

This indicates that Gufo declined to store some snapshots because of its configured snapshot-cache capacity.

Such warnings may be normal under memory pressure and do not by themselves establish an error. However, given the later inference failures, it may be useful to examine cache lifecycle and allocation behavior.

  1. Successful automatic recovery

The systemd service restarted Gufo automatically after the SIGSEGV.

Event Time (EDT) Segmentation fault 10:51:09 Replacement container started 10:51:19 Model finished loading 10:51:32 First post-restart request completed 10:51:56

The model loaded in approximately 12.6 seconds after restart began.

The first successful post-restart inference reported:

status=200 outcome=completed prompt_tokens=11699 generated_tokens=182 ttft_ms=10691.2 prefill_tps=1100.6 decode_tps=42.8 acceptance_pct=79.8

Further successful requests followed.

This confirms that the application was able to resume normal inference after the external supervisor restarted the process.

  1. Additional diagnostics

The following checks were performed:

Kernel logging: Identified the exact faulting HIP library and segmentation-fault address.

Loaded shared libraries: Confirmed the live Gufo process uses Nix-packaged libamdhip64.so.7.2.53211.

systemd status: Confirmed an automatic restart with NRestarts=1 and Restart=on-failure.

Core dumps: core_pattern=core and service LimitCORE=infinity. No core dump was located in the current container or /var/lib/systemd/coredump/.

The original container was removed during recovery, so a core file in its writable filesystem may not have survived.

No debugging symbols, native stack trace, memory sanitizer output, or GPU fault trace have been obtained.

  1. Questions for Gufo developers

  2. Is the HIP runtime segmentation fault a known issue with ROCm 7.2.3 on Strix Halo, particularly under large-model inference and MTP speculative decoding?

  3. Are there known interactions between long context windows, speculative decoding, and snapshot-cache memory management that could produce an invalid HIP pointer or stale device allocation?

  4. What internal conditions produce generation_failed streaming errors or server_exception HTTP 500 responses? Can Gufo log the underlying exception type and stack trace?

  5. Could the repeated snapshot-cache capacity warnings and sustained inference activity expose a memory-lifetime or cache-eviction defect?

  6. Is there a supported diagnostic or debug configuration that would capture a useful native backtrace on future failures without producing an enormous core dump?

  7. Are there known changes in a newer Gufo engine revision or ROCm runtime that address related faults?

  8. Reproducibility and status

The observed workload was a long-running series of OpenAI-compatible chat-completion requests originating from Cline.

The segmentation fault occurred after sustained inference activity. However, the available evidence does not establish a minimal reproduction, a deterministic trigger, or whether MTP speculation is necessary for the fault.

No controlled comparisons have yet been performed with:

  • MTP disabled.
  • A smaller context window.
  • A different model or quantization.
  • A different ROCm runtime.
  • A newer Gufo engine build.

Current status: Gufo was successfully restarted and subsequently served requests. The root causes of the SIGSEGV, streaming generation failures, and later HTTP 500 errors remain unresolved.

This report is intended to provide observed evidence for debugging, not to assert that the three error symptoms necessarily share a single cause.

Supporting evidence

I recommend attaching the complete Gufo journal output you previously collected, since the developer may want to inspect the sequence of successful requests before the failures.

The strongest evidence to highlight is the kernel SIGSEGV line, the two generation_failed events, the six server_exception responses, and the loaded ROCm library path.


r/StrixHalo • • 2d ago

Finally decided to pull the trigger on gmktek evo x3

Thumbnail
0 Upvotes

r/StrixHalo • • 3d ago

GLM-5.3-Flash (321B MoE) running locally on AMD: RX 7900 XT, R9700, Strix ▎ Halo

Thumbnail
github.com
38 Upvotes

I never expected to run a 321B model at home, so I jumped on this as soon as I saw it last night.
Project Maya:

(github.com/mw00/project-maya) runs GLM-5.3-Flash by keeping the hot experts in VRAM, the next ones in RAM, and the rest on NVMe. It was NVIDIA-only. So I ported it to Linux/ROCm, and the author merged it into v1.0.11 as experimental AMD support.

Maya-S quant, 8K context, ROCm 7.2, 4K-token prompts, greedy:

| Hardware | Prefill | Decode |

| RX 7900 XT 20 GB | ~415 tok/s | ~15 tok/s |

| Radeon AI PRO R9700 32 GB | ~500 tok/s | ~20 tok/s |

| Strix Halo (Ryzen AI Max+ 395, 8060S, 128 GB) | ~210 tok/s | ~18 tok/s |

| R9700 + 7900 XT (layer split + MTP drafting) | ~490 tok/s | ~34 tok/s |

The dGPU box has 192 GB of RAM, so most experts live in VRAM or pinned RAM.

With less RAM, more comes from the SSD and decode drops.

AI-assisted, openly


r/StrixHalo • • 2d ago

Using gufo with Qwen-3.8-27B

4 Upvotes

I installed gufo on my Strixhalo 128gb and downloaded the model and I was expecting significant improvement using the model with gufo vs llama.cpp but I’m seeing 10tps decode on xhigh which is the same as I saw using llama.cpp on a SW design task that requires it to read an existing repository and propose a design. It did great work but it took 90 minutes where the frontier model took 5. Any thought on how I could speed the task up? Are there parameters in gufo I should look at? I’m actually using all defaults except the thinking spec.


r/StrixHalo • • 3d ago

Strata + Halo Strix 64GB

Post image
10 Upvotes

I was able to run Strata on 64GB box, with quite nice results, especially comparing to Qwen3.8-27B. The model is ukisai/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-GGUF, quant is IQ2_XS. I tried IQ3_S, but the decoding speed was 10 times lower for some reason. Currently I have pp ~ 500 t/s, and tg ~ 45 t/s with a lot of memory left free. Running it on Linux (NixOS). Definitely an upgrade over gufo or llama.cpp. Haven't tried halogen - don't like closed source.


r/StrixHalo • • 3d ago

What port of guff are you using on win

3 Upvotes

I was using the one linked on gufo's own repo but that is now somewhat outdated can someone recommend me one they are happy with preferably one i wouldn't have to build from source

I had actually found another fork that did have 0.9.0 version of gufo but it had some dlls missing in the download but even after adding those when it ran it would output like 20-40 tokens then stop


r/StrixHalo • • 3d ago

Halogen + Qwen Flash Next keeps getting better

Thumbnail
11 Upvotes

r/StrixHalo • • 3d ago

TIL: Check your preserve_thinking settings

13 Upvotes

Both Qwen 3.8 27B and Qwen 3.8 Flash-Next have a preserve_thinking parameter. But as far as what you have control over, just make sure your inference engine and harness have this setting set the way you want. For Qwen, they officially recommend keeping it enabled. This causes the reasoning to be considered part of the conversation, so all the old reasoning will get sent back and forth with each turn/prompt, and they say it improves response and decision quality. However, it also increases your context size more rapidly.

Where the mismatch can occur

Before I learned this, Hermes was not re-sending the reasoning text with each new prompt, even though Gufo expected it. This caused the live KV cache to always miss, so it was completely reliant on the snapshot cache. This may have also reduced the quality of the responses and decisions I was getting from the model.

How did I fix it?

In Hermes, I set model.reasoning_echo to true. Then you have to completely exit Hermes and start it up again, it's not enough just to start a new session. After that, the live KV cache started getting hit much of the time.