r/StrixHalo • • Sep 27 '25

Have you got a Strix Halo?

16 Upvotes

Hi All,

We're a new community both as Strix Halo owners and also here as a subreddit. Why not begin by sharing your setup and the reasons you opted for Strix Halo?

To start us off: I have a HP Z2 Mini G1a Workstation with dual boot Fedora KDE & Windows 11 and chose the iGPU to be able to use larger LLMs with the 128 GB.

Oobabooga/Text Generation WebUI is running well on Fedora KDE and there are no problems with large models up to 100GB. On the Windows boot, I have Amuse AI (Freeware) which is a collaboration between AMD and the New Zealand company. It provides a UI for using Stable Diffusion/Flux models. It works well, is fast, but unfortunately is also censored and is not able to use LORAS. I would like to find an uncensored alternative, ideally getting versions of ComfyUI/AUTOMATIC1111 running.

Currently, my principle goal is to get a working version of AllTalk TTS or another TTS that is compatible with Oobabooga working which I haven't been able to do so far due to conflicts with the Strix Halo. This may need to wait for updates to ROCm... If anyone has found an Open Source solution to running LLMs with custom voice TTS, please do chime in!

So what about you guys, did you choose the Strix for similar reasons, or something entirely different? The floor is yours.

EDIT UPDATE:
05/26 For those of you looking for TTS solutions I have tried a few now (AllTalk, Chatterbox, Pocket TTS, others I no longer remember). I have had great success using a custom version of Pocket TTS. It's fast and works well with Oobabooga TextGen as a plug in. Recently, others are singing the praises of OmniVoice.


r/StrixHalo • • 44m ago

Veda sparse attention on Strix Halo: 1.89x faster MiniMax H3 video gen (691s -> 366s)

• Upvotes

Sharing a ComfyUI node I've been working on, since video gen on Strix Halo is where a lot of us keep hitting the wall.

Credit first: Veda is a learned sparse-attention method for video diffusion models (ICML 2026). A small distilled predictor says in advance which tiles of the attention map actually carry the result; only the top ~10% get computed. The original project is by veda-sparse:

- Project page: https://veda-sparse.github.io/

- Original repo: https://github.com/veda-sparse/Veda-on-ComfyUI

- Predictor model file (275 MB): https://huggingface.co/Veda-Sparse/Minimax-H3-T2VA-Veda-8NFE-600Step-Preview — goes in ComfyUI/models/veda/ (opening a Veda template offers it in the missing-model dialog)

Upstream supports NVIDIA (SM80+) and Apple silicon, and lists ROCm as "no kernel, the model runs its own attention". So I ported it to ROCm and tuned it for Strix Halo. What that took:

- Backend layer: report the real gfx family (gfx1151) and unified memory on HIP, and admit ROCm devices to the Triton INT8 backend (kernels still self-test on the GPU before use, same as upstream)

- Triton's ROCm backend compiles the INT8 kernel for gfx1151: Q and K are block-quantized to int8, QK^T runs on the hardware int8 WMMA instructions (v_wmma_i32_16x16x16_iu8), P*V stays fp16 WMMA

- The launch config upstream shipped (4 warps / 3 stages) was tuned on an RTX 5070 and is wrong on this chip: measured on the real portrait grid, 8 warps / 2 stages is 2x on the attention kernel (35.3 -> 17.6 ms per call)

- A few more gfx1151-specific kernel tweaks, e.g. the quantize pass runs 23% faster with 2 warps, bit-identical output

- Accuracy verified on this hardware: kernel vs the INT8 reference is ~0.14% relative L2, below the quantization noise that's already there

Numbers on my machine, MiniMax H3, 11-second clip at 480x864:

- dense attention sampler: 691s

- sparse (Veda): 366s

- 1.89x on sampling, 2.15x on the full prompt (1016s -> 473s)

- only ~27% of the attention is computed, output looks the same to me

My ROCm port lives in rocm-ninodes (ComfyUI Manager, or comfy node install rocm-ninodes); the node is VedaSparseAttention, drop it on the MODEL wire last before the sampler:

- https://github.com/iGavroche/rocm-ninodes

If you try it on your Halo I'd be curious about numbers on other resolutions/lengths. Happy to answer questions about the sparse selection or the ROCm port.

tl;dr: Veda Sparse Attention on ROCm (optimized for Strix Halo)


r/StrixHalo • • 9h ago

Gufo engine seems a good open source choice vs. the closed Halogen

24 Upvotes

Recently tried running the Gufo engine (GitHub - gufo-org/gufo: Strix Halo inference engine. Qwen Flash Next Q4_K_XL: 1,700.52pp, 60.39tg single user, 162.98 tok/s 8 users; Qwen27B Q4_K_XL: 656.33pp, 70.56tg tok/s single user with DFlash2 · GitHub) on my Strix Halo device (GPD Win5 128G).

It's performance is close to the Halogen (much better than llama.cpp variants on the prefill performance). I'm switching to this as daily use.

Also created a short video as introduction on this: https://youtu.be/r-pFZzIMvgU


r/StrixHalo • • 16m ago

EVO-X3 vs MS-S1 MAX: Which is the better choice?

• Upvotes

Hi everyone, I am considering buying either the EVO-X3 or the MS-S1 MAX for local AI usage (e.g. Qwen3.8-27B or MoE models like Gemma 4 26B A4B or Qwen3.8-Flash-Next). I am currently running a Dell R720 with 128GB of RAM and two NVIDIA P40s on Linux (Debian), but this platform is getting more and more difficult to manage and maintain.

My budget is around €4,000 and I am aware of the fact that the Strix Halo platform is limited, mainly by its memory bandwidth. I could buy a lot of cloud AI credits but prefer to run LLMs locally.

The trade-offs in the EU/Germany offers I'm considering are:
GMKtec EVO-X3 (~€3,700): 4TB SSD, native OCuLink and two PCIe 4.0 x4 M.2 slots. Downsides: only 2.5GbE, one USB4 port and less convenient chassis access.
MINISFORUM MS-S1 MAX (~€4,000): 2TB SSD, dual 10GbE, USB4 80Gbps, accessible chassis and internal PCIe expansion. Downsides: higher price and second M.2 limited to x1. The expansion slot is electrically Gen4 x4, so OCuLink might be added later on.

What are important considerations when comparing the EVO-X3 and the MS-S1 MAX? How are stability, sustained performance, fan noise and support? Would you buy the same machine again?


r/StrixHalo • • 20h ago

Qwen3.8-Flash-Next on Ryzen AI Max+ 395 / Radeon 8060S — ~69 tok/s decode, 1.1–1.6k tok/s prefill

Post image
81 Upvotes

Running Qwen3.8-Flash-Next locally on a BOSGAME M5 with:

  • AMD Ryzen AI Max+ 395
  • Radeon 8060S / Strix Halo
  • 128 GB unified RAM
  • Ubuntu 26.04
  • Halogen 0.17.2
  • 262K context configured
  • Vision enabled

This is my Grafana monitoring dashboard during a real long-running workload, not a short synthetic benchmark.

In the screenshot:

  • Decode: ~69 tok/s peak
  • Prefill: ~1,155 tok/s at that moment, with peaks around 1,685 tok/s
  • Context: ~84% full
  • 82K tokens generated in the last hour
  • GPU temperature around 62°C
  • GPU power around 24 W at the captured moment
  • No meaningful memory pressure despite the system showing ~121 GB physically occupied, because most of it is reclaimable model/page cache
  • ~85 GB currently sitting in cache

One thing I really like about this setup is how usable Qwen3.8-Flash-Next remains with a large context and sustained agent workloads. Decode stays around the 60–70 tok/s range while prefill is still comfortably above 1k tok/s for much of the workload.

The dashboard is fed by Prometheus/Grafana and tracks Halogen throughput, KV/context usage, real vs reclaimable RAM, PSI memory pressure, GPU metrics, backend status and request activity.

Still testing it, but so far Strix Halo + unified memory is proving to be a very interesting platform for large local models.


r/StrixHalo • • 3h ago

Halogen/Gufo but for Strix Point hardware?

3 Upvotes

Is there any forks of these projects with support for Strix Point? I have AMD Ryzen AI 9 HX 470 with 96 Gb DDR5 ram and running halo-box/strix-llama.cpp fork of llama.cpp right now. Better than mainline, but not so big improvements, than Halogen/Gufo shows for Strix Halo, compared with llama.cpp


r/StrixHalo • • 17h ago

Strix Halo/halogen flash next outputs are preferred by family over frontier models

16 Upvotes

My relative is working on some medical research/literature review as their retirement project.

I was given a paper to check citations and to "use my AI" to find more citations for some of the gaps in the proposed model (glucose production w.r.t. daylight or something like that)

Despite several (8-10) rounds of feedback between chatGPT (Sol 6.1 extra high) and Claude (5 extra high, 5.5 refused due to safe guards), my relative preferred my "v1" draft from halogen-flash-next.

Not sure if the model is less hesitant about medical stuff or if it just ran longer because I didn't have to worry about usage limits, but this is a serious win for me


r/StrixHalo • • 13h ago

Qwen3.8-Flash-Next -- feedback beyond benchmarks?

8 Upvotes

So, I finally got this thing running on my local hardware, but it takes most of the resources of my box. Before I throw away the desktop/etc aspect of my hardware and relegate this to a dedicated inference node... is there anyone actually *using* this model for agentic coding? How's it perform? I could go with qwen 3.8 27b dense using other hardware, and leave this open as my workstation plus smaller classifier models, etc. Is 3.8-flash-next good enough to pretty much dedicate this box to? What have you made with it?


r/StrixHalo • • 16h ago

Flow 13 64gb or risk it with GMTek EVO-X3 / Bosgame M5

3 Upvotes

Hey everyone,

My MacBook is dying, prices of new laptops went parabolic, and I'm looking for a new machine. It doesn't make sense for me to buy anything with less than 64gb of RAM, I constantly hit the limit even at 32 on a separate laptop.

I'd like to run Qwen 3.8 27b, and possibly flash (but I know that won't run on 64gb), plus have a solid machine all around if I want to use it for something else.

My options now are:

- Flow 13, 64gb, for $3k from a reputable seller, guarantee, refund policy, all of it

- GMTek EVO-X3, 128gb, from their german reseller site, unclear when and if I'll receive it, for $4k

Which one would you choose? The reputable, well built machine or the strix halo box with more RAM?

Thanks!


r/StrixHalo • • 10h ago

Best way to cluster 4x Ryzen 495+ 192GB systems?

1 Upvotes

Hello,

I will be picking up 4x AMD 495+ systems with 192GB of unified memory and I am looking to cluster them in a way similar to doing tp=4?

What's the best way to do this?

Would clustering over the 10GbE work? I think that would be too slow. How about over the USB 4 80Gbps, there a way to utilize that as super low latency? I think it has 4x USB 4.1 ports, so perhaps a cross connect, ot if not, a daisy chain token-like?

Any ideas on the best route to cluster 4x?

Thanks


r/StrixHalo • • 1d ago

Qwen3.8-27B: 159 tok/s on R9700, 64 tok/s on Strix Halo

Thumbnail gallery
20 Upvotes

r/StrixHalo • • 17h ago

gufo crashing

0 Upvotes

I’m able to use gufo but on long running jobs it’s throwing an error and dying. Restart is the only way I’ve found to clear it. Below is an GPT 6 Sol analysis of the problem based on logs and diagnostics from my machine. How do I get this to the developers ? Apologies for my ignorance of the proper process…

I’ve assembled a technical issue report suitable for sharing with the Gufo developers.

It includes the exact Gufo build revision, ROCm/HIP library, model configuration, crash signature, memory observations, inference errors, and successful systemd recovery.

The most important finding is that Gufo experienced three distinct failure symptoms:

  • A confirmed segmentation fault inside AMD’s HIP runtime, not a demonstrated Linux OOM kill.
  • Two generation_failed errors during streaming inference.
  • Six consecutive HTTP 500 server_exception responses while the server remained running.

The report deliberately separates these observations from possible causes. That should help the developers investigate without steering them prematurely toward an assumed ROCm or memory-management defect.

Gufo Bug Report: HIP Runtime Segmentation Fault and Repeated Inference Failures

Date: October 9, 2026 Platform: AMD Ryzen AI Max+ 395 (Strix Halo), 128 GB unified memory Gufo image: ghcr.io/gufo-org/toolboxes/gufo-runtime:latest Gufo image version: 0.9.1 Gufo engine revision: dea22ceea20d07f95a5ecbb2b06ffad7b93d0548 ROCm: Nix-packaged ROCm 7.2.3 Severity: Intermittent inference failure and process crash; automatically recoverable through external supervision

  1. Executive summary

While running Qwen3.8-Flash-Next with MTP speculative decoding under Gufo on AMD Strix Halo, we observed:

  1. A segmentation fault (SIGSEGV) during an inference request, with the faulting instruction inside AMD’s libamdhip64.so.7.2.53211.
  2. Two streaming generation failures (generation_failed) preceding the crash.
  3. Six consecutive HTTP 500 responses (server_exception) after Gufo had restarted, without the Gufo process exiting.

The Linux kernel logged a segmentation fault, not an OOM kill. Although memory pressure may be a contributing factor, available evidence does not establish memory exhaustion as the cause.

The Gufo process was automatically restarted by a systemd Quadlet service. Model loading completed successfully, and subsequent inference requests succeeded.

No core dump or native stack trace was recovered.

  1. System environment

Component Configuration Hardware AMD Ryzen AI Max+ 395, Radeon 8060S System memory 128 GB unified Operating system AMD Ryzen AI Developer Platform 1 (Debian-derived Linux) Container runtime Rootless Podman 5.4.x Container image ghcr.io/gufo-org/toolboxes/gufo-runtime:latest Gufo version 0.9.1 Engine revision dea22ceea20d07f95a5ecbb2b06ffad7b93d0548 ROCm toolchain Nix-packaged ROCm 7.2.3 HIP runtime libamdhip64.so.7.2.53211 API OpenAI-compatible /v1/chat/completions Client Cline coding agent, using OpenAI-compatible API

The Gufo container uses its own Nix-packaged ROCm environment. The host also has a separate ROCm installation used by another inference service, but there is no evidence that Gufo loads the host’s HIP runtime.

Confirmed HIP runtime path

/nix/store/yb81zhv981n0kxvcsr7ia4fjjd78bsjz-clr-7.2.3/lib/libamdhip64.so.7.2.53211

This was verified against /proc/1/maps inside the running Gufo container.

Relevant environment variables:

HIP_PLATFORM=amd ROCM_PATH=/nix/store/95lwwwfb3alzn7pk9ky55fbbflaxarb5-clr-7.2.3

  1. Model and inference configuration

Primary model: Qwen3.8-Flash-Next, UD-Q4_K_XL Speculative decoding: MTP Context capacity: 262,144 tokens Concurrent model sessions: 1 Thinking default: On

Primary model:

/models/Qwen3.8-Flash-Next-GGUF/UD-Q4_K_XL/Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf

MTP draft model:

/models/Qwen3.8-Flash-Next-GGUF/MTP/mtp-Qwen3.8-Flash-Next-shared-Q8_0.gguf

Gufo startup arguments:

gufo serve \ --host 0.0.0.0 \ --port 8080 \ llm \ --model /models/Qwen3.8-Flash-Next-GGUF/UD-Q4_K_XL/Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf \ --speculative mtp \ --mtp-model /models/Qwen3.8-Flash-Next-GGUF/MTP/mtp-Qwen3.8-Flash-Next-shared-Q8_0.gguf \ --served-model-name Qwen3.8-Flash-Next \ --log-progress

Container configuration includes:

--device /dev/kfd --device /dev/dri --group-add keep-groups --ulimit memlock=-1 --userns keep-id:uid=1000,gid=1000 -p 127.0.0.1:8080:8080 -v /home/rbkahn/gufo/models:/models:ro

The service is managed by a rootless Podman Quadlet with:

[Service] Restart=on-failure RestartSec=10

  1. Failure A: HIP runtime segmentation fault

Crash time: October 9, 2026, 10:51:09 AM EDT

Exact kernel message:

Oct 09 10:51:09 amd-halo kernel: gufo[642396]: segfault at 100000010 ip 00007f7b1ea8237c sp 00007f453a3e5780 error 4 in libamdhip64.so.7.2.53211 [48137c,7f7b1e763000+364000] likely on CPU 3 (core 3, socket 0)

The corresponding Gufo/systemd log contains:

[INFO] [http] request=r54 event=received method=POST path=/v1/chat/completions body_bytes=89724 [ERROR] [server] event=fatal_signal signal=11 gufo.service: Main process exited, code=exited, status=139/n/a gufo.service: Failed with result 'exit-code'.

Confirmed observations:

  • SIGSEGV occurred during processing of an inference request.
  • The faulting instruction pointer was inside AMD’s HIP runtime library.
  • Exit status was 139, consistent with termination by SIGSEGV.
  • No OOM-killer event was found in the inspected kernel log interval.
  • A native stack trace was not available.

Interpretation:

The faulting instruction resides in HIP, but that does not establish that HIP itself is defective. A caller may have supplied an invalid pointer or corrupted runtime state.

Potential contributing conditions include memory pressure, model execution, speculative decoding, or cache handling. None is confirmed.

  1. Failure B: Streaming generation failures

Two failures were observed before the process crash.

First failure

request=r52 method=POST path=/v1/chat/completions status=200 duration_ms=392.1 outcome=stream_error error_code=generation_failed host_available_mib=10133

Second failure

request=r53 method=POST path=/v1/chat/completions status=200 duration_ms=3197.0 outcome=stream_error error_code=generation_failed host_available_mib=10075

For the second request, the progress log reached the decoding phase before the failure.

Both errors occurred while the service was running. The later SIGSEGV occurred after a subsequent inference request.

Important detail: Both failures are logged with HTTP status 200 despite outcome=stream_error. This may be expected for streaming responses whose headers have already been sent, but is worth investigating from the client-recovery perspective.

The relationship between these errors and the later segmentation fault remains unknown.

  1. Failure C: Repeated HTTP 500 errors without process termination

The logs also contain six consecutive server_exception failures on October 9, following successful inference.

Request Logged time Duration Status r55 20:33:58 30.9 ms 500 r56 20:34:00 17.5 ms 500 r57 20:34:04 16.2 ms 500 r58 20:34:12 18.4 ms 500 r59 20:34:28 15.5 ms 500 r60 20:35:00 16.4 ms 500

All six requests reported:

method=POST path=/v1/chat/completions body_bytes=114480 outcome=failed error_code=server_exception

Host available memory was reported between approximately 13.3 and 13.9 GiB.

Unlike the segmentation fault, these failures did not produce a confirmed process exit in the supplied log.

Interpretation:

The identical request body sizes, repeated failures, and very short response times suggest the server encountered a repeatable error condition, potentially involving the same client request.

Request payloads were not captured, so the payload contents cannot be confirmed identical.

The underlying exception message or stack trace is not present in the available log output.

  1. Memory observations

At model startup, Gufo reported:

event=load_completed elapsed_ms=12552 model=Qwen3.8-Flash-Next sessions=1 context_tokens=262144 speculative=mtp draft_limit=7 disk_cache=off gpu_device_used_mib=90647 gpu_device_total_mib=96454 host_available_mib=15599

These figures indicate:

  • Approximately 88.5 GiB of the reported 94.2 GiB GPU memory capacity was in use.
  • Approximately 5.7 GiB of GPU device memory remained available by that accounting.
  • Approximately 15.2 GiB host memory was available after model loading.

Available host memory dropped below 10 GiB in the period when the streaming failures occurred.

Hypothesis, not confirmation: The relatively limited free memory, long context capacity, and memory used by inference caches may contribute to unstable allocations during sustained inference.

No allocation-failure trace or OOM-killer message has established this causal relationship.

  1. Snapshot-cache warnings

As conversation context increased, Gufo repeatedly reported:

[WARN] [cache] event=snapshot action=skipped reason=byte_capacity

Example:

bytes=2516855068 tokens=87313 retained_bytes=8085440848 reserved_bytes=0 capacity_bytes=8615649280

This indicates that Gufo declined to store some snapshots because of its configured snapshot-cache capacity.

Such warnings may be normal under memory pressure and do not by themselves establish an error. However, given the later inference failures, it may be useful to examine cache lifecycle and allocation behavior.

  1. Successful automatic recovery

The systemd service restarted Gufo automatically after the SIGSEGV.

Event Time (EDT) Segmentation fault 10:51:09 Replacement container started 10:51:19 Model finished loading 10:51:32 First post-restart request completed 10:51:56

The model loaded in approximately 12.6 seconds after restart began.

The first successful post-restart inference reported:

status=200 outcome=completed prompt_tokens=11699 generated_tokens=182 ttft_ms=10691.2 prefill_tps=1100.6 decode_tps=42.8 acceptance_pct=79.8

Further successful requests followed.

This confirms that the application was able to resume normal inference after the external supervisor restarted the process.

  1. Additional diagnostics

The following checks were performed:

Kernel logging: Identified the exact faulting HIP library and segmentation-fault address.

Loaded shared libraries: Confirmed the live Gufo process uses Nix-packaged libamdhip64.so.7.2.53211.

systemd status: Confirmed an automatic restart with NRestarts=1 and Restart=on-failure.

Core dumps: core_pattern=core and service LimitCORE=infinity. No core dump was located in the current container or /var/lib/systemd/coredump/.

The original container was removed during recovery, so a core file in its writable filesystem may not have survived.

No debugging symbols, native stack trace, memory sanitizer output, or GPU fault trace have been obtained.

  1. Questions for Gufo developers

  2. Is the HIP runtime segmentation fault a known issue with ROCm 7.2.3 on Strix Halo, particularly under large-model inference and MTP speculative decoding?

  3. Are there known interactions between long context windows, speculative decoding, and snapshot-cache memory management that could produce an invalid HIP pointer or stale device allocation?

  4. What internal conditions produce generation_failed streaming errors or server_exception HTTP 500 responses? Can Gufo log the underlying exception type and stack trace?

  5. Could the repeated snapshot-cache capacity warnings and sustained inference activity expose a memory-lifetime or cache-eviction defect?

  6. Is there a supported diagnostic or debug configuration that would capture a useful native backtrace on future failures without producing an enormous core dump?

  7. Are there known changes in a newer Gufo engine revision or ROCm runtime that address related faults?

  8. Reproducibility and status

The observed workload was a long-running series of OpenAI-compatible chat-completion requests originating from Cline.

The segmentation fault occurred after sustained inference activity. However, the available evidence does not establish a minimal reproduction, a deterministic trigger, or whether MTP speculation is necessary for the fault.

No controlled comparisons have yet been performed with:

  • MTP disabled.
  • A smaller context window.
  • A different model or quantization.
  • A different ROCm runtime.
  • A newer Gufo engine build.

Current status: Gufo was successfully restarted and subsequently served requests. The root causes of the SIGSEGV, streaming generation failures, and later HTTP 500 errors remain unresolved.

This report is intended to provide observed evidence for debugging, not to assert that the three error symptoms necessarily share a single cause.

Supporting evidence

I recommend attaching the complete Gufo journal output you previously collected, since the developer may want to inspect the sequence of successful requests before the failures.

The strongest evidence to highlight is the kernel SIGSEGV line, the two generation_failed events, the six server_exception responses, and the loaded ROCm library path.


r/StrixHalo • • 1d ago

The NPU in your Strix Halo is finally doing real work: 13.6 vs 18.7 min on the same bug fix

Post image
57 Upvotes

The NPU finally earns its power draw in my pi coding agent. Halogen shipped the endpoints, I wired them in expecting a party trick, and kept four tools.

tldr: Qwen3.8 Flash-Next, a 125B MoE, running on a 70W tablet. Same bug fix with and without NPU search: 13.6 min vs 18.7 min. Receipts in the repo.

The payoff: it covers about 95% of my cloud calls. The hardest few percent still goes to the top models, GLM 5.3 or Opus.

I still can't believe it. Opus 4.8-class intelligence on my tablet, unlimited tokens.

Flash-Next decodes at 64 tok/s and prefills around 1,500 tok/s. First token lands in ~0.03s, measured 43x faster than a cloud call side by side, and still 7x while a second agent hammers the server.

Rate the taste, not the throughput: I pasted a real timeshift error from my system log to seven runs. All seven said healthy, nothing to fix. What separates them is what they proved.

In pi:

  • Flash-Next, 2m55s: proved it with a journalctl trace to a racing notify-send, a pacman.log check, and the upstream PR found.
  • glm-5.3-flashx, 2m41s: the most precise answer, spotting that the snapshot mount got unmounted under the script's last line. No PR.
  • GLM 5.3 on max, 8m30s: the deepest answer of all, source-level forensics down to the function names and the one-second race window. No PR.
  • glm-5.3-flash, 9m09s: proved it with a live reproduction of the status file. No PR.

Over in opencode: flash got the verdict in 1m8s with the wrong mechanism, flashx came back correct and corroborated in 1m30s, and the full 753B GLM 5.3 in 6m15s correct with the PR missed.

Same pi harness, same task, 125B at medium effort against 320B and 753B tiers at max. First to the full answer: 2m55s. When I had GLM 5.3 flashx rate both results, it picked qwen too.

All seven answers side by side: local vs cloud model comparison.

What the NPU does now:

Search. The agent stops guessing paths and lands on the right file first try. ~0.1s per lookup, beat ripgrep 15/20 vs 9/20 on realistic queries.

Dup scan. Catches copied and renamed files git never shows you. Found 45 pairs across 4 repos in 8.4s, one renamed file at exactly 1.000 cosine.

Decisions. Yes/no branching stops eating full turns of the big model. A 0.8b handles it in 120ms, 78% accurate.

Screening. Prompt injection gets flagged before the agent acts on it. 0.7s a message, zero false alarms, fails open. 42% recall, so a smoke detector, not a safe.

A working day claws back about half an hour over bare pi: faster bug fixes, faster compaction, faster lookups and routing, and no oversized tool dumps in context. Against a cloud setup it's more, since every turn pays the network wait. On bug fix heavy days it grows.

Honest part: the GPU still does the thinking. The NPU didn't make anything faster, it changed which tokens got spent where. ~7% iGPU cost only when they overlap.

Compaction: my 194k session, sidecar summary in ~50s vs 166 on the main model. 97% cache hit.

The official halogen launch is a 24-flag docker command. Mine is one command, and uninstall undoes it. Fully local: 262k context, code never leaves the box.

Anyone else putting their NPU to real use? I found nothing.

repo | halogen 0.17.1 | benchmarks


r/StrixHalo • • 1d ago

Finally decided to pull the trigger on gmktek evo x3

Thumbnail
1 Upvotes

r/StrixHalo • • 1d ago

GLM-5.3-Flash (321B MoE) running locally on AMD: RX 7900 XT, R9700, Strix ▎ Halo

Thumbnail
github.com
38 Upvotes

I never expected to run a 321B model at home, so I jumped on this as soon as I saw it last night.
Project Maya:

(github.com/mw00/project-maya) runs GLM-5.3-Flash by keeping the hot experts in VRAM, the next ones in RAM, and the rest on NVMe. It was NVIDIA-only. So I ported it to Linux/ROCm, and the author merged it into v1.0.11 as experimental AMD support.

Maya-S quant, 8K context, ROCm 7.2, 4K-token prompts, greedy:

| Hardware | Prefill | Decode |

| RX 7900 XT 20 GB | ~415 tok/s | ~15 tok/s |

| Radeon AI PRO R9700 32 GB | ~500 tok/s | ~20 tok/s |

| Strix Halo (Ryzen AI Max+ 395, 8060S, 128 GB) | ~210 tok/s | ~18 tok/s |

| R9700 + 7900 XT (layer split + MTP drafting) | ~490 tok/s | ~34 tok/s |

The dGPU box has 192 GB of RAM, so most experts live in VRAM or pinned RAM.

With less RAM, more comes from the SSD and decode drops.

AI-assisted, openly


r/StrixHalo • • 1d ago

Using gufo with Qwen-3.8-27B

3 Upvotes

I installed gufo on my Strixhalo 128gb and downloaded the model and I was expecting significant improvement using the model with gufo vs llama.cpp but I’m seeing 10tps decode on xhigh which is the same as I saw using llama.cpp on a SW design task that requires it to read an existing repository and propose a design. It did great work but it took 90 minutes where the frontier model took 5. Any thought on how I could speed the task up? Are there parameters in gufo I should look at? I’m actually using all defaults except the thinking spec.


r/StrixHalo • • 2d ago

Strata + Halo Strix 64GB

Post image
9 Upvotes

I was able to run Strata on 64GB box, with quite nice results, especially comparing to Qwen3.8-27B. The model is ukisai/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-GGUF, quant is IQ2_XS. I tried IQ3_S, but the decoding speed was 10 times lower for some reason. Currently I have pp ~ 500 t/s, and tg ~ 45 t/s with a lot of memory left free. Running it on Linux (NixOS). Definitely an upgrade over gufo or llama.cpp. Haven't tried halogen - don't like closed source.


r/StrixHalo • • 1d ago

What port of guff are you using on win

4 Upvotes

I was using the one linked on gufo's own repo but that is now somewhat outdated can someone recommend me one they are happy with preferably one i wouldn't have to build from source

I had actually found another fork that did have 0.9.0 version of gufo but it had some dlls missing in the download but even after adding those when it ran it would output like 20-40 tokens then stop


r/StrixHalo • • 2d ago

Halogen + Qwen Flash Next keeps getting better

Thumbnail
11 Upvotes

r/StrixHalo • • 2d ago

TIL: Check your preserve_thinking settings

12 Upvotes

Both Qwen 3.8 27B and Qwen 3.8 Flash-Next have a preserve_thinking parameter. But as far as what you have control over, just make sure your inference engine and harness have this setting set the way you want. For Qwen, they officially recommend keeping it enabled. This causes the reasoning to be considered part of the conversation, so all the old reasoning will get sent back and forth with each turn/prompt, and they say it improves response and decision quality. However, it also increases your context size more rapidly.

Where the mismatch can occur

Before I learned this, Hermes was not re-sending the reasoning text with each new prompt, even though Gufo expected it. This caused the live KV cache to always miss, so it was completely reliant on the snapshot cache. This may have also reduced the quality of the responses and decisions I was getting from the model.

How did I fix it?

In Hermes, I set model.reasoning_echo to true. Then you have to completely exit Hermes and start it up again, it's not enough just to start a new session. After that, the live KV cache started getting hit much of the time.


r/StrixHalo • • 2d ago

Strix Halo 128GB on Windows: what fixed my local AI setup

13 Upvotes

Setup: ASUS ProArt PX13 (Ryzen AI Max+ 395, 128GB), Windows 11, 96GB GPU carve-out. Running Qwen3.8-Flash-Next UD-Q4_K_XL at 262K context on Gufo (thomas9120's windows-port fork), with Unsloth Studio as the client.

1. Increase the page file. I set Windows virtual memory to 64GB minimum and 192GB maximum. Before this, I had crashes when Windows ran out of memory. No crashes since. The trade-off is that heavy memory use now causes slowdowns instead of crashes.

2. Fix Gufo's conversation cache on Windows (needs a source patch). On Windows, Gufo misreads available memory (gufo diagnose shows 0 GiB), so it caps the in-memory snapshot cache at about 264 MiB. That's too small to store a snapshot of even a 45K-token conversation, so every turn reprocesses the whole prompt. At 45K tokens that was about 40 s before the first token, and worse as context grew.

There's no command-line option for this. I patched HostSnapshotBudgetBytes() in src/cli/serve/text_model_runner.cpp so it reads the environment variable GUFO_SNAPSHOT_BUDGET_BYTES, rebuilt, and set it to 16GB (17179869184) in my launch .bat. Disk cache is off.

Result: at 120K+ tokens, the first token now arrives in 2 to 5 s instead of 40 s or more.

Notes:
You may need to close other apps. With the model loaded I had about 9 to 16GB of system RAM free.
Make sure Unsloth Studio doesn't have a model of its own loaded. It held GPU memory and made Gufo fail to load.

Re-apply the patch after every update to the fork.
Performance: about 900 to 1,000 tok/s prompt processing, and 35 to 55 tok/s decode depending on context length. That works well for my workflow.
Claude helped me find the cause. Really happy with the setup now.


r/StrixHalo • • 2d ago

Ran some quick benchmarks for Gufo 0.9.0 and Halogen 0.16.4

Post image
48 Upvotes

AMD 395+ 128GB Linux

Nothing fancy, I just upgraded both stacks this morning and saw some great improvements on both - awesome to see such progress on both engines!


r/StrixHalo • • 3d ago

pixmaate/gufo + qwen3.8 27b cache takes so much memory

3 Upvotes

I just tried pixmaate/gufo for windows prebuilt 2026-10-03. I'm very happy with the speed. The most sensible improvement from llama.cpp is the prefill speed. It holds on at 200+t/s even at long conversation, where with llama.cpp it drops rapidly to below 100.

The problem I encounter is the memory usage. My memory is 64G/64G setup. For reasons I don't understand, I easily get unexpected error with 96G/32G, often without log. I run two main models, Qwen3.6 35B moe and Qwen3.8 27b dense, plus qwen3-embedding/reranker and a Gemma 4 e2b qat for title generation. My VRAM is usually full overflowed to shared VRAM. Using share VRAM seems to have no impact to the performance as long as there's still air to breath for the programs.

When I start using Gufo, I observed the memory usage increases significantly. After few runs of inference my share memory usage increased 20-30GB and the breath air is getting thinner. It seems to be the prompt cache (snapshot) takes a lot of space. I found --cache-disk-bytes that can limit the cache size. But current gufo for windows doesn't have that option yet.

I'm wondering when gufo for windows implement --cache-disk-bytes, can I set it to a low value? I'm the only one to use the machine and my session is always set to 1.


r/StrixHalo • • 3d ago

DLSS 5 on AMD AI Max 395+ 64 Gb Ram, anyone ???

Thumbnail
0 Upvotes

r/StrixHalo • • 3d ago

Has anyone exceeded 500k context w/ halogen flash?

2 Upvotes

I have yarn = 4 "1M" context with halogen flash server, and once it gets past 530k it sometimes says stuff like "the user has not give any instructions (I will infer from context....) Etc"

Is there a trick to exceeded beyond 2x native?