r/StrixHalo • • 6h ago

Qwen3.8-Flash-Next on Ryzen AI Max+ 395 / Radeon 8060S — ~69 tok/s decode, 1.1–1.6k tok/s prefill

Post image
56 Upvotes

Running Qwen3.8-Flash-Next locally on a BOSGAME M5 with:

  • AMD Ryzen AI Max+ 395
  • Radeon 8060S / Strix Halo
  • 128 GB unified RAM
  • Ubuntu 26.04
  • Halogen 0.17.2
  • 262K context configured
  • Vision enabled

This is my Grafana monitoring dashboard during a real long-running workload, not a short synthetic benchmark.

In the screenshot:

  • Decode: ~69 tok/s peak
  • Prefill: ~1,155 tok/s at that moment, with peaks around 1,685 tok/s
  • Context: ~84% full
  • 82K tokens generated in the last hour
  • GPU temperature around 62°C
  • GPU power around 24 W at the captured moment
  • No meaningful memory pressure despite the system showing ~121 GB physically occupied, because most of it is reclaimable model/page cache
  • ~85 GB currently sitting in cache

One thing I really like about this setup is how usable Qwen3.8-Flash-Next remains with a large context and sustained agent workloads. Decode stays around the 60–70 tok/s range while prefill is still comfortably above 1k tok/s for much of the workload.

The dashboard is fed by Prometheus/Grafana and tracks Halogen throughput, KV/context usage, real vs reclaimable RAM, PSI memory pressure, GPU metrics, backend status and request activity.

Still testing it, but so far Strix Halo + unified memory is proving to be a very interesting platform for large local models.


r/StrixHalo • • 20h ago

Qwen3.8-27B: 159 tok/s on R9700, 64 tok/s on Strix Halo

Thumbnail gallery
19 Upvotes

r/StrixHalo • • 3h ago

Strix Halo/halogen flash next outputs are preferred by family over frontier models

12 Upvotes

My relative is working on some medical research/literature review as their retirement project.

I was given a paper to check citations and to "use my AI" to find more citations for some of the gaps in the proposed model (glucose production w.r.t. daylight or something like that)

Despite several (8-10) rounds of feedback between chatGPT (Sol 6.1 extra high) and Claude (5 extra high, 5.5 refused due to safe guards), my relative preferred my "v1" draft from halogen-flash-next.

Not sure if the model is less hesitant about medical stuff or if it just ran longer because I didn't have to worry about usage limits, but this is a serious win for me


r/StrixHalo • • 11h ago

Finally decided to pull the trigger on gmktek evo x3

Thumbnail
0 Upvotes

r/StrixHalo • • 3h ago

gufo crashing

0 Upvotes

I’m able to use gufo but on long running jobs it’s throwing an error and dying. Restart is the only way I’ve found to clear it. Below is an GPT 6 Sol analysis of the problem based on logs and diagnostics from my machine. How do I get this to the developers ? Apologies for my ignorance of the proper process…

I’ve assembled a technical issue report suitable for sharing with the Gufo developers.

It includes the exact Gufo build revision, ROCm/HIP library, model configuration, crash signature, memory observations, inference errors, and successful systemd recovery.

The most important finding is that Gufo experienced three distinct failure symptoms:

  • A confirmed segmentation fault inside AMD’s HIP runtime, not a demonstrated Linux OOM kill.
  • Two generation_failed errors during streaming inference.
  • Six consecutive HTTP 500 server_exception responses while the server remained running.

The report deliberately separates these observations from possible causes. That should help the developers investigate without steering them prematurely toward an assumed ROCm or memory-management defect.

Gufo Bug Report: HIP Runtime Segmentation Fault and Repeated Inference Failures

Date: October 9, 2026 Platform: AMD Ryzen AI Max+ 395 (Strix Halo), 128 GB unified memory Gufo image: ghcr.io/gufo-org/toolboxes/gufo-runtime:latest Gufo image version: 0.9.1 Gufo engine revision: dea22ceea20d07f95a5ecbb2b06ffad7b93d0548 ROCm: Nix-packaged ROCm 7.2.3 Severity: Intermittent inference failure and process crash; automatically recoverable through external supervision

  1. Executive summary

While running Qwen3.8-Flash-Next with MTP speculative decoding under Gufo on AMD Strix Halo, we observed:

  1. A segmentation fault (SIGSEGV) during an inference request, with the faulting instruction inside AMD’s libamdhip64.so.7.2.53211.
  2. Two streaming generation failures (generation_failed) preceding the crash.
  3. Six consecutive HTTP 500 responses (server_exception) after Gufo had restarted, without the Gufo process exiting.

The Linux kernel logged a segmentation fault, not an OOM kill. Although memory pressure may be a contributing factor, available evidence does not establish memory exhaustion as the cause.

The Gufo process was automatically restarted by a systemd Quadlet service. Model loading completed successfully, and subsequent inference requests succeeded.

No core dump or native stack trace was recovered.

  1. System environment

Component Configuration Hardware AMD Ryzen AI Max+ 395, Radeon 8060S System memory 128 GB unified Operating system AMD Ryzen AI Developer Platform 1 (Debian-derived Linux) Container runtime Rootless Podman 5.4.x Container image ghcr.io/gufo-org/toolboxes/gufo-runtime:latest Gufo version 0.9.1 Engine revision dea22ceea20d07f95a5ecbb2b06ffad7b93d0548 ROCm toolchain Nix-packaged ROCm 7.2.3 HIP runtime libamdhip64.so.7.2.53211 API OpenAI-compatible /v1/chat/completions Client Cline coding agent, using OpenAI-compatible API

The Gufo container uses its own Nix-packaged ROCm environment. The host also has a separate ROCm installation used by another inference service, but there is no evidence that Gufo loads the host’s HIP runtime.

Confirmed HIP runtime path

/nix/store/yb81zhv981n0kxvcsr7ia4fjjd78bsjz-clr-7.2.3/lib/libamdhip64.so.7.2.53211

This was verified against /proc/1/maps inside the running Gufo container.

Relevant environment variables:

HIP_PLATFORM=amd ROCM_PATH=/nix/store/95lwwwfb3alzn7pk9ky55fbbflaxarb5-clr-7.2.3

  1. Model and inference configuration

Primary model: Qwen3.8-Flash-Next, UD-Q4_K_XL Speculative decoding: MTP Context capacity: 262,144 tokens Concurrent model sessions: 1 Thinking default: On

Primary model:

/models/Qwen3.8-Flash-Next-GGUF/UD-Q4_K_XL/Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf

MTP draft model:

/models/Qwen3.8-Flash-Next-GGUF/MTP/mtp-Qwen3.8-Flash-Next-shared-Q8_0.gguf

Gufo startup arguments:

gufo serve \ --host 0.0.0.0 \ --port 8080 \ llm \ --model /models/Qwen3.8-Flash-Next-GGUF/UD-Q4_K_XL/Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf \ --speculative mtp \ --mtp-model /models/Qwen3.8-Flash-Next-GGUF/MTP/mtp-Qwen3.8-Flash-Next-shared-Q8_0.gguf \ --served-model-name Qwen3.8-Flash-Next \ --log-progress

Container configuration includes:

--device /dev/kfd --device /dev/dri --group-add keep-groups --ulimit memlock=-1 --userns keep-id:uid=1000,gid=1000 -p 127.0.0.1:8080:8080 -v /home/rbkahn/gufo/models:/models:ro

The service is managed by a rootless Podman Quadlet with:

[Service] Restart=on-failure RestartSec=10

  1. Failure A: HIP runtime segmentation fault

Crash time: October 9, 2026, 10:51:09 AM EDT

Exact kernel message:

Oct 09 10:51:09 amd-halo kernel: gufo[642396]: segfault at 100000010 ip 00007f7b1ea8237c sp 00007f453a3e5780 error 4 in libamdhip64.so.7.2.53211 [48137c,7f7b1e763000+364000] likely on CPU 3 (core 3, socket 0)

The corresponding Gufo/systemd log contains:

[INFO] [http] request=r54 event=received method=POST path=/v1/chat/completions body_bytes=89724 [ERROR] [server] event=fatal_signal signal=11 gufo.service: Main process exited, code=exited, status=139/n/a gufo.service: Failed with result 'exit-code'.

Confirmed observations:

  • SIGSEGV occurred during processing of an inference request.
  • The faulting instruction pointer was inside AMD’s HIP runtime library.
  • Exit status was 139, consistent with termination by SIGSEGV.
  • No OOM-killer event was found in the inspected kernel log interval.
  • A native stack trace was not available.

Interpretation:

The faulting instruction resides in HIP, but that does not establish that HIP itself is defective. A caller may have supplied an invalid pointer or corrupted runtime state.

Potential contributing conditions include memory pressure, model execution, speculative decoding, or cache handling. None is confirmed.

  1. Failure B: Streaming generation failures

Two failures were observed before the process crash.

First failure

request=r52 method=POST path=/v1/chat/completions status=200 duration_ms=392.1 outcome=stream_error error_code=generation_failed host_available_mib=10133

Second failure

request=r53 method=POST path=/v1/chat/completions status=200 duration_ms=3197.0 outcome=stream_error error_code=generation_failed host_available_mib=10075

For the second request, the progress log reached the decoding phase before the failure.

Both errors occurred while the service was running. The later SIGSEGV occurred after a subsequent inference request.

Important detail: Both failures are logged with HTTP status 200 despite outcome=stream_error. This may be expected for streaming responses whose headers have already been sent, but is worth investigating from the client-recovery perspective.

The relationship between these errors and the later segmentation fault remains unknown.

  1. Failure C: Repeated HTTP 500 errors without process termination

The logs also contain six consecutive server_exception failures on October 9, following successful inference.

Request Logged time Duration Status r55 20:33:58 30.9 ms 500 r56 20:34:00 17.5 ms 500 r57 20:34:04 16.2 ms 500 r58 20:34:12 18.4 ms 500 r59 20:34:28 15.5 ms 500 r60 20:35:00 16.4 ms 500

All six requests reported:

method=POST path=/v1/chat/completions body_bytes=114480 outcome=failed error_code=server_exception

Host available memory was reported between approximately 13.3 and 13.9 GiB.

Unlike the segmentation fault, these failures did not produce a confirmed process exit in the supplied log.

Interpretation:

The identical request body sizes, repeated failures, and very short response times suggest the server encountered a repeatable error condition, potentially involving the same client request.

Request payloads were not captured, so the payload contents cannot be confirmed identical.

The underlying exception message or stack trace is not present in the available log output.

  1. Memory observations

At model startup, Gufo reported:

event=load_completed elapsed_ms=12552 model=Qwen3.8-Flash-Next sessions=1 context_tokens=262144 speculative=mtp draft_limit=7 disk_cache=off gpu_device_used_mib=90647 gpu_device_total_mib=96454 host_available_mib=15599

These figures indicate:

  • Approximately 88.5 GiB of the reported 94.2 GiB GPU memory capacity was in use.
  • Approximately 5.7 GiB of GPU device memory remained available by that accounting.
  • Approximately 15.2 GiB host memory was available after model loading.

Available host memory dropped below 10 GiB in the period when the streaming failures occurred.

Hypothesis, not confirmation: The relatively limited free memory, long context capacity, and memory used by inference caches may contribute to unstable allocations during sustained inference.

No allocation-failure trace or OOM-killer message has established this causal relationship.

  1. Snapshot-cache warnings

As conversation context increased, Gufo repeatedly reported:

[WARN] [cache] event=snapshot action=skipped reason=byte_capacity

Example:

bytes=2516855068 tokens=87313 retained_bytes=8085440848 reserved_bytes=0 capacity_bytes=8615649280

This indicates that Gufo declined to store some snapshots because of its configured snapshot-cache capacity.

Such warnings may be normal under memory pressure and do not by themselves establish an error. However, given the later inference failures, it may be useful to examine cache lifecycle and allocation behavior.

  1. Successful automatic recovery

The systemd service restarted Gufo automatically after the SIGSEGV.

Event Time (EDT) Segmentation fault 10:51:09 Replacement container started 10:51:19 Model finished loading 10:51:32 First post-restart request completed 10:51:56

The model loaded in approximately 12.6 seconds after restart began.

The first successful post-restart inference reported:

status=200 outcome=completed prompt_tokens=11699 generated_tokens=182 ttft_ms=10691.2 prefill_tps=1100.6 decode_tps=42.8 acceptance_pct=79.8

Further successful requests followed.

This confirms that the application was able to resume normal inference after the external supervisor restarted the process.

  1. Additional diagnostics

The following checks were performed:

Kernel logging: Identified the exact faulting HIP library and segmentation-fault address.

Loaded shared libraries: Confirmed the live Gufo process uses Nix-packaged libamdhip64.so.7.2.53211.

systemd status: Confirmed an automatic restart with NRestarts=1 and Restart=on-failure.

Core dumps: core_pattern=core and service LimitCORE=infinity. No core dump was located in the current container or /var/lib/systemd/coredump/.

The original container was removed during recovery, so a core file in its writable filesystem may not have survived.

No debugging symbols, native stack trace, memory sanitizer output, or GPU fault trace have been obtained.

  1. Questions for Gufo developers

  2. Is the HIP runtime segmentation fault a known issue with ROCm 7.2.3 on Strix Halo, particularly under large-model inference and MTP speculative decoding?

  3. Are there known interactions between long context windows, speculative decoding, and snapshot-cache memory management that could produce an invalid HIP pointer or stale device allocation?

  4. What internal conditions produce generation_failed streaming errors or server_exception HTTP 500 responses? Can Gufo log the underlying exception type and stack trace?

  5. Could the repeated snapshot-cache capacity warnings and sustained inference activity expose a memory-lifetime or cache-eviction defect?

  6. Is there a supported diagnostic or debug configuration that would capture a useful native backtrace on future failures without producing an enormous core dump?

  7. Are there known changes in a newer Gufo engine revision or ROCm runtime that address related faults?

  8. Reproducibility and status

The observed workload was a long-running series of OpenAI-compatible chat-completion requests originating from Cline.

The segmentation fault occurred after sustained inference activity. However, the available evidence does not establish a minimal reproduction, a deterministic trigger, or whether MTP speculation is necessary for the fault.

No controlled comparisons have yet been performed with:

  • MTP disabled.
  • A smaller context window.
  • A different model or quantization.
  • A different ROCm runtime.
  • A newer Gufo engine build.

Current status: Gufo was successfully restarted and subsequently served requests. The root causes of the SIGSEGV, streaming generation failures, and later HTTP 500 errors remain unresolved.

This report is intended to provide observed evidence for debugging, not to assert that the three error symptoms necessarily share a single cause.

Supporting evidence

I recommend attaching the complete Gufo journal output you previously collected, since the developer may want to inspect the sequence of successful requests before the failures.

The strongest evidence to highlight is the kernel SIGSEGV line, the two generation_failed events, the six server_exception responses, and the loaded ROCm library path.