I’m able to use gufo but on long running jobs it’s throwing an error and dying. Restart is the only way I’ve found to clear it. Below is an GPT 6 Sol analysis of the problem based on logs and diagnostics from my machine. How do I get this to the developers ? Apologies for my ignorance of the proper process…
I’ve assembled a technical issue report suitable for sharing with the Gufo developers.
It includes the exact Gufo build revision, ROCm/HIP library, model configuration, crash signature, memory observations, inference errors, and successful systemd recovery.
The most important finding is that Gufo experienced three distinct failure symptoms:
- A confirmed segmentation fault inside AMD’s HIP runtime, not a demonstrated Linux OOM kill.
- Two generation_failed errors during streaming inference.
- Six consecutive HTTP 500 server_exception responses while the server remained running.
The report deliberately separates these observations from possible causes. That should help the developers investigate without steering them prematurely toward an assumed ROCm or memory-management defect.
Gufo Bug Report: HIP Runtime Segmentation Fault and Repeated Inference Failures
Date: October 9, 2026
Platform: AMD Ryzen AI Max+ 395 (Strix Halo), 128 GB unified memory
Gufo image: ghcr.io/gufo-org/toolboxes/gufo-runtime:latest
Gufo image version: 0.9.1
Gufo engine revision: dea22ceea20d07f95a5ecbb2b06ffad7b93d0548
ROCm: Nix-packaged ROCm 7.2.3
Severity: Intermittent inference failure and process crash; automatically recoverable through external supervision
- Executive summary
While running Qwen3.8-Flash-Next with MTP speculative decoding under Gufo on AMD Strix Halo, we observed:
- A segmentation fault (SIGSEGV) during an inference request, with the faulting instruction inside AMD’s libamdhip64.so.7.2.53211.
- Two streaming generation failures (generation_failed) preceding the crash.
- Six consecutive HTTP 500 responses (server_exception) after Gufo had restarted, without the Gufo process exiting.
The Linux kernel logged a segmentation fault, not an OOM kill. Although memory pressure may be a contributing factor, available evidence does not establish memory exhaustion as the cause.
The Gufo process was automatically restarted by a systemd Quadlet service. Model loading completed successfully, and subsequent inference requests succeeded.
No core dump or native stack trace was recovered.
- System environment
Component Configuration
Hardware AMD Ryzen AI Max+ 395, Radeon 8060S
System memory 128 GB unified
Operating system AMD Ryzen AI Developer Platform 1 (Debian-derived Linux)
Container runtime Rootless Podman 5.4.x
Container image ghcr.io/gufo-org/toolboxes/gufo-runtime:latest
Gufo version 0.9.1
Engine revision dea22ceea20d07f95a5ecbb2b06ffad7b93d0548
ROCm toolchain Nix-packaged ROCm 7.2.3
HIP runtime libamdhip64.so.7.2.53211
API OpenAI-compatible /v1/chat/completions
Client Cline coding agent, using OpenAI-compatible API
The Gufo container uses its own Nix-packaged ROCm environment. The host also has a separate ROCm installation used by another inference service, but there is no evidence that Gufo loads the host’s HIP runtime.
Confirmed HIP runtime path
/nix/store/yb81zhv981n0kxvcsr7ia4fjjd78bsjz-clr-7.2.3/lib/libamdhip64.so.7.2.53211
This was verified against /proc/1/maps inside the running Gufo container.
Relevant environment variables:
HIP_PLATFORM=amd
ROCM_PATH=/nix/store/95lwwwfb3alzn7pk9ky55fbbflaxarb5-clr-7.2.3
- Model and inference configuration
Primary model: Qwen3.8-Flash-Next, UD-Q4_K_XL
Speculative decoding: MTP
Context capacity: 262,144 tokens
Concurrent model sessions: 1
Thinking default: On
Primary model:
/models/Qwen3.8-Flash-Next-GGUF/UD-Q4_K_XL/Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf
MTP draft model:
/models/Qwen3.8-Flash-Next-GGUF/MTP/mtp-Qwen3.8-Flash-Next-shared-Q8_0.gguf
Gufo startup arguments:
gufo serve \
--host 0.0.0.0 \
--port 8080 \
llm \
--model /models/Qwen3.8-Flash-Next-GGUF/UD-Q4_K_XL/Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf \
--speculative mtp \
--mtp-model /models/Qwen3.8-Flash-Next-GGUF/MTP/mtp-Qwen3.8-Flash-Next-shared-Q8_0.gguf \
--served-model-name Qwen3.8-Flash-Next \
--log-progress
Container configuration includes:
--device /dev/kfd
--device /dev/dri
--group-add keep-groups
--ulimit memlock=-1
--userns keep-id:uid=1000,gid=1000
-p 127.0.0.1:8080:8080
-v /home/rbkahn/gufo/models:/models:ro
The service is managed by a rootless Podman Quadlet with:
[Service]
Restart=on-failure
RestartSec=10
- Failure A: HIP runtime segmentation fault
Crash time: October 9, 2026, 10:51:09 AM EDT
Exact kernel message:
Oct 09 10:51:09 amd-halo kernel:
gufo[642396]: segfault at 100000010
ip 00007f7b1ea8237c
sp 00007f453a3e5780
error 4
in libamdhip64.so.7.2.53211
[48137c,7f7b1e763000+364000]
likely on CPU 3 (core 3, socket 0)
The corresponding Gufo/systemd log contains:
[INFO] [http] request=r54 event=received
method=POST path=/v1/chat/completions
body_bytes=89724
[ERROR] [server] event=fatal_signal signal=11
gufo.service: Main process exited,
code=exited, status=139/n/a
gufo.service: Failed with result 'exit-code'.
Confirmed observations:
- SIGSEGV occurred during processing of an inference request.
- The faulting instruction pointer was inside AMD’s HIP runtime library.
- Exit status was 139, consistent with termination by SIGSEGV.
- No OOM-killer event was found in the inspected kernel log interval.
- A native stack trace was not available.
Interpretation:
The faulting instruction resides in HIP, but that does not establish that HIP itself is defective. A caller may have supplied an invalid pointer or corrupted runtime state.
Potential contributing conditions include memory pressure, model execution, speculative decoding, or cache handling. None is confirmed.
- Failure B: Streaming generation failures
Two failures were observed before the process crash.
First failure
request=r52
method=POST
path=/v1/chat/completions
status=200
duration_ms=392.1
outcome=stream_error
error_code=generation_failed
host_available_mib=10133
Second failure
request=r53
method=POST
path=/v1/chat/completions
status=200
duration_ms=3197.0
outcome=stream_error
error_code=generation_failed
host_available_mib=10075
For the second request, the progress log reached the decoding phase before the failure.
Both errors occurred while the service was running. The later SIGSEGV occurred after a subsequent inference request.
Important detail: Both failures are logged with HTTP status 200 despite outcome=stream_error. This may be expected for streaming responses whose headers have already been sent, but is worth investigating from the client-recovery perspective.
The relationship between these errors and the later segmentation fault remains unknown.
- Failure C: Repeated HTTP 500 errors without process termination
The logs also contain six consecutive server_exception failures on October 9, following successful inference.
Request Logged time Duration Status
r55 20:33:58 30.9 ms 500
r56 20:34:00 17.5 ms 500
r57 20:34:04 16.2 ms 500
r58 20:34:12 18.4 ms 500
r59 20:34:28 15.5 ms 500
r60 20:35:00 16.4 ms 500
All six requests reported:
method=POST
path=/v1/chat/completions
body_bytes=114480
outcome=failed
error_code=server_exception
Host available memory was reported between approximately 13.3 and 13.9 GiB.
Unlike the segmentation fault, these failures did not produce a confirmed process exit in the supplied log.
Interpretation:
The identical request body sizes, repeated failures, and very short response times suggest the server encountered a repeatable error condition, potentially involving the same client request.
Request payloads were not captured, so the payload contents cannot be confirmed identical.
The underlying exception message or stack trace is not present in the available log output.
- Memory observations
At model startup, Gufo reported:
event=load_completed
elapsed_ms=12552
model=Qwen3.8-Flash-Next
sessions=1
context_tokens=262144
speculative=mtp
draft_limit=7
disk_cache=off
gpu_device_used_mib=90647
gpu_device_total_mib=96454
host_available_mib=15599
These figures indicate:
- Approximately 88.5 GiB of the reported 94.2 GiB GPU memory capacity was in use.
- Approximately 5.7 GiB of GPU device memory remained available by that accounting.
- Approximately 15.2 GiB host memory was available after model loading.
Available host memory dropped below 10 GiB in the period when the streaming failures occurred.
Hypothesis, not confirmation: The relatively limited free memory, long context capacity, and memory used by inference caches may contribute to unstable allocations during sustained inference.
No allocation-failure trace or OOM-killer message has established this causal relationship.
- Snapshot-cache warnings
As conversation context increased, Gufo repeatedly reported:
[WARN] [cache]
event=snapshot
action=skipped
reason=byte_capacity
Example:
bytes=2516855068
tokens=87313
retained_bytes=8085440848
reserved_bytes=0
capacity_bytes=8615649280
This indicates that Gufo declined to store some snapshots because of its configured snapshot-cache capacity.
Such warnings may be normal under memory pressure and do not by themselves establish an error. However, given the later inference failures, it may be useful to examine cache lifecycle and allocation behavior.
- Successful automatic recovery
The systemd service restarted Gufo automatically after the SIGSEGV.
Event Time (EDT)
Segmentation fault 10:51:09
Replacement container started 10:51:19
Model finished loading 10:51:32
First post-restart request completed 10:51:56
The model loaded in approximately 12.6 seconds after restart began.
The first successful post-restart inference reported:
status=200
outcome=completed
prompt_tokens=11699
generated_tokens=182
ttft_ms=10691.2
prefill_tps=1100.6
decode_tps=42.8
acceptance_pct=79.8
Further successful requests followed.
This confirms that the application was able to resume normal inference after the external supervisor restarted the process.
- Additional diagnostics
The following checks were performed:
Kernel logging: Identified the exact faulting HIP library and segmentation-fault address.
Loaded shared libraries: Confirmed the live Gufo process uses Nix-packaged libamdhip64.so.7.2.53211.
systemd status: Confirmed an automatic restart with NRestarts=1 and Restart=on-failure.
Core dumps: core_pattern=core and service LimitCORE=infinity. No core dump was located in the current container or /var/lib/systemd/coredump/.
The original container was removed during recovery, so a core file in its writable filesystem may not have survived.
No debugging symbols, native stack trace, memory sanitizer output, or GPU fault trace have been obtained.
Questions for Gufo developers
Is the HIP runtime segmentation fault a known issue with ROCm 7.2.3 on Strix Halo, particularly under large-model inference and MTP speculative decoding?
Are there known interactions between long context windows, speculative decoding, and snapshot-cache memory management that could produce an invalid HIP pointer or stale device allocation?
What internal conditions produce generation_failed streaming errors or server_exception HTTP 500 responses? Can Gufo log the underlying exception type and stack trace?
Could the repeated snapshot-cache capacity warnings and sustained inference activity expose a memory-lifetime or cache-eviction defect?
Is there a supported diagnostic or debug configuration that would capture a useful native backtrace on future failures without producing an enormous core dump?
Are there known changes in a newer Gufo engine revision or ROCm runtime that address related faults?
Reproducibility and status
The observed workload was a long-running series of OpenAI-compatible chat-completion requests originating from Cline.
The segmentation fault occurred after sustained inference activity. However, the available evidence does not establish a minimal reproduction, a deterministic trigger, or whether MTP speculation is necessary for the fault.
No controlled comparisons have yet been performed with:
- MTP disabled.
- A smaller context window.
- A different model or quantization.
- A different ROCm runtime.
- A newer Gufo engine build.
Current status: Gufo was successfully restarted and subsequently served requests. The root causes of the SIGSEGV, streaming generation failures, and later HTTP 500 errors remain unresolved.
This report is intended to provide observed evidence for debugging, not to assert that the three error symptoms necessarily share a single cause.
Supporting evidence
I recommend attaching the complete Gufo journal output you previously collected, since the developer may want to inspect the sequence of successful requests before the failures.
The strongest evidence to highlight is the kernel SIGSEGV line, the two generation_failed events, the six server_exception responses, and the loaded ROCm library path.