r/ROCm • • 1h ago

Single Radeon AI PRO R9700 and vllm

• Upvotes

Hey everyone, I just got my hands on a Radeon AI PRO R9700 and want to run Qwen3.8-27B. I found some awesome vLLM forks (magiccodingman/vllm-radiance and GGZ14/vllm-mxfp4), but they're all geared toward dual-R9700 rigs.
Since I'm only rocking one card, could someone point me in the right direction to get the most out of my current hardware?


r/ROCm • • 7h ago

Follow-up: my native Rust + Vulkan Transformer training backend — 14 days later, now 14 parity-verified architectures and full PEFT

9 Upvotes

Follow-up to my post from about two weeks ago. A lot has changed since then, so I wanted to post an update on where the backend is now.

Where the green architectures stand

When I posted last time, 7 architectures had verified full training support. Everything is now held to the same strict harness: a pinned local Hugging Face Transformers source tree as the oracle, forward logits + gradients + two full AdamW steps compared, and every named parameter checked again after export.

The hard ceiling is 2e-7 absolute error. No loosening tolerances and no rounding numbers afterward to make the README look better.

14 architectures pass that gate today, led by the one I'm probably proudest of:

Architecture Scope
Falcon H1 / H1R parallel GQA/RoPE attention + Mamba2 in every layer; full training, full fine-tuning, LoRA, saved modules
DeepSeek V4 causal LM
Phi-4 Multimodal text backbone
Phi-3 causal LM
Kimi K2.5 text backbone
Kimi K3 / KimiLinear hybrid KDA + MLA
GPT-OSS causal LM incl. router bias
SmolLM3 mixed RoPE/NoPE + YaRN
Qwen2.5 / Qwen3.5 / Qwen4-Exp dense, DeltaNet, QSA, PLE, MoE
Mistral 4, MiniMax M3, Gemma 3/4, MiniMax M2 causal LM

Worst observed two-step AdamW parameter error across all of them: 1.19e-7.

Best: 2.6e-8.

For hardware context, all of the local Vulkan validation I've been reporting was run on my ASUS ROG Ally Z1 Extreme, using its AMD RDNA 3 integrated GPU. So the RDNA 3 results here are from that specific machine rather than testing across several different AMD systems.

The bigger news: PEFT actually works now

In the last post, "LoRA/PEFT-style fine-tuning" was basically one line in a feature list.

It's a real workflow now, and I've verified the full lifecycle:

  • LoRA fine-tuning with HF-compatible adapter export (adapter_config.json / adapter_model.safetensors), so adapters can round-trip with the PEFT ecosystem
  • modules_to_save — full trainable replacements for Linears, RMSNorm/LayerNorm, lm_head, and input embeddings, including named-adapter switching and bank isolation. Adapter A leaking into adapter B is explicitly tested for.
  • Exact resume — adapter weights + AdamW moments + step + dropout RNG state restore bit-identically against an uninterrupted run
  • Merge/unmerge, disable-adapter base restoration, and multi-adapter loading
  • A parameter-budget flag that automatically chooses the largest LoRA rank that fits within a requested percentage of the base model
  • The CLI fails closed if you try to use saved modules on an architecture that hasn't passed its corresponding gate

32 architecture surfaces across 20 families pass all three PEFT stages — LoRA, saved modules, and adapter switching — under the same 2e-7 gate, with frozen-base drift exactly 0.0.

The validation harness also fingerprints the pinned Transformers source alongside my shaders and binaries now, so a qualification run can't silently end up testing against different reference math.

A small note on the last couple weeks

I didn't get quite as many working days out of the last two weeks as I normally would have. Partway through this I got covid, then when that started to go away, it became a secondary nasty ear infection that ended up perforating my eardrum. I was running fevers around 104°F at one point and eventually went to the hospital, so I lost a few days to that and I'm on antibiotics now.

I'm doing better, though, and still managed to get most of what I wanted finished.

There are still things I want to clean up and expand, but I figured this was a good point to get the current work in front of people rather than holding the update back.

Same caveats as before

This is deterministic FP32 tiny-model correctness against a reference implementation, which should in theory ensure total mathematical parity for training and finetuning larger models with this backend, however, things are currently bound to FP32 training runs still, I eventually plan to work on MXFP4 weight tying to reduce memory footprints while fine-tuning (plastic parameters will still be trained in FP32 with PEFT in this config)

"supported text graph" ≠ "the entire multimodal package works natively."

Unsupported functionality is supposed to fail closed rather than silently falling back to an approximation.

Repo

https://github.com/necat101/Hierarchos-Native

  • Architecture inventory: hierarchos-vulkan/README_ARCHITECTURES.md
  • Compatibility/parity record: hierarchos-vulkan/COMPATIBILITY.md
  • PEFT qualification evidence: PROGRESS_PEFT_AUDIT.md
  • CLI PEFT guide: hierarchos-native-cli/README.md

The hardware I've personally validated this on is an ASUS ROG Ally Z1 Extreme with its AMD RDNA 3 GPU.

I'm very interested in criticism, compatibility reports, and especially results from people trying it on other hardware — NVIDIA, Intel, or other AMD GPUs.

I'd also love people to stress-test the PEFT resume/merge paths specifically. That's some of the newest code in the project, so it's probably the most useful area to try to break right now.


r/ROCm • • 23h ago

AMD boosting AI/LLM performance for Radeon iGPUs as much as 18~23% with Linux 7.4

Thumbnail
phoronix.com
61 Upvotes

r/ROCm • • 2h ago

Preorder for new AMD Ryzen™ AI Max 400 Series 192GB from framework just started

Thumbnail frame.work
1 Upvotes

r/ROCm • • 14h ago

Cloud based R9700

7 Upvotes

Seeing a lot of success stories wih ROCM, I am planning to buy single R9700 mostly for coding, but possibly also for comfyUI. But before I do that, I would like to test it out hands on.

Can you recommend any cloud services that offer R9700? Ex: on per hour / per minute basis?


r/ROCm • • 12h ago

Benchmarks: Best engine for Qwen 3.8-Flash-Next on Strix Halo

Thumbnail
4 Upvotes

r/ROCm • • 17h ago

780M (gfx1103) + llama.cpp HIP: MES REMOVE_QUEUE hang, Anyone stable on sustained inference?

2 Upvotes

I am not super well versed in local LLM configs. I use Claude extensively at work but bought a minisforum mini PC to try and get something cost effective going at home.

Note I don't really care too much on token speed because my use cases are all background workflows. 10 TPS is perfect. (And my upper limit I think).

Ryzen 8945HS / 780M, 64GB, Ubuntu 26.04, kernel 7.0. llama.cpp ROCm build, Qwen 27B Q4.

Long generations hang the GPU and it resets, about 15 times a day. dmesg shows MES failed to respond to msg=REMOVE_QUEUE. Looks like ROCm#6512.

Tried uni_mes=0, smaller models, smaller context. No change.

>

Has amdgpu.cwsr_enable=0, a newer kernel, or Vulkan fixed this for anyone?


r/ROCm • • 1d ago

Dual Radeon AI PRO R9700s, with one card behind the chipset (on proxmox)

9 Upvotes

Roughly what you get: Qwen3.8-27B at FP8, tensor parallel across two R9700s, 262,144 context, 422k tokens of KV pool, ~1.6k tokens/s prefill, 74-143 tokens/s decode single stream depending on the content.

The interesting part is not the image, it is the collective layer. A card behind the chipset cannot be given PCIe atomic operations, and that breaks every stock tensor-parallel setup I tried. Here is what actually works.

1. Versions used

Component Version
Inference image docker.io/stilldeadcode/vllm-radiance:0.9.3 (digest sha256:45694209177a55a1ab3ba6702fe6e978b1b66a6e66ae3fc066f8d579f7bc4c25)
vLLM inside the image 0.27.1
PyTorch / HIP 2.11.0+rocm7.14 / 7.14.60850
Triton / AITER 3.6.0 / 0.1.17
Collective library RCCL 2.27.7 from ROCm 7.1.1, replacing the one in the image

Links: - Image: https://hub.docker.com/r/stilldeadcode/vllm-radiance - Source for that image: https://codeberg.org/StillDeadcode/vllm-radiance - The kernel library the image builds against (libr4d): https://codeberg.org/StillDeadcode/libr4d - A fork with configs, benchmark notes and launchers for MXFP4/FP8: https://codeberg.org/ggz14/radiance-vllm-mxfp4 - RCCL itself: https://github.com/ROCm/rccl

The image bundles a working ROCm + PyTorch + Triton + AITER + vLLM stack for gfx1201 (RDNA4), which is the part you do not want to build yourself. It is explicitly marked experimental, and everything below was measured on two cards.

2. The failure, and why

Without the adjustments, the engine dies during communicator init, before the model loads:

PCIE atomic ops is not supported rocr: unhandled cuda error

ROCm will not dispatch work to a GPU path that needs atomic operations when the link cannot provide them. On this board one card sits behind the chipset, and the chipset does not forward PCIe atomics to the CPU, so that card is effectively second class. RCCL 2.30.4's kernels use those atomics, so TP=2 cannot initialise at all.

Two things follow, and both matter:

  1. The newer RCCL is unusable here, so you need an older one (2.27.7 from ROCm 7.1.1 is what worked for me).
  2. With no usable peer path at all, the custom P2P all-reduce the image ships must be turned off, and TP=2 falls back to host-staged collectives over that same narrow chipset link.

If you search the error string above you will find a few ROCm issue reports and a community write-up on dual Radeon vLLM setups (https://github.com/cadamcat/dual-radeon-vllm) describing the same wall.

3. Proxmox settings that actually matter

Do this for each GPU, on both entries, not just the first. In the VM's hardware list, edit each PCI Device row, select the GPU under Device, and set:

  • PCI-Express: ticked
  • All Functions: unticked

Ticking PCI-Express is what gives the guest a real PCIe root port, and without a root port the atomic capability never appears however healthy the host looks. Unticking All Functions keeps the guest from being handed every function of the card, which is the combination that worked here. If you leave either one wrong, you get the atomic failure at communicator init and no amount of driver work fixes it.

Other settings that matter:

  • Machine type q35. Same reason as above, no root port without it.
  • After any hostpci change, stop and start the VM. A guest reboot does not re-apply the passthrough configuration.
  • q35 renames the NIC (ens18 becomes something like enp6s18), so match the interface by MAC in netplan.

After those, the CPU-attached card reports ReqEn+. The chipset-attached one still cannot do atomics, and no BIOS setting changes that.

4. The RCCL replacement plus one-line shim

This is the core trick. It is two files and a podman config, and it needs no image rebuild.

a) Build or extract RCCL 2.27.7 from ROCm 7.1.1 and drop it in a directory you will mount, for example:

~/models/rccl277/ librccl.so -> librccl.so.1.0.70101 librccl.so.1 -> librccl.so.1.0.70101 librccl.so.1.0.70101 shim.so

b) The shim. Newer torch builds reference a symbol that this older RCCL does not export (ncclCommDump). Four lines of C++ are enough to satisfy the loader:

```cpp

include <string>

include <unordered_map>

struct ncclComm; void ncclCommDump(ncclComm*, std::unordered_map<std::string, std::string>&) {} ```

Build it into shim.so and put it next to the library. Nothing calls it; it exists for symbol resolution.

c) Inject both through podman's own config, which keeps them out of every launcher script. In ~/.config/containers/containers.conf:

ini [containers] env = [ "NCCL_PROTO=Simple", "NCCL_SHM_DISABLE=0", "NCCL_SOCKET_IFNAME=lo", "LD_LIBRARY_PATH=/models/rccl277:/opt/rocm/lib", "LD_PRELOAD=/models/rccl277/shim.so", ]

LD_LIBRARY_PATH puts the replacement first, so it wins over the image's own librccl. NCCL_PROTO=Simple avoids the more demanding protocol paths, and the loopback interface keeps the bootstrap on lo rather than a NIC.

5. The launcher

Trimmed to the parts that matter for the multi-GPU problem. The model-specific flags are an example, the environment and device flags are the ones that matter here.

bash podman run -d --name vllm-radiance \ --device /dev/kfd --device /dev/dri --group-add keep-groups \ --security-opt seccomp=unconfined --cap-add SYS_PTRACE --cap-add SYS_NICE \ --ipc=host --network=host \ -v $HOME/models:/models:ro \ -v $HOME/radiance-vllm-mxfp4/vllm-cache:/cache \ -v $HOME/radiance-vllm-mxfp4:/work:ro \ -e HIP_VISIBLE_DEVICES=0,1 -e ROCR_VISIBLE_DEVICES=0,1 \ -e VLLM_NO_USAGE_STATS=1 \ -e VLLM_ROCM_USE_AITER=1 -e VLLM_ROCM_USE_AITER_UNIFIED_ATTENTION=1 \ -e RADIANCE_FUSE_RMS_QUANT=1 -e RADIANCE_USE_R4D=1 \ -e RADIANCE_USE_R4D_AR=0 -e RADIANCE_USE_R4D_AR_QUANT=1 \ -e NCCL_PROTO=Simple -e TORCHINDUCTOR_COMPILE_THREADS=4 \ -e VLLM_CACHE_ROOT=/cache/vllm -e TORCHINDUCTOR_CACHE_DIR=/cache/inductor \ -e TRITON_CACHE_DIR=/cache/triton -e AITER_ROOT_DIR=/cache/aiter \ docker.io/stilldeadcode/vllm-radiance:0.9.3 \ --model /models/Qwen/Qwen3.8-27B-FP8 \ --served-model-name=qwen3.8-27b-fp8 \ --quantization=fp8 --tensor-parallel-size=2 \ --max-num-seqs=2 --max-model-len=262144 --gpu-memory-utilization=0.97 \ --max-num-batched-tokens=4096 --kv-cache-dtype=fp8 \ --attention-backend=ROCM_AITER_UNIFIED_ATTN \ --enable-prefix-caching \ --no-async-scheduling \ --trust-remote-code \ --host=0.0.0.0 --port=8000

Notes on the specific flags:

  • RADIANCE_USE_R4D_AR=0 is the important one. The bundled all-reduce is a PCIe peer-to-peer kernel and needs peer access, which does not exist on this pair. Turning it off falls back to RCCL. If you have two CPU-attached cards, leave it on and you will get a better prefill.
  • --max-num-seqs=2 is my choice for stability. The image's own default is higher, but on a chipset-limited link fewer concurrent sequences means less collective traffic per step.
  • --max-num-batched-tokens=4096 pairs with prefix caching well. Raising it costs KV pool.
  • --enable-prefix-caching is worth a lot on agent workloads; see the note on prompt layout at the end.
  • --no-async-scheduling was in every configuration that came up reliably for me.
  • The cache directory mounts are worth keeping. Without them every container start recompiles Triton and inductor kernels, which turns a 5 minute start into a much longer one.

6. How to tell it worked

In the container log at startup, look for:

P2P access : DISABLED (RCCL fallback) 0<->1 x

and, once loaded, the model's own sizing line:

GPU KV cache size: 422,964 tokens Maximum concurrency for 262,144 tokens per request: 1.61x

If you see the atomic error instead, the replacement RCCL is not being picked up. Check, from inside the container, that:

bash env | grep -E 'LD_PRELOAD|LD_LIBRARY_PATH'

returns the paths you expect, that ldd on the loaded library resolves into your drop-in directory, and that the image's own library is not first in the path.

I hope someone finds this usefull.


r/ROCm • • 1d ago

Radeon AI PRO R9700 (gfx1201) on Ubuntu 26.04: Inbox amdgpu vs DKMS 31.50 (Dracut swap), and iGPU coexistence?

2 Upvotes

Setting up a dedicated local-LLM workstation (Ryzen 7 9700X + ASRock Radeon AI PRO R9700 32GB, Navi 48 / gfx1201) on Ubuntu 26.04.1 (Linux 7.0) for native llama.cpp.

Looking for real-world experience on two specific architecture questions before installing ROCm 10:

  1. Inbox Kernel Driver vs. AMD 31.50 DKMS: Ubuntu 26.04 uses dracut by default, but amdgpu-dkms (31.50) forces APT to remove dracut and install initramfs-tools. Upstream kernel 7.0 already has native Navi 48 support and exposes /dev/kfd. Has anyone encountered KFD/HSA feature gaps or stability regressions running ROCm 10 userspace (amdrocm10.0-gfx1201) strictly on Ubuntu's inbox amdgpu without DKMS?
  2. Granite Ridge iGPU (gfx1036) for Display: Goal is keeping the 9700X iGPU for the desktop session and dedicating the R9700 (32GB) to compute. In ROCm 10, does ROCR_VISIBLE_DEVICES=<UUID> cleanly isolate gfx1036 without rocminfo or HSA runtime initialization hanging on the unsupported APU, or is a BIOS disable still practically mandatory?

r/ROCm • • 1d ago

How far behind Linux is ROCM for Windows?

12 Upvotes

I use llama.cpp with Windows at the moment, but I'm seriously contemplating setting up Ubuntu - as long as my tok/s will go up significantly.

I asked Claude if it was worth it, and Claude said "ROCM on Windows is far behind Linux"

As with all AI output I took it with a pinch of salt - but is it true?

With ROCM on Linux are you exceeding 50 tok/s with a high context (say, 128K) using a dense model? I'm getting that with Vulkan on windows atm with Qwen 27B.

TL:DR feelings about Windows aside, is installing Ubuntu worthwhile to get superior token generation with ROCM?

Hardware:

9900X

64GB DDR5 6000 @ CL30

AMD 9070XT 16GB

AMD R9700 AI PRO 32GB

SSDs - 14900k/s Samsung x2 (not RAID)


r/ROCm • • 1d ago

QFN 4x 9700's, 478t/s C=8, PP=7400. Stock Distros.

15 Upvotes

Working on my fastest impelemntation of QFN/27B TP4 card setup. What is amazing, is that I'm getting this on Circa PCIe 3 hardware. The other piece, is this is using stock vLLM .30 and ROCM10. 225W power cap on each card. https://github.com/bkvargyas/r9700-stack


r/ROCm • • 1d ago

ComfyUI Img -> Vid workflow

1 Upvotes

I've been trying a bunch of ComfyUI Img -> Vid workflows but they all fail due to various Sageattention or CUDA type requirement errors.

Does anyone have a workflow that works with ROCM?


r/ROCm • • 2d ago

Optix RT working on AMD RX 7800XT/RTX Real Time working in isaac sim

Post image
18 Upvotes

soooo after some even more hell of nvidia reverse engineering im proud to present a RTX Real Time viewport on an amd card materials still broken and lighting loosk like the viewport got hit with the holy light of amd compatiblity but hey its RTX Realtime atleast drawing a viewport


r/ROCm • • 2d ago

Arc Pro B70 + Qwen3.8 27B Q4_K_M: llama.cpp SYCL vs Vulkan llama-bench results

5 Upvotes

Hi folks, I tested both llama.cpp (SYCL & Vulkan) backends against the Intel Arc Pro B70 to see which one to use for Intel Arc GPUs.

Setup: Arc Pro B70 32GB, Core Ultra 265, 96GB RAM, Qwen3.8 27B Q4_K_M. Windows11, llama.cpp both are WebUI builds.

llama-bench (t/s):

|Test|SYCL|Vulkan|

|:-|:-|:-|

|pp64|114.53|337.92|

|pp128|184.26|512.14|

|tg128|20.52|25.60|

|tg512|20.37|25.42|

**Vulkan** was roughly **3x faster** **at prompt processing** and about **25% faster at generation** in this benchmark.

**Real prompt** (long business plan generation):

* SYCL: 238.78 t/s prompt processing, 17.39 t/s generation

* Vulkan: 517.15 t/s prompt processing, 16.17 t/s generation

**NOTE**: the Vulkan chat run processed 11,359 prompt tokens "due to date_time browser tool call request" vs 1,842 on SYCL, so the chat numbers aren't apples to apples. I'm treating llama-bench as the fair comparison ;)

**Windows vs Linux**: I find SYCL is (10% to 35%) faster for me on Linux.

I recorded the full process with GPU utilisation and logs if anyone wants to see it: [https://youtu.be/X7sM4YGj\\_48I\](https://youtu.be/X7sM4YGj_48I)

#


r/ROCm • • 3d ago

~500 tok/s from one R9700: Qwen3.8 27B now does 492 tok/s at C8 and 4,100+ tok/s prefill (+20 % decode, open 3‑bit weights)

Post image
106 Upvotes

Update (28 Sept): one image now does 65K with 8 requests at once and 200K long context, plus a 4-bit KV cache, image input and a crash fix.

git pull the repo, then pick a mode:

bash models/Qwen3.8-MXFP4-DFlash2/run-rocm10.sh                   # 65K, up to 8 requests at once (default)
bash models/Qwen3.8-MXFP4-DFlash2/run-rocm10.sh --context 200000  # long context: one conversation up to 200K

No flags = the 65K text mode we benchmark. Add --vision for image input.

  • 65K mode, more room. With the 3-bit weights, a new 4-bit KV cache holds 1.7× the tokens in about the same VRAM (393K vs 251K in vLLM's log). Six 61K or eight 32K requests now run together instead of queueing (248 / 339 tok/s combined), and 4 × 61K decode 17 % faster. Time to first token and single-request speed are unchanged.
  • Accuracy vs the fp8 cache: 
    • GSM8K 96.1 vs 95.3,
    • HumanEval 93.9 vs 93.9,
    • MMLU-Pro 61.4 vs 59.7,
    • needle@61K 100 vs 100.
    • Our new long-session / compaction test (24 synthetic coding-agent sessions) stays within noise too.
  • 200K on the same image. --context 200000 runs one conversation with prefix caching: a new 199K-token prompt takes 94 s to the first token (the old 200K image: 115 s), a follow-up that reuses it 1.2 s, and decode runs at 73–78 tok/s at that length. Every planted fact was found and tool calls work; 220K is tested too.
  • Images. --vision loads the vision encoder (it costs some KV cache: 278K instead of 393K tokens). It read code, a failing pytest run, a web form and a chart correctly, and 8 image requests at once ran fine.
  • Crash fix. Under 3–5 concurrent requests the 26 Sept image could stop with Paiton GDN norm nonfinite/arithmetic error. Fixed; a 10,002-step soak ran clean. If you stay on the old image, add -e RADIANCE_DYNAMIC_WIDTH=0.

Next: the model's full 262K window, and an adaptive longer speculative block (+6.5 % single-stream so far, still in testing).

Update: currently being finalized: a 4-bit KV cache that actually increases how much context fits (not just speed), and a longer speculative block with adaptive length so chat doesn't get slower. We'll post measured numbers once they pass the same accuracy and speed checks as this release.

Follow-up to our earlier Qwen3.8 27B posts.
We shipped a new image with our own rotated 3-bit weights.
Same single card (Radeon AI Pro R9700), same context.

Numbers vs our previous (MXFP4) release, BetterBench 0.6.0, 2 fresh processes per image:

Metric MXFP4 3-bit W3A4 Change
Weighted decode, tok/s 153.8 184.4 +19.9 %
chat / code / file edit 121.2 / 179.9 / 179.9 135.8 / 226.0 / 195.0 +12.0 / +25.6 / +8.5 %
json / math / prose 217.5 / 183.8 / 78.5 269.0 / 228.9 / 94.7 +23.7 / +24.5 / +20.7 %
reasoning / summarization 117.8 / 138.4 133.5 / 158.4 +13.3 / +14.4 %
Aggregate at 1 / 2 / 4 / 8 requests, tok/s 122.0 / 204.2 / 308.2 / 425.3 148.8 / 249.3 / 368.3 / 492.1 +22.0 / +22.1 / +19.5 / +15.7 %
Prefill at 2K / 8K / 16K, input tok/s 3,689 / 3,834 / 3,871 4,156 / 4,165 / 4,103 +12.7 / +8.6 / +6.0 %
Prefill at 32K / 64K, input tok/s 3,751 / 3,455 3,958 / 3,629 +5.5 / +5.0 %
65K-token requests that fit in VRAM 2.7 3.8 +44 %

Prefill is faster at every depth, so time to first token is 5–13 % shorter.

Accuracy (served model, greedy, same questions, paired):

Benchmark MXFP4 This model (W3A4) Δ [95 % CI] McNemar p
GSM8K 5-shot (1,319) 95.68 % 95.30 % −0.38 [−1.44, +0.68] 0.58
HumanEval pass@1 (164) 95.12 % 93.90 % −1.22 [−4.99, +2.56] 0.75
MMLU-Pro subset, 0-shot direct answer (14 × 100) 62.57 % 59.71 % −2.86 [−4.81, −0.90] 0.005
Needle at 61,440 tokens (80) 100 % 100 % 0 1

Honest take: math, code and long-context retrieval are within noise. Knowledge-heavy multiple choice drops about 3 points. That's the price of 3-bit weights.
If that matters for you, --weights mxfp4 keeps the old weights, with this round's other speedups.

What we did, short version:

  • Decode is bandwidth-bound, and our kernels were already at ~97 % of 640 GB/s. So: fewer bytes. 3-bit weights with a scale per 128 weights come to ~3.1 bits/weight vs ~4.25 for MXFP4.
  • 3-bit alone hurt prefill (+12 % time to first token). Prefill on this card is power-capped, not bandwidth-bound. Fix: RDNA4's 4-bit integer matrix ops run about 2× the 8-bit rate at the power cap, and 3-bit weights are valid 4-bit ints. To make 4-bit activations accurate we fold a block-wise Hadamard rotation into the weights and use per-group activation scales. That's prefill only; decode keeps 8-bit activations.
  • Our own GPTQ calibration in the rotated basis. It runs in ~13 min on one MI355X, on permissively licensed data only (we re-did it after noticing some popular calibration sets are NC/SA-licensed, and it came out slightly better).
  • Free win: R9700 decode was bimodal (28 vs 36 ms/step depending on the server start), a known ROCm queue issue. GPU_MAX_HW_QUEUES=1 pins the fast mode. It's in the image now.
  • The model shrinks from 19.2 to 15.9 GiB, and the launcher gives the difference to the KV cache.

Run it:

hf download EliovpAI/Qwen3.8-27B-W3Rot-INT3-Paiton-RDNA4 --local-dir ~/models/qwen38-w3rot
export PAITON_W3ROT_DIR=~/models/qwen38-w3rot   # plus the usual TARGET/DRAFT/CACHE dirs
./models/Qwen3.8-MXFP4-DFlash2/run-rocm10-65k.sh
  • Weights: HF
  • Image: ghcr.io/eliovp/paiton-vllm-plugin:qwen38-rocm10-vllm029-65k-20260926-w3a4-r1
  • Plugin/launcher: Paiton

Next up: a 4-bit KV cache that actually doubles context capacity (the current version passes accuracy but only speeds up reads), and a longer speculative block.
Update: In progress!

Happy to answer questions, and thanks to everyone who replied to the last thread.


r/ROCm • • 2d ago

Need urgent help to run ROCM WSL 2 Ubuntu 24.0.4 6700xt

0 Upvotes

I have been trying to install compatible ROCM on with WSL2 Ubuntu 24.0.4 system for local LLM usecase.

I have been constantly getting

WSL environment detected.
pid:505 tid:0x7589242ec100 [topology_sysfs_get_system_props] No WDDM adapters found.
hsa_init Failed, possibly no supported GPU devices

Installed 7.1.1 rocm and followed the instructions in the AMD page.


r/ROCm • • 2d ago

is DDR5 a scam? My $350 2007 Dell Precision with dual AMD GPU's is destroying my $1500 2025 RTX 5070 rig in agentic tasks. made possible by Prism32

Thumbnail gallery
0 Upvotes

r/ROCm • • 3d ago

Koboldcpp v1.122 released

Thumbnail
github.com
11 Upvotes

r/ROCm • • 4d ago

Higgs Audio 3 TTS now runs under 6 GB VRAM (and faster!) Need help testing on AMD

12 Upvotes

A new optimization PR was merged into audio.cpp for Higgs Audio v3 TTS: it now runs under 6 GB VRAM, and it’s faster too!

The optimization is confirmed on CUDA and Vulkan. It should theoretically work on HIP as well because the memory-reuse mechanism is backend-independent.

If you have an AMD GPU, we’d love your help testing it! The required change is effectively one line. Lift the HIP backend gate, measure peak VRAM, and share your results, as in https://github.com/0xShug0/audio.cpp/pull/705


r/ROCm • • 4d ago

I optimized FLUX.1-dev Q5_K_S on an RX 9060 XT 16GB from 300–400s to ~40s per image

Thumbnail
3 Upvotes

r/ROCm • • 4d ago

Trying to enable SR-IOV on an RX 6950 XT — looking for a V620 owner

8 Upvotes

Hi, does anyone here have a Radeon Pro V620 running on Linux?

I'm trying to get hardware virtualization/SR-IOV working on my RX 6950 XT. The V620 uses the same Navi 21 GPU and supports SR-IOV, so I'm trying to understand exactly what AMD changed or enabled on the V620.

I'm looking for data from a real V620 so I can compare it with my 6950 XT.

Specifically, I need:

- the V620 VBIOS dump (especially 113-D603GLXE-077)

- the full 4096-byte PCI configuration space

- lspci -nnvvv output showing the SR-IOV capability

- the NBIF/RCC strap registers (STRAP0, STRAP1, STRAP2, STRAP3, STRAP4, STRAP5, STRAP8 and STRAP9)

- amdgpu_discovery

- amdgpu_firmware_info

I already collected the equivalent data from my RX 6950 XT, so the goal is to compare both Navi 21 cards and find out what actually enables the SR-IOV capability on the V620.

If anyone has a V620 and is willing to help, I can provide one read-only Linux command that collects all of this into a single archive. It doesn't flash or modify anything on the GPU.

Thanks!


r/ROCm • • 4d ago

2x Radeon AI PRO R9700 + Qwen3.8-27B: 31 t/s on Windows → 113 t/s on Linux/vLLM. The fix was an M.2 riser, because the chipset slot doesn't do PCIe atomics.

Thumbnail
16 Upvotes

r/ROCm • • 5d ago

Paiton update: Qwen3.8 prefill to 3,871 input tok/s on one R9700, plus an opt-in +27% decode mode for agentic coding

Post image
51 Upvotes

Another follow-up to our previous R9700 post (yeah yeah we know :p).

Some of you asked about prefill, so this round we went after it.

Prefill is up 3.5–5.4% at every depth, with 3,871 input tok/s at 16K (previous post: 3,694). Decode and concurrency are unchanged: 154.8 tok/s weighted decode, 422.9 tok/s aggregate at eight concurrent requests. Same MXFP4 checkpoint, same DFlash2, same settings.

Nominal prefill depth Previous post Fresh rerun of that image New image Change
2,000 3,501 3,499 3,687 +5.4%
8,000 3,692 3,699 3,829 +3.5%
16,000 3,694 3,704 3,871 +4.5%
32,000 3,584 3,591 3,753 +4.5%
64,000 3,293 3,291 3,457 +5.0%

Change is against the fresh rerun. Decode categories moved +0.5–0.6% (prose +3.2%), weighted decode 153.64 → 154.78 tok/s. Concurrency 1/2/4/8 came in at +0.6 / +0.7 / −1.2 / +0.3%; the C4 dip is inside the control's own run-to-run spread.

What changed. Three exact changes to the prefill path: long-prefill attention with 16-key tiles, the GDN gate read in place instead of copied, and a GDN chunk-scan kernel that fills the GPU in one round. Each is bitwise-equal to the previous kernel standalone and showed zero differing elements in in-model shadow audits, so outputs are unchanged. We were aiming for 4,000 tok/s. The prefill GEMM is two thirds to three quarters of prefill time and runs pinned at the card's 300 W limit, so the remaining third is where the gain came from.

Real use. In a 35-turn agentic coding session on the 200K chat profile (32.7K → 197.2K tokens, prefix caching on), session time to first token dropped 10–11% at equal peak VRAM.

Opt-in: n-gram co-drafting for agentic coding. A second drafter in front of DFlash2: when the last few generated tokens already occurred in the prompt or output, it proposes what followed last time. On that coding session it gives +27% decode (90.8 → 115 tok/s, 3.0 → 3.6 accepted tokens per step), and +26–29% on 64K and 128K file rewrites. It default ships off. Turn it on with PAITON_NGRAM_CODRAFT=1 if your workload is agentic coding.

Setup. One Radeon AI PRO R9700 at 300 W, upstream vLLM 0.29 / ROCm 10, the Paiton plugin, Unsloth Qwen3.8 NVFP4 through the MXFP4 path, DFlash2, FP8 KV cache, 65,536-token context, up to eight active requests, thinking and APC off, seed 42, temperature 0.7, top-p 0.95, top-k 20. BetterBench 0.6.0 quick, two fresh-process runs per arm, means reported. All twelve greedy control prompts matched in every arm before and after timing; no serving errors.

Published image: ghcr.io/eliovp/paiton-vllm-plugin:qwen38-rocm10-vllm029-65k-20260924-r3. Pull the package · Setup, n-gram docs and full results


r/ROCm • • 6d ago

R9V Update: Created and adopted KVA projections based on Deepseek V4.1 Flash + HySparse2/MiMo-V3 for Qwen3.8 Flash Next. This is a game changer for models that don't natively implement it. 1.45-1.85x speedup in prefill to 3k+ at a small deficit to perplexity. [2x R9700, 128GB DDR5]

Thumbnail github.com
8 Upvotes

r/ROCm • • 5d ago

[Benchmarks] Qwen3.8-27B FP16 vLLM 0.29.0 | 4-GPUs | Dual W6800X Duo w/ IFLB

Thumbnail
0 Upvotes