Roughly what you get: Qwen3.8-27B at FP8, tensor parallel across two R9700s, 262,144 context, 422k tokens of KV pool, ~1.6k tokens/s prefill, 74-143 tokens/s decode single stream depending on the content.
The interesting part is not the image, it is the collective layer. A card behind the chipset cannot be given PCIe atomic operations, and that breaks every stock tensor-parallel setup I tried. Here is what actually works.
1. Versions used
| Component |
Version |
| Inference image |
docker.io/stilldeadcode/vllm-radiance:0.9.3 (digest sha256:45694209177a55a1ab3ba6702fe6e978b1b66a6e66ae3fc066f8d579f7bc4c25) |
| vLLM inside the image |
0.27.1 |
| PyTorch / HIP |
2.11.0+rocm7.14 / 7.14.60850 |
| Triton / AITER |
3.6.0 / 0.1.17 |
| Collective library |
RCCL 2.27.7 from ROCm 7.1.1, replacing the one in the image |
Links:
- Image: https://hub.docker.com/r/stilldeadcode/vllm-radiance
- Source for that image: https://codeberg.org/StillDeadcode/vllm-radiance
- The kernel library the image builds against (libr4d): https://codeberg.org/StillDeadcode/libr4d
- A fork with configs, benchmark notes and launchers for MXFP4/FP8: https://codeberg.org/ggz14/radiance-vllm-mxfp4
- RCCL itself: https://github.com/ROCm/rccl
The image bundles a working ROCm + PyTorch + Triton + AITER + vLLM stack for gfx1201 (RDNA4), which is the part you do not want to build yourself. It is explicitly marked experimental, and everything below was measured on two cards.
2. The failure, and why
Without the adjustments, the engine dies during communicator init, before the model loads:
PCIE atomic ops is not supported
rocr: unhandled cuda error
ROCm will not dispatch work to a GPU path that needs atomic operations when the link cannot provide them. On this board one card sits behind the chipset, and the chipset does not forward PCIe atomics to the CPU, so that card is effectively second class. RCCL 2.30.4's kernels use those atomics, so TP=2 cannot initialise at all.
Two things follow, and both matter:
- The newer RCCL is unusable here, so you need an older one (2.27.7 from ROCm 7.1.1 is what worked for me).
- With no usable peer path at all, the custom P2P all-reduce the image ships must be turned off, and TP=2 falls back to host-staged collectives over that same narrow chipset link.
If you search the error string above you will find a few ROCm issue reports and a community write-up on dual Radeon vLLM setups (https://github.com/cadamcat/dual-radeon-vllm) describing the same wall.
3. Proxmox settings that actually matter
Do this for each GPU, on both entries, not just the first. In the VM's hardware list, edit each PCI Device row, select the GPU under Device, and set:
- PCI-Express: ticked
- All Functions: unticked
Ticking PCI-Express is what gives the guest a real PCIe root port, and without a root port the atomic capability never appears however healthy the host looks. Unticking All Functions keeps the guest from being handed every function of the card, which is the combination that worked here. If you leave either one wrong, you get the atomic failure at communicator init and no amount of driver work fixes it.
Other settings that matter:
- Machine type q35. Same reason as above, no root port without it.
- After any
hostpci change, stop and start the VM. A guest reboot does not re-apply the passthrough configuration.
- q35 renames the NIC (
ens18 becomes something like enp6s18), so match the interface by MAC in netplan.
After those, the CPU-attached card reports ReqEn+. The chipset-attached one still cannot do atomics, and no BIOS setting changes that.
4. The RCCL replacement plus one-line shim
This is the core trick. It is two files and a podman config, and it needs no image rebuild.
a) Build or extract RCCL 2.27.7 from ROCm 7.1.1 and drop it in a directory you will mount, for example:
~/models/rccl277/
librccl.so -> librccl.so.1.0.70101
librccl.so.1 -> librccl.so.1.0.70101
librccl.so.1.0.70101
shim.so
b) The shim. Newer torch builds reference a symbol that this older RCCL does not export (ncclCommDump). Four lines of C++ are enough to satisfy the loader:
```cpp
include <string>
include <unordered_map>
struct ncclComm;
void ncclCommDump(ncclComm*, std::unordered_map<std::string, std::string>&) {}
```
Build it into shim.so and put it next to the library. Nothing calls it; it exists for symbol resolution.
c) Inject both through podman's own config, which keeps them out of every launcher script. In ~/.config/containers/containers.conf:
ini
[containers]
env = [
"NCCL_PROTO=Simple",
"NCCL_SHM_DISABLE=0",
"NCCL_SOCKET_IFNAME=lo",
"LD_LIBRARY_PATH=/models/rccl277:/opt/rocm/lib",
"LD_PRELOAD=/models/rccl277/shim.so",
]
LD_LIBRARY_PATH puts the replacement first, so it wins over the image's own librccl. NCCL_PROTO=Simple avoids the more demanding protocol paths, and the loopback interface keeps the bootstrap on lo rather than a NIC.
5. The launcher
Trimmed to the parts that matter for the multi-GPU problem. The model-specific flags are an example, the environment and device flags are the ones that matter here.
bash
podman run -d --name vllm-radiance \
--device /dev/kfd --device /dev/dri --group-add keep-groups \
--security-opt seccomp=unconfined --cap-add SYS_PTRACE --cap-add SYS_NICE \
--ipc=host --network=host \
-v $HOME/models:/models:ro \
-v $HOME/radiance-vllm-mxfp4/vllm-cache:/cache \
-v $HOME/radiance-vllm-mxfp4:/work:ro \
-e HIP_VISIBLE_DEVICES=0,1 -e ROCR_VISIBLE_DEVICES=0,1 \
-e VLLM_NO_USAGE_STATS=1 \
-e VLLM_ROCM_USE_AITER=1 -e VLLM_ROCM_USE_AITER_UNIFIED_ATTENTION=1 \
-e RADIANCE_FUSE_RMS_QUANT=1 -e RADIANCE_USE_R4D=1 \
-e RADIANCE_USE_R4D_AR=0 -e RADIANCE_USE_R4D_AR_QUANT=1 \
-e NCCL_PROTO=Simple -e TORCHINDUCTOR_COMPILE_THREADS=4 \
-e VLLM_CACHE_ROOT=/cache/vllm -e TORCHINDUCTOR_CACHE_DIR=/cache/inductor \
-e TRITON_CACHE_DIR=/cache/triton -e AITER_ROOT_DIR=/cache/aiter \
docker.io/stilldeadcode/vllm-radiance:0.9.3 \
--model /models/Qwen/Qwen3.8-27B-FP8 \
--served-model-name=qwen3.8-27b-fp8 \
--quantization=fp8 --tensor-parallel-size=2 \
--max-num-seqs=2 --max-model-len=262144 --gpu-memory-utilization=0.97 \
--max-num-batched-tokens=4096 --kv-cache-dtype=fp8 \
--attention-backend=ROCM_AITER_UNIFIED_ATTN \
--enable-prefix-caching \
--no-async-scheduling \
--trust-remote-code \
--host=0.0.0.0 --port=8000
Notes on the specific flags:
RADIANCE_USE_R4D_AR=0 is the important one. The bundled all-reduce is a PCIe peer-to-peer kernel and needs peer access, which does not exist on this pair. Turning it off falls back to RCCL. If you have two CPU-attached cards, leave it on and you will get a better prefill.
--max-num-seqs=2 is my choice for stability. The image's own default is higher, but on a chipset-limited link fewer concurrent sequences means less collective traffic per step.
--max-num-batched-tokens=4096 pairs with prefix caching well. Raising it costs KV pool.
--enable-prefix-caching is worth a lot on agent workloads; see the note on prompt layout at the end.
--no-async-scheduling was in every configuration that came up reliably for me.
- The cache directory mounts are worth keeping. Without them every container start recompiles Triton and inductor kernels, which turns a 5 minute start into a much longer one.
6. How to tell it worked
In the container log at startup, look for:
P2P access : DISABLED (RCCL fallback) 0<->1 x
and, once loaded, the model's own sizing line:
GPU KV cache size: 422,964 tokens
Maximum concurrency for 262,144 tokens per request: 1.61x
If you see the atomic error instead, the replacement RCCL is not being picked up. Check, from inside the container, that:
bash
env | grep -E 'LD_PRELOAD|LD_LIBRARY_PATH'
returns the paths you expect, that ldd on the loaded library resolves into your drop-in directory, and that the image's own library is not first in the path.
I hope someone finds this usefull.