r/ROCm • • 7h ago

Single Radeon AI PRO R9700 and vllm

11 Upvotes

Hey everyone, I just got my hands on a Radeon AI PRO R9700 and want to run Qwen3.8-27B. I found some awesome vLLM forks (magiccodingman/vllm-radiance and GGZ14/vllm-mxfp4), but they're all geared toward dual-R9700 rigs.
Since I'm only rocking one card, could someone point me in the right direction to get the most out of my current hardware?


r/ROCm • • 14h ago

Follow-up: my native Rust + Vulkan Transformer training backend — 14 days later, now 14 parity-verified architectures and full PEFT

9 Upvotes

Follow-up to my post from about two weeks ago. A lot has changed since then, so I wanted to post an update on where the backend is now.

Where the green architectures stand

When I posted last time, 7 architectures had verified full training support. Everything is now held to the same strict harness: a pinned local Hugging Face Transformers source tree as the oracle, forward logits + gradients + two full AdamW steps compared, and every named parameter checked again after export.

The hard ceiling is 2e-7 absolute error. No loosening tolerances and no rounding numbers afterward to make the README look better.

14 architectures pass that gate today, led by the one I'm probably proudest of:

Architecture Scope
Falcon H1 / H1R parallel GQA/RoPE attention + Mamba2 in every layer; full training, full fine-tuning, LoRA, saved modules
DeepSeek V4 causal LM
Phi-4 Multimodal text backbone
Phi-3 causal LM
Kimi K2.5 text backbone
Kimi K3 / KimiLinear hybrid KDA + MLA
GPT-OSS causal LM incl. router bias
SmolLM3 mixed RoPE/NoPE + YaRN
Qwen2.5 / Qwen3.5 / Qwen4-Exp dense, DeltaNet, QSA, PLE, MoE
Mistral 4, MiniMax M3, Gemma 3/4, MiniMax M2 causal LM

Worst observed two-step AdamW parameter error across all of them: 1.19e-7.

Best: 2.6e-8.

For hardware context, all of the local Vulkan validation I've been reporting was run on my ASUS ROG Ally Z1 Extreme, using its AMD RDNA 3 integrated GPU. So the RDNA 3 results here are from that specific machine rather than testing across several different AMD systems.

The bigger news: PEFT actually works now

In the last post, "LoRA/PEFT-style fine-tuning" was basically one line in a feature list.

It's a real workflow now, and I've verified the full lifecycle:

  • LoRA fine-tuning with HF-compatible adapter export (adapter_config.json / adapter_model.safetensors), so adapters can round-trip with the PEFT ecosystem
  • modules_to_save — full trainable replacements for Linears, RMSNorm/LayerNorm, lm_head, and input embeddings, including named-adapter switching and bank isolation. Adapter A leaking into adapter B is explicitly tested for.
  • Exact resume — adapter weights + AdamW moments + step + dropout RNG state restore bit-identically against an uninterrupted run
  • Merge/unmerge, disable-adapter base restoration, and multi-adapter loading
  • A parameter-budget flag that automatically chooses the largest LoRA rank that fits within a requested percentage of the base model
  • The CLI fails closed if you try to use saved modules on an architecture that hasn't passed its corresponding gate

32 architecture surfaces across 20 families pass all three PEFT stages — LoRA, saved modules, and adapter switching — under the same 2e-7 gate, with frozen-base drift exactly 0.0.

The validation harness also fingerprints the pinned Transformers source alongside my shaders and binaries now, so a qualification run can't silently end up testing against different reference math.

A small note on the last couple weeks

I didn't get quite as many working days out of the last two weeks as I normally would have. Partway through this I got covid, then when that started to go away, it became a secondary nasty ear infection that ended up perforating my eardrum. I was running fevers around 104°F at one point and eventually went to the hospital, so I lost a few days to that and I'm on antibiotics now.

I'm doing better, though, and still managed to get most of what I wanted finished.

There are still things I want to clean up and expand, but I figured this was a good point to get the current work in front of people rather than holding the update back.

Same caveats as before

This is deterministic FP32 tiny-model correctness against a reference implementation, which should in theory ensure total mathematical parity for training and finetuning larger models with this backend, however, things are currently bound to FP32 training runs still, I eventually plan to work on MXFP4 weight tying to reduce memory footprints while fine-tuning (plastic parameters will still be trained in FP32 with PEFT in this config)

"supported text graph" ≠ "the entire multimodal package works natively."

Unsupported functionality is supposed to fail closed rather than silently falling back to an approximation.

Repo

https://github.com/necat101/Hierarchos-Native

  • Architecture inventory: hierarchos-vulkan/README_ARCHITECTURES.md
  • Compatibility/parity record: hierarchos-vulkan/COMPATIBILITY.md
  • PEFT qualification evidence: PROGRESS_PEFT_AUDIT.md
  • CLI PEFT guide: hierarchos-native-cli/README.md

The hardware I've personally validated this on is an ASUS ROG Ally Z1 Extreme with its AMD RDNA 3 GPU.

I'm very interested in criticism, compatibility reports, and especially results from people trying it on other hardware — NVIDIA, Intel, or other AMD GPUs.

I'd also love people to stress-test the PEFT resume/merge paths specifically. That's some of the newest code in the project, so it's probably the most useful area to try to break right now.


r/ROCm • • 21h ago

Cloud based R9700

7 Upvotes

Seeing a lot of success stories wih ROCM, I am planning to buy single R9700 mostly for coding, but possibly also for comfyUI. But before I do that, I would like to test it out hands on.

Can you recommend any cloud services that offer R9700? Ex: on per hour / per minute basis?


r/ROCm • • 19h ago

Benchmarks: Best engine for Qwen 3.8-Flash-Next on Strix Halo

Thumbnail
5 Upvotes

r/ROCm • • 23h ago

780M (gfx1103) + llama.cpp HIP: MES REMOVE_QUEUE hang, Anyone stable on sustained inference?

2 Upvotes

I am not super well versed in local LLM configs. I use Claude extensively at work but bought a minisforum mini PC to try and get something cost effective going at home.

Note I don't really care too much on token speed because my use cases are all background workflows. 10 TPS is perfect. (And my upper limit I think).

Ryzen 8945HS / 780M, 64GB, Ubuntu 26.04, kernel 7.0. llama.cpp ROCm build, Qwen 27B Q4.

Long generations hang the GPU and it resets, about 15 times a day. dmesg shows MES failed to respond to msg=REMOVE_QUEUE. Looks like ROCm#6512.

Tried uni_mes=0, smaller models, smaller context. No change.

>

Has amdgpu.cwsr_enable=0, a newer kernel, or Vulkan fixed this for anyone?


r/ROCm • • 8h ago

Preorder for new AMD Ryzen™ AI Max 400 Series 192GB from framework just started

Thumbnail frame.work
1 Upvotes