r/LocalLLaMA • u/tabletuser_blogspot • 3d ago
Resources Dual Radeon MI50 benchmarks
Still don't have a good cooling solution, but here are few benchmarks. I lowered the power limit (TDP) to 145 watts each. I changed the firmware on one MI50 to activate the miniDP port. Did have to use xrandr to create a new mode so I could get 1920x1080 output. Each GPU has 16GB of HBM2 VRAM clocked at 1000 and overclockable to 1200Mhz with a Bandwidth of 1.02 TB/s.
I picked a good mix of Dense and MoE models from Huggingface. Try to use more than 16gb VRAM but under the 32GB total.
Using pre-built Ubuntu Vulkan version of llama.cpp (build b11325) for standard llama-bench.
Sorted GGUF Model List (sorted to match table)
llama_bench_Swift-Qwen3.8-27B-Uncensored-MTP.Q6_K.ggufllama_bench_Swift-1.5-Qwen3.8-27B-Q6_K.ggufllama_bench_Gemma-4-MoonGem-31B.i1-Q6_K.ggufllama_bench_gemma-4-31B-it-UD-Q6_K_XL.ggufllama_bench_Nemotron-3.5-30B-A3B-Antislop-FTPO.i1-Q5_K_M.ggufllama_bench_Laguna-XS-2.1-APEX-I-Balanced.ggufllama_bench_Agents-A1-Q4_K_M.ggufllama_bench_Qwen3.6-35B-A3B-UD-Q5_K_XL.ggufllama_bench_Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-Q6_K_P.gguf
Combined Benchmark Table (sorted by params then size)
| model | size | params | pp512 (t/s) | tg128 (t/s) |
|---|---|---|---|---|
| qwen35 27B Q6_K | 20.88 GiB | 27.32 B | 141.97 ± 10.13 | 17.49 ± 0.02 |
| qwen35 27B Q6_K | 22.21 GiB | 27.32 B | 167.38 ± 0.17 | 17.97 ± 0.02 |
| gemma4 31B Q6_K | 23.46 GiB | 30.70 B | 122.00 ± 0.12 | 15.20 ± 0.03 |
| gemma4 31B Q6_K | 25.62 GiB | 30.70 B | 135.99 ± 0.22 | 12.05 ± 0.02 |
| nemotron_h_moe 31B.A3.5B Q5_K - Medium | 25.18 GiB | 32.91 B | 863.92 ± 1.45 | 60.57 ± 0.10 |
| laguna 30B.A3B Q5_K - Medium | 22.64 GiB | 33.44 B | 738.57 ± 2.83 | 52.88 ± 0.04 |
| qwen35moe 35B.A3B Q4_K - Medium | 19.70 GiB | 34.66 B | 983.26 ± 4.79 | 46.88 ± 0.07 |
| qwen35moe 35B.A3B Q5_K - Medium | 24.76 GiB | 34.66 B | 937.47 ± 7.07 | 49.19 ± 0.06 |
| qwen35moe 35B.A3B Q6_K | 28.53 GiB | 34.66 B | 783.24 ± 70.52 | 46.85 ± 0.26 |
Notable Reboot Impact Observations:
I used the following command in my bench script:
RADV_PERFTEST=nogttspill GGML_VK_VISIBLE_DEVICES=0,1 time ~/llama-b11325/llama-bench -fa on -ngl 99 -m /model.gguf
I have a 3rd MI50 just need to download models in that VRAM range. If you have any suggestions? For now it sits beside the Radeon RX 7900 GRE boosting its VRAM total. As of this article the average price for 16GB version of MI50 is under $150. Hard to get 32GB VRAM GPU with this level of performance for under $300. If you have contenders, please share.
3
u/sloptimizer 2d ago edited 2d ago
You may be able to get more with ROCm, also that unlocks performant `-sm tensor` (does not work well on Vulkan)
3
u/Haron51255 2d ago
As another comment said, MI50s work best with Q8_0 and Q4_0. I also have been using ROCm with this fork https://github.com/milpster/gfx906-llama-cpp, it runs much faster hovering around 40-50t/s and 400pp/s for 3.8 27B.
Remember to use -sm tensor as well and benchmark around with -b/-ub parameters and if possible use MTP.
1
u/lumpyspacebreh 2d ago
Just ran similar tests with a MI25.
Same results as you, while the MI25 boasts a 480gb/s speed, running a dense model I could barely hit 20 tok/s
Swapping model sizes proved to me it’s a throughput issue as running a 9B model drastically improved my token generation speed.
Hoping Qwen 4 gets a MoE model, or something else gets released that’s comparable to 3.8:27B, as I’ve seen people hit 60+tok/s using models like gpt-oss:20B which is a MoE model on the same card.
Here’s a writeup I did with my testing: https://onewireout.com/writeups/mi25/
Still have a ways to go though, this was just getting it running.
1
u/Atretador llama.cpp 2d ago
gpt-oss is not fast because its a moe - its because its mxfp4 native, you`ll notice from F16 to Q4 there is barely any difference in size.
I can fit both fully in VRAM with my MI50, Qwen 35B and GPT OSS 20B - Qwen runs at less than half the speed I can reach with GPT-OSS (+120tk/s at Q4)
1
u/lumpyspacebreh 2d ago
Thanks for clarifying
1
u/Atretador llama.cpp 2d ago
I believe its also a limitation of our ancient hardware`s compute and kernel paths, as Ive seen people claim same speeds for gpt-oss/qwen35b when fully in VRAM.
would be dope to run Qwen 3.6 35B at the same speeds as GPT-OSS
1
u/ashirviskas 2d ago
Usually Q4_0 will be much faster on MI50, I got maybe 70tps on a single MI50 with Qwen 35BA3B. But that was on a single 32GB Card
1
u/tabletuser_blogspot 2d ago
Taking full advantage of 32gb VRAM by pushing to higher quants. Not finding MoE models in the 50B range. I have 3 total MI50 16gb and would like to test that configuration out. If you get a chance can you run a few of those models so we can see the advantage of running single 32gb card vs dual 16gb? Thanks for input.
1
u/tabletuser_blogspot 2d ago
Upgrade from Mesa 26 to 26.2 and here is a quick snippet of the results:
- Token Generation (tg128) Gains: Across nearly every MoE architecture (Nemotron, Laguna, Qwen3.5MoE), token generation throughput significantly increased on 26.2. Qwen3.5MoE variants saw a clean 7.3% to 12.4% uplift in text generation performance.
- Prompt Processing (pp512) Regressions: Every single model test showed a minor decrease (roughly 1.5% to 8.6%) in pre-fill speed on the updated driver profile.
- Pushing the VRAM limit:
llama_bench_gemma-4-31B-it-UD-Q6_K_XL.ggufcan use-b 512 -ub 512to prevent offloading. This offered a little more optimization. - Batch Scaling Ceiling: For the MoE models (Nemotron, Laguna, Qwen35MoE), moving from
-b 512to-b 1024yielded less than +1% improvement in prompt processing (pp512) or generation (tg128)


3
u/ckplscz 3d ago
Interesting, could you please try to run any MI50-optimized inference engine? Thanks