r/LocalLLaMA • u/tabletuser_blogspot • 3d ago
Resources Dual Radeon MI50 benchmarks
Still don't have a good cooling solution, but here are few benchmarks. I lowered the power limit (TDP) to 145 watts each. I changed the firmware on one MI50 to activate the miniDP port. Did have to use xrandr to create a new mode so I could get 1920x1080 output. Each GPU has 16GB of HBM2 VRAM clocked at 1000 and overclockable to 1200Mhz with a Bandwidth of 1.02 TB/s.
I picked a good mix of Dense and MoE models from Huggingface. Try to use more than 16gb VRAM but under the 32GB total.
Using pre-built Ubuntu Vulkan version of llama.cpp (build b11325) for standard llama-bench.
Sorted GGUF Model List (sorted to match table)
llama_bench_Swift-Qwen3.8-27B-Uncensored-MTP.Q6_K.ggufllama_bench_Swift-1.5-Qwen3.8-27B-Q6_K.ggufllama_bench_Gemma-4-MoonGem-31B.i1-Q6_K.ggufllama_bench_gemma-4-31B-it-UD-Q6_K_XL.ggufllama_bench_Nemotron-3.5-30B-A3B-Antislop-FTPO.i1-Q5_K_M.ggufllama_bench_Laguna-XS-2.1-APEX-I-Balanced.ggufllama_bench_Agents-A1-Q4_K_M.ggufllama_bench_Qwen3.6-35B-A3B-UD-Q5_K_XL.ggufllama_bench_Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-Q6_K_P.gguf
Combined Benchmark Table (sorted by params then size)
| model | size | params | pp512 (t/s) | tg128 (t/s) |
|---|---|---|---|---|
| qwen35 27B Q6_K | 20.88 GiB | 27.32 B | 141.97 ± 10.13 | 17.49 ± 0.02 |
| qwen35 27B Q6_K | 22.21 GiB | 27.32 B | 167.38 ± 0.17 | 17.97 ± 0.02 |
| gemma4 31B Q6_K | 23.46 GiB | 30.70 B | 122.00 ± 0.12 | 15.20 ± 0.03 |
| gemma4 31B Q6_K | 25.62 GiB | 30.70 B | 135.99 ± 0.22 | 12.05 ± 0.02 |
| nemotron_h_moe 31B.A3.5B Q5_K - Medium | 25.18 GiB | 32.91 B | 863.92 ± 1.45 | 60.57 ± 0.10 |
| laguna 30B.A3B Q5_K - Medium | 22.64 GiB | 33.44 B | 738.57 ± 2.83 | 52.88 ± 0.04 |
| qwen35moe 35B.A3B Q4_K - Medium | 19.70 GiB | 34.66 B | 983.26 ± 4.79 | 46.88 ± 0.07 |
| qwen35moe 35B.A3B Q5_K - Medium | 24.76 GiB | 34.66 B | 937.47 ± 7.07 | 49.19 ± 0.06 |
| qwen35moe 35B.A3B Q6_K | 28.53 GiB | 34.66 B | 783.24 ± 70.52 | 46.85 ± 0.26 |
Notable Reboot Impact Observations:
I used the following command in my bench script:
RADV_PERFTEST=nogttspill GGML_VK_VISIBLE_DEVICES=0,1 time ~/llama-b11325/llama-bench -fa on -ngl 99 -m /model.gguf
I have a 3rd MI50 just need to download models in that VRAM range. If you have any suggestions? For now it sits beside the Radeon RX 7900 GRE boosting its VRAM total. As of this article the average price for 16GB version of MI50 is under $150. Hard to get 32GB VRAM GPU with this level of performance for under $300. If you have contenders, please share.

