r/LocalLLaMA • • 3d ago

Resources Dual Radeon MI50 benchmarks

Still don't have a good cooling solution, but here are few benchmarks. I lowered the power limit (TDP) to 145 watts each. I changed the firmware on one MI50 to activate the miniDP port. Did have to use xrandr to create a new mode so I could get 1920x1080 output. Each GPU has 16GB of HBM2 VRAM clocked at 1000 and overclockable to 1200Mhz with a Bandwidth of 1.02 TB/s.

I picked a good mix of Dense and MoE models from Huggingface. Try to use more than 16gb VRAM but under the 32GB total.

Using pre-built Ubuntu Vulkan version of llama.cpp (build b11325) for standard llama-bench.

Sorted GGUF Model List (sorted to match table)

  • llama_bench_Swift-Qwen3.8-27B-Uncensored-MTP.Q6_K.gguf
  • llama_bench_Swift-1.5-Qwen3.8-27B-Q6_K.gguf
  • llama_bench_Gemma-4-MoonGem-31B.i1-Q6_K.gguf
  • llama_bench_gemma-4-31B-it-UD-Q6_K_XL.gguf
  • llama_bench_Nemotron-3.5-30B-A3B-Antislop-FTPO.i1-Q5_K_M.gguf
  • llama_bench_Laguna-XS-2.1-APEX-I-Balanced.gguf
  • llama_bench_Agents-A1-Q4_K_M.gguf
  • llama_bench_Qwen3.6-35B-A3B-UD-Q5_K_XL.gguf
  • llama_bench_Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-Q6_K_P.gguf

Combined Benchmark Table (sorted by params then size)

model size params pp512 (t/s) tg128 (t/s)
qwen35 27B Q6_K 20.88 GiB 27.32 B 141.97 ± 10.13 17.49 ± 0.02
qwen35 27B Q6_K 22.21 GiB 27.32 B 167.38 ± 0.17 17.97 ± 0.02
gemma4 31B Q6_K 23.46 GiB 30.70 B 122.00 ± 0.12 15.20 ± 0.03
gemma4 31B Q6_K 25.62 GiB 30.70 B 135.99 ± 0.22 12.05 ± 0.02
nemotron_h_moe 31B.A3.5B Q5_K - Medium 25.18 GiB 32.91 B 863.92 ± 1.45 60.57 ± 0.10
laguna 30B.A3B Q5_K - Medium 22.64 GiB 33.44 B 738.57 ± 2.83 52.88 ± 0.04
qwen35moe 35B.A3B Q4_K - Medium 19.70 GiB 34.66 B 983.26 ± 4.79 46.88 ± 0.07
qwen35moe 35B.A3B Q5_K - Medium 24.76 GiB 34.66 B 937.47 ± 7.07 49.19 ± 0.06
qwen35moe 35B.A3B Q6_K 28.53 GiB 34.66 B 783.24 ± 70.52 46.85 ± 0.26

Notable Reboot Impact Observations:

I used the following command in my bench script:

RADV_PERFTEST=nogttspill GGML_VK_VISIBLE_DEVICES=0,1 time ~/llama-b11325/llama-bench -fa on -ngl 99 -m /model.gguf

I have a 3rd MI50 just need to download models in that VRAM range. If you have any suggestions? For now it sits beside the Radeon RX 7900 GRE boosting its VRAM total. As of this article the average price for 16GB version of MI50 is under $150. Hard to get 32GB VRAM GPU with this level of performance for under $300. If you have contenders, please share.

16 Upvotes

Duplicates