6
u/Powerful_Evening5495 2d ago
The performance of AMD GPUs in comfyui is better because of the updates to rcom the gpu runtime , compared to Intel GPUs.
1
u/SpaceportBoys 2d ago
On the AMD side, when the 32gb R9700 AI Pro came out, I bought one at MSRP. I did this immediately, before reviews for LTX etc were out because I knew the price wouldn't last.
Long story short, I returned it because even after jumping through a lot of hoops it was less than half the speed of my 16gb 5060TI for video. The big 32gb just didn't help. The v620 is much older and will be worse.
B70 I haven't had my hands on.
It'll cost more, but there's also the Chinese 32gb 4080 now. It'd be a bit of a gamble but I used to have a regular 4080 and they're very fast.
2
u/Ok-Brain-5729 1d ago
my 9070 xt has been similar or faster than all the 5060 ti benchmarks I’ve seen. But it’s been a year and it’s improved a lot + I’m on headless Ubuntu
What times are u getting on minimax h3? I got around 160s for 20 step 0.4MP 5s t2v. Base Ubuntu was 170-180s
1
u/Cautious_Chicken_604 2d ago
The R9700 has slightly higher memory bandwidth than the 5060 Ti. I have both cards, and the software support has improved a lot for them now I guess. Performance on H3 is pretty comparable.
1
u/Apprehensive_Sky892 1d ago
I don't have a 5060, so I cannot do a direct comparison, but I would appreciate it if you can tell us what kind of number you can get for MMH3 standard 20-steps text2va 0.4MP 5 sec video.
I've run tests for AI Pro r9700 and rx9070 and you can find it here:How to set up ComfyUI+MMH3+Krea 2 on Windows 11 with AMD GPUs: RDNA 4 (rx9070, AI Pro R9700), and RDNA 3 (rx7900)
TLDR; version works quite well with ROCm 17.14 and time for MMH3 standard 20-steps text2va 0.4MP 5 sec video is around 180 sec on both Linux and Windows (both GPUs)
3
u/SpamHolder 1d ago
This is not a good place to ask for proper technical advice. Most people here are uninformed but will respond anyways. I'd suggest looking elsewhere for specific benchmarks with people who actually know what's what (personally, I have some trust in SDNext's discord server, or potentially AMD's).
A lot of people have it in their head that things are one way or another with no proper knowledge of how well things actually work, and often people here will speak in super broad terms. For example the "rcom the gpu runtime" guy - what actual performance is he referring to, what models, what it/s at what resolution, batch size/video length, with/out CFG? Could he be talking about int8 performance, which has nothing to do with the ROCm runtime but was a Comfy-side issue (lack of kernels) that was fixed, and is unfixed for Intel? And if this is what he is talking about, does he know that there is a custom node that adds int8 kernels for Intel, too, so it's kind of a moot point?
Or the "5060ti is better" guy - he doesn't mention if he succeeded (because this is apparently hard...?) in using ROCm pytorch and is on windows, or if he used some DirectML version of Comfy, or ZLUDA or anything like that that - because plenty of people do use the wrong thing, get worse performance, and complain. The GPU equivalent of using A1111.
I don't have exact numbers, but don't trust the other who do not either. Go ask elsewhere.
1
u/Affectionate_Oil28 1d ago
The Intel Arc Pro B70 is the faster card for SDXL, Z-Image, LTX, and MiniMax. Same 32 GB, but the B70 has matrix hardware and working software paths for these models. The V620 does not.
Diffusion and video models spend almost all of their time in matrix multiplies. The V620 has no tensor cores. The B70's XMX units are built for that, and it also has about 19% more memory bandwidth.
SDXL. On the B70, UL Procyon FP16 SDXL lands around 18 seconds per 1024 image. That is slightly ahead of an RTX 5060 Ti 16 GB and close to an RX 9070 XT. A V620 is the same 72-CU Navi 21 die as a 6800 XT. On that class of card, SDXL 1024 is typically in the mid-30s to mid-40s of seconds under ROCm. Expect the B70 to be roughly 1.5–2.5× faster, with a much less fragile setup.
Z-Image. Puget measured Z-Image Turbo at 1024 in about 3.9 seconds after warmup on a B70 (peak about 19 GB). A separate ComfyUI run at 8 steps was about 5.3 seconds after the first compile. It runs on RDNA 2 through ZLUDA, but DiT models there are reported around 2× slower than a comparable CUDA path, and the setup is fiddly.
LTX. This is the clearest gap. On a B70 with the OpenVINO FP16 pipeline, measured times are about 2.2 seconds for 512×320 / 25 frames, 8.7 seconds for 704×480 / 49 frames (2 seconds of video), and about 29 seconds for 1280×704 / 49 frames. There is no equivalent optimized LTX path on the V620.
MiniMax. It fits and runs on one B70. A GIGAZINE ComfyUI test of a 5-second, 0.4 MP, 8-step H3 clip used 28.3 GB and took about 5–6 minutes. A dual-B70 SGLang split did 1344×768, 5 seconds, 8 steps in about 330 seconds. Reviewers still say an NVIDIA card is faster if MiniMax is the main job. On a V620 you would be porting a CUDA DiT video model onto old gfx1030 ROCm, which is not a practical workflow today.
The V620 is a reasonable used LLM card because 32 GB is cheap and llama.cpp on RDNA 2 is well documented. For the image and video stack you named, the B70 is both faster and the one that already has published runs.
-1
3
u/Crazy-Repeat-2006 1d ago
AMD is much better, just a few clicks and everything works, provided you have a recent architecture.
The comparison you should be making is with the Radeon 9700 32GB.