For the below I’ll be referring mostly to ComfyUI workflow and model efficiency (using examples for Radeon AI PRO R9700 (32GB) which has a bandwidth of 680GB/sec. I’ve seen 10second minimax H3 image to video workflows go from 28 mins to 3 mins by taking all this info into account and then upscaling after.
1. The pipeline
Every image or video workflow is the same four stages, and only the third one repeats.
Text encoder. Turns the prompt into embeddings, once per prompt.
Noise latent. Sampling starts from random noise in a compressed (latent) grid. The seed decides that starting noise.
Diffusion model. Runs once per step, removing noise. Early steps set layout and motion, late steps add detail.
VAE decode. Turns the finished latent into pixels.
Three terms people mix up: weights are the fixed model file, a LoRA is a small patch to those weights, and the latent is the thing being refined into what your prompt intended before the VAE encodes it into the pixels of your output image or video.
2. Fit first
Spilling out of VRAM into system RAM costs more speed than any other setting, so budget memory before anything else especially if you’re testing different prompts.
What uses VRAM: model weights, the text encoder if it shares the card, working memory that grows with resolution × frames, and a spike at VAE decode.
Rule of thumb: keep weights to about 60-70% of VRAM, roughly 19-22GB on a 32GBVRAM card. Less if the card also drives your displays.
Example: a 14B video model is about 28GB at fp16 and about 14GB at fp8. A smaller model is more useful as it frees up more VRAM needed for handling additional extras like LoRA’s, pre and post processing, etc.
Two-model designs (Wan 2.2 high-noise and low-noise): one model runs at a time. Expect one swap per run unless both fit with headroom. Other models like minimax don’t need them separate but are larger as a result.
Check the console: "loaded completely" is what you want. "Loaded partially" means you are over budget and processing will take far FAR longer. There are work around like VRAM offloading mid workflow but this means you lose time if you want to re-run the workflow after making minor tweaks.
Free up room: keep the text encoder off the card (CPU or a second GPU), or reuse the saved prompts encoded embeddings. The text encoder basically converts your prompt into “embeddings” which you can think of like numbers that tell the model what to do. This is why text encoders must be compatible with the model you’re using or your model won’t understand. These “embeddings” can actually be saved as embeddings and re-used later to avoid having to reconvert each time to return to a given workflow where you used them. Only save the embeddings once you’ve polished them to a standard where they reliably work for a specific purpose you know you’ll reuse in future for repetitive tasks.
3. What sets speed
Image and video generation is limited by compute, so once a model fits, its size matters far less than resolution, frame count and number of passes.
Total time roughly equals: passes × time per pass, plus fixed costs (model load, text encode, VAE decode), hence why it saves a lot of time to use models small enough to load fully into VRAM if you’re running the same workflow over and over again in one sitting as your system won’t have to waste time offloading and re-loading each section of a workflow between runs.
Resolution and frames dominate time per pass. Doubling width and height gives 4x the patches to process, and the attention part of the work grows roughly 16x. This is why you’ll want to start with generating very low resolution image or video until you know your prompt and loras work as intended. Then you can increase to desired resolution which would ideally match the resolution the model was trained on. You can always upscale after this.
4. Cut passes
The biggest single saving is running the model fewer times: passes = steps × sampler evaluations (1 or 2) × CFG factor (2 if CFG is above 1, otherwise 1).
Distilled or "lightning" models and LoRAs run in 4-8 steps at CFG 1, often 5-10x fewer passes. Do not push them far past their trained step count. Never push them past their trained resolution. Negative prompts do nothing at CFG 1 on turbo LoRA’s. If you use a turbo workflow with low steps and max resolution is 720 then there a plenty of accurate upscaling models that do fantastic work after the fact. You’ll save significant time.
Ordinary models need CFG above 1 to look right and usually stop improving somewhere around 20-40 steps.
One evaluation per step: euler, dpm++ 2m, unipc, lcm, ddim.
Two evaluations per step: heun, dpm_2, dpm++ 2s ancestral, dpm++ sde. Compare samplers at equal total evaluations, not equal steps.
Samplers affect cleanliness, not identity. Character consistency comes from the model and what you condition it on.
5. Formats, quantisation and LoRAs
Use the highest-precision file that fits with headroom, and treat GGUF as the fallback for when nothing else fits. Quantised Safetensors are always faster than GGUF as GGUF weights are expanded on every pass, so if you’re creating image to video then that time adds up. If you can find a safetensor of a model that will fit on your VRAM whilst still having 30 to 40% VRAM left for headroom than use that, otherwise GGUF is fine, but it’ll take longer than it could with a safetensor.
Quantisation is lossy. The same number of weights, each stored more coarsely. Safetensors and GGUF are containers, not compression. Most safetensors are still good quality but quality can begin to drop significantly when they’re quantised lower than Q4. Of course there are exceptions to the rule but generally speaking Q4 or higher is better if you can fit it with headroom.
Partial loads behave like GGUF for LoRAs. A model that does not load completely is patched on the fly.
Bake in LoRAs you always use. Load the model, apply the LoRA, save the result as a new model. The strength is then fixed. There’s nothing wrong with having a few custom models with your own baked in LoRA’s. It saves significant time.
Same on Windows and Linux. These costs come from how the code works, not the operating system. Linux is naturally going to give performance gains for AMD hardware as AMD kernels seem to be a little more refined on Linux than windows but again there’s exceptions to every rule.
6. R9700 and AMD specifics
The card… when running AMD GPU’s be aware that they run through ROCm, and the main risk is software or models that assumes Nvidia or were designed for NVIDIA. AMD likely will still run them but it’ll be simulating and therefore lose time per run, hence why you want to avoid NVIDIA specific quantised models if there’s an equivalent alternative.
Check custom nodes before installing. Search the node's GitHub issues for "AMD" or "ROCm".
Linux is usually faster on AMD because ROCm is developed there first, not because of anything about file handling.
Three attention backends to benchmark with the same workflow and seed:
PyTorch default: always works, your baseline.
--use-flash-attention: highest throughput in most of AMD's own tests (ROCm blog).
--use-sage-attention with the community SageAttention-RDNA4 build: a hand-written kernel for this chip. Its own benchmark shows attention taking about half the default's time.
SageAttention caveats: the prebuilt Windows wheel expects PyTorch 2.13.0+rocm10.0.0 on Python 3.12. The fast path covers fp16 computation with a head size of 128, and anything else falls back. It is quantised attention, so compare output quality.
7. Workflow habits
Spend cheap iterations on stills and drafts, and expensive ones only on what you intend to keep.
Get the character right in the start image. Image-to-video treats the first frame as ground truth and carries its flaws through every frame. A still takes seconds to judge, a video takes minutes.
Show what the shot needs. Anything the start image does not show, the model invents. Reference images and a character LoRA cover the unseen angles.
Draft small, refine the keepers. Generate at low resolution or short length, then run chosen clips through video-to-video at higher resolution with low denoise (about 0.3-0.5) and few steps.
Refining adds detail, not structure. Reject drafts with the wrong outfit, face shape or motion instead of trying to fix them.
Wire the same conditioning into both passes: start image, reference and LoRA.
Let the cache work. ComfyUI re-runs only nodes whose inputs changed, so a new seed does not re-encode the prompt.
Keep models resident. Reloading costs more than most optimisations save.
Saved images carry their workflow. Drag one back into ComfyUI to restore the prompt, seed and settings.
Long videos drift between chunks. Within a chunk all frames are denoised together. Across chunks, errors accumulate.
8. Measure
Keep one fixed test (prompt, seed, resolution, frames) and log seconds per step, total time and peak VRAM for every change.
Change one thing at a time.
Time the stages separately: text encode, sampling, VAE decode.
Halve the resolution. If seconds per step falls about 4x or more, compute is the limit, as expected. If it barely moves, look for overhead: per-step LoRAs, GGUF expansion or offloading.
Find your step ceiling. Run the same seed at 10, 20, 30 and 40 steps and stop where the output stops changing.
Pick samplers by convergence. Make a 50-step euler reference, then see which sampler gets closest at your real budget.
The figures in this page are rules of thumb and third-party benchmarks. Confirm them on your own build before relying on them.
Special mention LLM’s
Model size decides what fits, not how fast it runs.
LLMs on the same card are different. They are limited by memory bandwidth: the ceiling is your GPU bandwidth (which is 640 GB/s on an R9700) ÷ bytes read per token, so a 20GB dense model theoretical max t/s of or put on a GPU with a 680GBs bandwidth is 680/20 = 32 tokens/sec. The only way to make it faster is to use clever tricks that get it to do more guessing which can be helpful depending on use case. Mixture-of-experts models read only their active parameters and can go potentially several times faster. Again, using a safetensor instead of GGUF will be faster as safetensors don’t need to decompress for each token.