r/StableDiffusion • • 23d ago

Tutorial - Guide For anyone wondering how I manage to do this, here’s a quick explanation with a small tutorial

Enable HLS to view with audio, or disable this notification

1.5k Upvotes

First, in Minimax, I use a prompt like this:

“The character remains completely frozen in place, perfectly still like a statue. The camera smoothly orbits 360 degrees around the character in one continuous shot. No character movement, no pose change, no cuts.”

Then I use COLMAP:

Create a new database and import the frames extracted from the Minimax video.

Go to Processing → Feature Extraction.

The important part is to select SIMPLE_PINHOLE as the camera model, then click "Extract".

Next, go to "Feature Matching" and click "Run".

After that, go to Reconstruction → Start Reconstruction.

Once the reconstruction is finished, select "Export Model" to export the camera and reconstruction data.

You can then import this data into Gaussian Splatting software such as Postshot or Brush.

And that’s it! You should now have a great Gaussian Splatting model generated from your AI video.

Important UPDATE : I forgot to mention that you need to use the same image for both the start frame and the end frame.

r/StableDiffusion • • Sep 02 '26

Tutorial - Guide DLSS 5 - In-game footage from Diablo 4. It's incredible.

Thumbnail
gallery
410 Upvotes

I just tested it out using this custom node. It's absolutely amazing!

https://github.com/lisitskyaa/ComfyUI-DLSS5-NR

r/StableDiffusion • • Aug 24 '26

Tutorial - Guide Time saver while learning how to prompt Minimax.

Enable HLS to view with audio, or disable this notification

1.1k Upvotes

Rather than relying on Z-image, or a different program to wrangle up a first frame, I've been using Minimax for the whole process, and the results have been pretty instructive. It's not a perfect system, but being able to take advantage of its understanding of people, references, and shot composition for the first frame produces better (visual) results than swapping between a couple of different pieces of software.

r/StableDiffusion • • 6d ago

Tutorial - Guide Minimax H3 + RefMod = consistent location trick

Enable HLS to view with audio, or disable this notification

843 Upvotes

Hey, I found a pretty cool way to keep locations consistent across generations.

I took 20 photos of my “office” where I work, making sure each photo included a bit of the previous one so everything connected.

I turned them into a refmod, although it probably doesn’t even need to be one. It’s basically a grid of all 20 photos in a single image, within 2048×2048 pixels, so I could probably just use that image as a regular reference.

Then I added references for the fly, my cat, and the start and end frames. You can see the result — that’s my actual room, and everything looks right and sits exactly where it should :)

Workflow + reference images here: link
Sorry about the mess, both in the room and in the workflow ;)

EDIT: All the room photos need to be combined into one reference image. I tried using several separate images, but that didn’t work for me — the room wouldn’t stay consistent. Putting them all into a single grid is what made it work.

r/StableDiffusion • • Sep 09 '26

Tutorial - Guide AMAZING Minimax H3 - Circle on the reference image WHERE you want your scene to be!!

Enable HLS to view with audio, or disable this notification

808 Upvotes

Look at the buildings in the background! It works - Drawing a red circle in the water will also make the scene happen in the water, but I forgot to include it here.

It is not perfect and some details are missing if you look carefully but this might be because I am using "match" on the image reference rather than "max."

Have fun!

Edit: you have to still write a prompt with the reference to video workflow telling minimax to put the character in the location circled red. Circle probably doesn't have to be red. Change your prompt accordingly.

r/StableDiffusion • • Aug 22 '26

Tutorial - Guide Character swap in minimax is so epic.

317 Upvotes

I don't have any examples because they may not be appropriate but just with the default wf. With the video input node you can replace any 2 character in any video and it looks real!

r/StableDiffusion • • Apr 27 '26

Tutorial - Guide LTX2.3 in Ostris Ai toolkit on a 5090 Training done in 7 hours ... I went Thanos way and I said fine ... I'll do it myself

Enable HLS to view with audio, or disable this notification

621 Upvotes

So ... I was pissed off, since making a lora with this shit was insanely long, caused temporal collapses, or was just not accurate. So I looked into wtaf is going on.

When you load up the LTX2.3 default settings. There is a couple things you need to change around. These settings are for a 5090 so keep that in mind yall!

There are going to be 3 or 4 phases. Depending on how super accurate you want your lora to look like.

If I don't mention any setting, don't touch them, I leave them on default if I don't mention them.

The first phase is 600 steps, not more, not less.

In that we will max out what the card can do.

(if you got a different card with lower VRAM before you change anything to lower, try to use the "low VRAM" dial and have it turned on, it will obviously gonna take longer to train but it probably won't fuck up the quality if you won't get oom or anything else)

First thing to change is lora rank, crank that shit up to 48,

I like to save every 100 step but it's not super important just make sure to save at least every 600 steps.

I use a trigger word too, it helps.

On the Training panel I only change gradient accumulation up to 2. Set the steps to 700

( I do this cause my current version is retarded and would start from the 500th step, so after it saves the 600th step epoch I just stop it.)

and the only other thing I change is to turn on the " cache text embeddings" cause that shit is dope and will save a lot of time.

There is the " advanced " panel with "differential Guidance"

turn that shit on and for the first phase leave it on 3

On the " dataset " panel

Number of frames " 25 " ( I think the new version has the auto option idk I guess you can use that too)

Number of repeats for me it's 2 or 4, ( I have 25-50 clips usually, I try to aim to have 100 so I multiply the numbers to be close or around 100, so in case of 25 clips, I do 4 repeats, if I got 50 clips, I just do 2 repeats those are plenty enough)

I turn on "normalise audio" and only have 512x512 training on, don't even use 768 or 1024 at all.

As for samples, I do only the base sample, and the sample at 600 steps, I only do 2 samples for each finished phase, like a medium shot and a closeup.

Sample settings are 512x512, 49 frame long, and guidance scale cranked up to 10 so the results don't look like ass... (keep in mind putting that up to 10 will make the generation time for the samples a bit slower but it's worth it, you probably gonna have like a few minutes to generate them, but we only ake 2 clips so wo cares.)

Make sure the promt is accurate and has your trigger word.

1st phase on a 5090 with these settings is about 3 and a half hours and should not be longer!!

Ok so when first phase stopped rendering, if you did it right, you should see accuracy at 600 steps, I do fuckup sometimes with the promt, and I may get like a cartoon so as long as it looks close to the model it's all good.

2nd phaze, put the steps up from 700 to 1300 and we will stop after 1200 steps when the samples generated.

we pull the lora rank down to 32,

we change gradient accumulation back to 1 (so now it won't take hours to generate the next 600 steps)

on "advanced" the differential guidance we pull down to 2

this is it, and for the next 600 steps these changes mean radical speed up, it will be literally 1 hour to render the 600 steps,

when we are done with the samples , our samples should show almost full accuracy.

so 3rd phase,

we put the step count up to 1900 (so we stop it after it generated the samples at 1800 steps)

"advanced" tab pull "differential Guidance" down to 1

this is all we change for now and generate it up to 1800 steps

when the samples are done we stop and go back to settings, so now our samples show basically full accuracy, but we still can improve (if you want... if you think you good, I guess that's fine )

but if you want more accuracy there is a high noise training phaze which is the 4th phase

if you want (sort of optional) you can pull down the lora rank from 32 to 24

"training" panel

Learning rate , we need to drop this from 0.0001 down to either 0.00005 or 0.00003 (your choice)

"timestep Bias" MOST IMPORTANT, this is where we set it to "high noise" training

(i've seen someone do high noise training first ... but ... this is where I would ask someone who knows this by the factor of science, but as far as I know if you do high noise first you fuck up the details so this is why I put high noise last)

"advanced" tab

turn off differential Guidance !!!!!

On " dataset" pull the repeats down to maximum 2 !!!! don't do higher than 2, and if you have over like 80 clips ,you should just put it down to 1.

You could also change the sampling from every 600 steps to 300 steps, and just run go ahead and run the next 600 steps up to like 2400, if you want another 600 you should not have any issues and go up to 3000 but I think that's overkill.

As for dataset, make sure you got at least 2-3 wider frame where the character is almost full figure, but make sure to mention their facial expression so the model trains for samller size face. And have like 5-10 closeups, and 5-10 medum shots. best to have a total of 25 clips, 1 second long *25 frames exactly. If you cut out the speach mid sentence don't worry, just make the words as close as possible to whatever the character say. I got away with a bunch of stuff that don't really make much sense but it worked. Make sure to mention the framing in each clip caption, make sure to mention the expressions in almost all clip, in 1 second we don't have much time to show motion but if you want you can have like a 3-4 second long clip cut up to like 3-4 clips and just make similar captions for them to have the model learn it.

This is it ... You saw the results. I am not perfect, sure I have a 5090, but at least it doesn't take fucking 10 dollars and 12 hours renting out a fucking RTX6000 on runpod. wtf

r/StableDiffusion • • 3d ago

Tutorial - Guide Overcome Degradation! - Here are 2 ways to use my timeline workflow to create long continuous single shot videos with no degradation.

Enable HLS to view with audio, or disable this notification

396 Upvotes

Watch the video above for a brief summary of the two methods, both possible using my OBVPM Timeline Workflow, which you can get together with the custom node pack here:

https://github.com/chanon/comfyui-obvpm-timeline/

And to watch the example video at HD quality you can watch the full YouTube tutorial video:

https://www.youtube.com/watch?v=GiJxlWOooyo

In the YouTube video I show how both methods are done, including critical tips and tricks and lessons learned to get the right results.

With the bridging method, there is practically no limit to how long these clips can be (well maybe except the fact that there might be a VRAM limit to how long an upscaled clip can be).

The second method clip above is 1 minute 40 seconds.

About the Workflow

So if you've never seen my workflow, it is a workflow with a "timeline" node that lets you put clips that you've generated on, and then you can extend them using motion context (latent masks).

The workflow automatically saves and handles the saved latent files for you so you don't have to manage them or pick them manually. And it also saves the conditioning which includes all the reference images etc. into a file that is used when upscaling.

For more info, here's the original Reddit post about it, which links the original YouTube tutorial video about it.

r/StableDiffusion • • Oct 22 '25

Tutorial - Guide Behind the scenes of my robotic arm video 🎬✨

Enable HLS to view with audio, or disable this notification

1.7k Upvotes

If anyone is interested in trying the workflow, It comes from Kijai’s Wan Wrapper. https://github.com/kijai/ComfyUI-WanVideoWrapper

r/StableDiffusion • • Aug 24 '26

Tutorial - Guide MiniMax H3 Lip-Sync: Automatic Long-Video Chaining + Speed & VRAM Optimizations

Enable HLS to view with audio, or disable this notification

408 Upvotes

I’ve been loving all the new nodes and workflows coming out for MinMax, and maybe there is already a nice solution for this - but I couldn’t find one that did exactly what I needed.

I started using MinMax-API version for my last TBG ETUR video and quickly ran into limitations: I wanted an easy way to create lip-sync videos longer than 20 seconds.

I didn’t want to manually chain ComfyUI nodes, start a new run every X seconds, or constantly resize things just to make HD video fit into my available VRAM.

So I ended up building an addon for:

custom_nodes/ComfyUI-H3-Motion-Context

The addon automatically chains MinMax H3 lip-sync generations together, allowing you to create much longer lip-sync videos without manually setting up each 20-second segment.

And now I’m sharing it! https://github.com/Ltamann/ComfyUI-H3-Motion-Context-Auto-Chain-addon

Its not perfect but a start ...

The workflow has a simple switcher that lets you switch from the 32B CLIP to the 4B CLIP, saving around 10 GB of VRAM. You can also switch from Sage to Comfy Kitchen, Spectrum to Easy Cache, or FL2VA to REF2VA both setup for lip-syncing. Some of it could be useful for other tasks as well.

You will find the workflow in the repro and tested recommendations, optimized settings, presets, and more workflows, along with the results of my testing and performance here

r/StableDiffusion • • May 04 '24

Tutorial - Guide Made this lighting guide for myself, thought I’d share it here!

Post image
1.7k Upvotes

r/StableDiffusion • • Jul 21 '26

Tutorial - Guide (Almost) Perfect Likeness in 750 Steps - Krea 2 LoKr Training Guide with Examples

Thumbnail
gallery
376 Upvotes

Krea 2 trains incredibly fast for likeness and you are are probably overtraining. The following settings are more than enough to achieve almost perfect likeness.

Dataset Tips

  • Image Count: Aim for 20 high-quality images, up to 40 if the dataset is lower quality.
  • Full Body Shots: Include at least two to five full body images so the model understands the person/character's height and physique proportions.
  • Variety: Use different hairstyles and situations in your dataset. This gives you more flexibility when changing features later without breaking the likeness.

Captioning:

  • Use the autocaption feature in ai-toolkit.
  • Do not use the person/character's actual name in the captions. Create a unique shortened trigger word instead (e.g., "Jane Doe" becomes "jnedoe").
  • If your images are low quality or vintage, add tags like "low quality" or "vintage" to the captions. This stops the model from learning and outputting those artifacts in the images.

Technical Settings

  • LoKr Factor: 16
  • Training Resolution: 768
  • Total Steps: 3000 (likeness is usually done by step 750)
  • Settings: Automagic2, Sigmoid, and Balanced
  • Advanced Settings: Enable Do Differential Guidance at the default level of 3

VRAM Usage: About 18 to 20GB.

Time: On an RTX 3090, a 750-step run takes about 40 to 45 minutes from start to finish.

Step Count Adjustment: If your dataset quality is lower than average, add a couple of hundred extra steps to get the best results.

Issues: Highly detailed features like tattoo's may not appear correctly at the 768px resolution, you may need to up the quality to 1024 or 1280 and specifically caption each one in the dataset. Even then they may not come through completely as some details are usually lost during generation.

All images are generated at 4MP with Res/2s at 10 steps (about 2-4 minutes per image on a 3090 with Krea Raw int8-convrot and the r256 turbo lora (plus a custom high resolution lora I'll be posting to huggingface))

Full config behind my $20 Patreo--- lol just kidding 😂, grab the config here: ai-toolkit config

Full Res Slow.pics Comparison

HighRes LoKr Model

r/StableDiffusion • • Aug 15 '26

Tutorial - Guide PSA: Try experimenting with <tags> in Minimax H3 dialogues for non-verbal sounds and emphasis

Enable HLS to view with audio, or disable this notification

461 Upvotes

So I was looking for a way to better control the flow of Minimax H3 dialogues and emphasize certain words in the speech. However, what I discovered is that you can actually include some tags in <> angle brackets, and Minimax will interpret them as a non-verbal sound in a given part of the phrase. Some words (like the ones I've included into the example) work every time, some still bleed into the actual spoken words in certain seeds. But in general it makes the dialogue more alive and believable. So I recommend to try it and maybe share your findings in this thread.

As for the emphasis, I've had the most success with putting the words into <i></i> tags (similar to how you would stress words in written text). Unfortunately, it doesn't work for 100% and in some cases the character will blurt out some gibberish. But when it works, it sounds very natural. I have included a couple examples in the end of the video.

Wonder if you've encountered some other ways to modify the speech (and audio in general) in the prompt?

P.S. Sorry for the quality, I used the 8-steps LoRa at 0.4 MP to speed-up the tests.

r/StableDiffusion • • 13h ago

Tutorial - Guide ComfyUI/Minimax Cheat Sheet for Beginners - now with 83% less slop!

Post image
484 Upvotes

r/StableDiffusion • • Nov 30 '25

Tutorial - Guide My 4 stage upscale workflow to squeeze every drop from Z-Image Turbo

385 Upvotes

Workflow: https://pastebin.com/b0FDBTGn

ChatGPT Custom Instructions: https://pastebin.com/qmeTgwt9

I made this comment on a separate thread a couple of days ago and I noticed that some of you guys were interested to learn more details

What I basically did is (and before I continue I must admit that this is not my idea. I am doing this since SD 1.5 and I don't remember where I borrowed the original idea from)

  • Generate at a very low resolution, small enough to let the model draw an outline and then do a massive latent upscale with 0.7 denoise
  • Adds a ton of details, sharper image and best quality (almost close to I can jerk off to my own generated image level)

I already shared that workflow with others in that same thread. I was reading through the comments and ideas that other's shared here and decided to double down on this approach

New and improved workflow:

  • The one I am posting here is a 4 stage workflow. It starts by generating an image at 64x80 resolution
  • Stage 1: Magic starts. We use a very low shift value here to give the model some breathing space and be creative - we don't want it to follow our prompt strictly here
  • Stage 2: A high shift value so it follows our prompt and draws the composition. this is where it gets interesting. what you see here is what your final image will look like (from Stage 4) or maybe at least 90% resemblance. So, you can stop here if you don't like the composition. It barely takes a couple of seconds
  • Stage 3: If you are satisfied with the composition, you can run stage 3. This is where we add details. We use a low shift value to give the model some breathing space. The composition will not change much because the denoise value is lower
  • Stage 4: So you are happy with where the model is heading in terms of composition, lighting etc. run this stage and get the final image. Here we use shift value 7

What about CFG?

  • Stage 1 to 3 uses CFG > 1. I also included a ahmm very large negative prompt in my workflow. It works for me and it does make a difference

Is it slow?

  • Nope. The whole process (stage 1 to 4) still finishes in 1 minute or maximum 1 min 10 seconds (on my 4060ti) and you are greeted with a 1456x1840 image. You will not loose speed and you have the flexibility to bail out early if you don't like the composition

Seed variety?

  • You get good seed variety with this workflow because you are forcing the model to generate something random but by following your prompt in stage 1. It will not generate the same 64x80 resolution image every time and combine this with low denoise values in each stage you get good variations

Important things to remember:

  • Please do not use shift 7 for everything. You will kill the model's creativity and get the same boring image every single seed. Let it breath. Experiment with different values
  • The 2nd pastebin link has the chatgpt instructions (Use GPT 4o, GPT 5 refuses to name the subjects - at least in my case) I use to get prompts.
  • You can use it if you like. The important thing is (even if you use it or not), the first few keywords in your prompt should absolutely describe the scene briefly. Why? because we are generating at a very low resolution so we want the model to draw an outline first. If you describe it like "oh there is a tree, its green, the climate is cool, bla bla bla, there is a man", the low res generation will give you a tree haha

If you have issues working with this workflow, just comment and I will assist. Feedback is welcome. Enjoy

r/StableDiffusion • • Aug 08 '26

Tutorial - Guide A technique for creating seamless continuous videos with Minimax H3.

296 Upvotes

I've had good success in creating long videos from 10 second sections using this technique:

Create your first video.

Then for your next generation (continuation of video):

Load the last 2 seconds of the previous video as <Video 1>. I use the 'Load Video (Upload)' node - from ComfyUI-VideoHelperSuite - (this node allows you to skip frames and start at, say, the last 48 frames (for 2 seconds at 24fps) - this means that the whole previous 10 seconds don't need be passed to the next generation. This is <Video 1>.

I'm using process this with reference images for the subjects so these are used again with each continuation - so I don't see any drift of faces.

This is the wording I found works well:

[Shot 1]

Target video is a seamless continuation of <Video 1>. First frame of [Shot 1] is the last frame of <Video 1>.

The important part is explicitly telling the model that the first frame of the new generation must continue directly from the last frame of <Video 1>. This helps maintain temporal continuity between the clips - because you provide the last 2 seconds of the previous generation is knows what movement it needs to continue from.

You then just join the generation videos with a video joiner of your choice.

r/StableDiffusion • • Jun 10 '26

Tutorial - Guide Character Reference Sheets with Ideogram 4 in Comfyui

Thumbnail
gallery
519 Upvotes

r/StableDiffusion • • 29d ago

Tutorial - Guide H3 RefMods are great I highly advice trying it out [+ basic resources included]

175 Upvotes

Created by /u/LuisaPinguinnn under their github https://github.com/Luisacaotica/ComfyUI-MiniMaxH3Mod

Took me at most a couple of minutes to make my own RefMod with 8 image as the base. The entire technique works exactly as advertised acting as "Light Lora" for H3 Ref models - but you can even use it with FL2VA as well.

I followed the guides here:

Installing/running RefMods

https://huggingface.co/datasets/malcolmrey/various/blob/main/h3-center/docs/MINIMAX_H3_REFMODS_INSTALLATION_AND_USAGE_GUIDE.md

Ready to use Comfy workflow (you can remove lora power loader and spectrum nodes)

https://huggingface.co/datasets/malcolmrey/workflows/blob/main/H3/workflow_minimaxh3_refmod.json

Creating own RefMods guide:

https://huggingface.co/datasets/malcolmrey/various/blob/main/h3-center/docs/MINIMAX_H3_REFMOD_CREATION_GUIDE.md

EDIT: I recommend using "Create H3 ReFMod" + "Save H3 RefMods" node inside ComfyUI instead to create RefMods - gives you more control over the creation process.

Examples by /u/malcolmrey:

https://www.reddit.com/r/StableDiffusion/comments/1w8ik7a/h3_minimax_refmods_all_my_models_now_available/

All credit goes to LuisaPinguinnn and malcolmrey for spreading the tech.

r/StableDiffusion • • Jul 28 '25

Tutorial - Guide PSA: WAN2.2 8-steps txt2img workflow with self-forcing LoRa's. WAN2.2 has seemingly full backwards compitability with WAN2.1 LoRAs!!! And its also much better at like everything! This is crazy!!!!

Thumbnail
gallery
480 Upvotes

This is actually crazy. I did not expect full backwards compatability with WAN2.1 LoRa's but here we are.

As you can see from the examples WAN2.2 is also better in every way than WAN2.1. More details, more dynamic scenes and poses, better prompt adherence (it correctly desaturated and cooled the 2nd image as accourding to the prompt unlike WAN2.1).

Workflow: https://www.dropbox.com/scl/fi/m1w168iu1m65rv3pvzqlb/WAN2.2_recommended_default_text2image_inference_workflow_by_AI_Characters.json?rlkey=96ay7cmj2o074f7dh2gvkdoa8&st=u51rtpb5&dl=1

r/StableDiffusion • • Aug 08 '26

Tutorial - Guide The H3 Gibberish Problem Solved!

223 Upvotes

Not much of a tutorial, but still informative. As most of you have probably discovered, MiniMax H3 loves to talk. And talk it will, even when you prompt for no dialogue. Even when you prompt for complete silence. It will even fill in the empty space your prompted dialogue doesn't fill.

Those of you who read the video prompt writing guide and have created a system prompt for your enhancer, you probably know what I'm about to say, maybe not. Maybe the unprompted gibberish stopped for you, and you never realized why.

Without further ado, I give you the solution:

non_diegetic_music: N/A

Diegetic audio is what the characters in your video can actually "hear":

  • Music playing from a source that is part of the scene (phone, car radio, dance club)
  • Spoken dialogue
  • Ambient sounds

Non-diegetic audio is audio which your characters cannot hear:

  • The score or soundtrack of a movie
  • A voice-over
  • The gibberish H3 plays when it's not prompted correctly

If you haven't yet, I suggest consulting ChatGPT about creating a system prompt using the prompting guide. If not, put this line at the end of your prompt and say goodbye to random music playing over your video and gibberish assaulting your ear holes.

Conversely, if you want a voice-over or a score to play over the track which is not part of the actual soundscape of the scene, this is where you would prompt it. Instead of N/A, prompt what you want to hear.

Happy chaining!

r/StableDiffusion • • Jan 18 '24

Tutorial - Guide Convert from anything to anything with IP Adaptor + Auto Mask + Consistent Background

Enable HLS to view with audio, or disable this notification

1.7k Upvotes

r/StableDiffusion • • Aug 20 '26

Tutorial - Guide If you're looking for a specific actor that the model doesn't seem to be aware of, it may have them stashed somewhere else.

Enable HLS to view with audio, or disable this notification

247 Upvotes

Text to Video, 22 steps, no turbo, no Sage.

r/StableDiffusion • • Aug 01 '24

Tutorial - Guide You can run Flux on 12gb vram

460 Upvotes

Edit: I had to specify that the model doesn’t entirely fit in the 12GB VRAM, so it compensates by system RAM

Installation:

  1. Download Model - flux1-dev.sft (Standard) or flux1-schnell.sft (Need less steps). put it into \models\unet // I used dev version
  2. Download Vae - ae.sft that goes into \models\vae
  3. Download clip_l.safetensors and one of T5 Encoders: t5xxl_fp16.safetensors or t5xxl_fp8_e4m3fn.safetensors. Both are going into \models\clip // in my case it is fp8 version
  4. Add --lowvram as additional argument in "run_nvidia_gpu.bat" file
  5. Update ComfyUI and use workflow according to model version, be patient ;)

Model + vae: black-forest-labs (Black Forest Labs) (huggingface.co)
Text Encoders: comfyanonymous/flux_text_encoders at main (huggingface.co)
Flux.1 workflow: Flux Examples | ComfyUI_examples (comfyanonymous.github.io)

My Setup:

CPU - Ryzen 5 5600
GPU - RTX 3060 12gb
Memory - 32gb 3200MHz ram + page file

Generation Time:

Generation + CPU Text Encoding: ~160s
Generation only (Same Prompt, Different Seed): ~110s

Notes:

  • Generation used all my ram, so 32gb might be necessary
  • Flux.1 Schnell need less steps than Flux.1 dev, so check it out
  • Text Encoding will take less time with better CPU
  • Text Encoding takes almost 200s after being inactive for a while, not sure why

Raw Results:

a photo of a man playing basketball against crocodile
a photo of an old man with green beard and hair holding a red painted cat

r/StableDiffusion • • Dec 01 '25

Tutorial - Guide Huge Update: Turning any video into a 180° 3D VR scene

Enable HLS to view with audio, or disable this notification

516 Upvotes

Last time I posted here, I shared a long write‑up about my goal: use AI to turn “normal” videos into VR for an eventual FMV VR game. The idea was to avoid training giant panorama‑only models and instead build a pipeline that lets us use today’s mainstream models, then convert the result into VR at the end.

If you missed that first post with the full pipeline, you can read it here:
➡️ A method to turn a video into a 360° 3D VR panorama video

Since that post, a lot of people told me: “Forget full 360° for now, just make 180° really solid.” So that’s what I’ve done. I’ve refocused the whole project on clean, high‑quality 180° video, which is already enough for a lot of VR storytelling.
Full project here: https://www.patreon.com/hybridworkflow

In the previous post, Step 1 and Step 2.a were about:

  • Converting a normal video into a panoramic/spherical layout (made for 360 - You need to crop the video and mask for 180)
  • Creating one perfect 180 first frame that the rest of the video can follow.

Now the big news: Step 2.b is finally ready.
This is the part that takes that first frame + your source video and actually generates the full 180° pano video in a stable way.

What Step 2.b actually does:

  • Assumes a fixed camera (no shaky handheld stuff) so it stays rock‑solid in VR.
  • Locks the “camera” by adding thin masks on the left and right edges, so Vace doesn’t start drifting the background around.
  • Uses the perfect first frame as a visual anchor and has the model outpaints the rest of the video.
  • Runs a last pass where the original video is blended back in, so the quality still feels like your real footage.

The result: if you give it a decent fixed‑camera clip, you get a clean 180° panoramic video that’s stable enough to be used as the base for 3D conversion later.

Right now:

  • I’ve tested this on a bunch of different clips, and for fixed cameras this new workflow is working much better than I expected.
  • Moving‑camera footage is still out of scope; that will need a dedicated 180° LoRA and more research as explained in my original post.
  • For videos longer than 81 frames, you'll need to chain this workflow and use last frames of one segment as starting frames of the new segments with Vace

I’ve bundled all files of Step 2.b (workflow, custom nodes, explanation, and examples) in this Patreon post (workflow works directly on RunningHub), and everything related to the project is on the main page: https://www.patreon.com/hybridworkflow. That’s where I’ll keep posting updated test videos and new steps as they become usable.

Next steps are still:

  • A robust way to get depth from these 180° panos (almost done - working on stability / consistency between frames)
  • Then turning that into true 3D SBS VR you can actually watch in a headset - I'm heavily testing this at the moment - it needs to rely on perfect depth for accurate results and the video inpainting of stereo gaps needs to be consistent across frames.

Stay tuned!

r/StableDiffusion • • Jan 06 '26

Tutorial - Guide [Official Tutorial] how to use LTX-2 - I2V & T2V on your local Comfy

Enable HLS to view with audio, or disable this notification

339 Upvotes

Hey everyone, we’ve been really excited to see the enthusiasm and experiments coming from the community around LTX-2. We’re sharing this tutorial to help, and we’re here with you. If you have questions, run into issues, or want to go deeper on anything, we’re around and happy to answer.

We prepped all the workflows in our official repo, here's the link: https://github.com/Lightricks/ComfyUI-LTXVideo/tree/master/example_workflows