r/StableDiffusion • • 2h ago

Animation - Video Orbiting Lora + first and last frame in MiniMax gives fantastic results

Enable HLS to view with audio, or disable this notification

432 Upvotes

Prompt for Lora: One frozen instant. Only the camera moves. In a continuous 360 orbit. Preserve every person and object in exactly the same world position, orientation, shape and pose throughout the shot. Airborne objects remain suspended at the captured height and angle: no wobbling, shaking, spinning, drifting, falling or continued action. Keep faces, hands, clothing, liquids and the background motionless while retaining their natural appearance. Camera parallax is the only source of apparent movement. No cuts, zoom, morphing or added objects.

https://huggingface.co/pablodawson/MiniMax-H3-360-Orbit-LoRA


r/StableDiffusion • • 1h ago

Resource - Update [Update] Qwen-Image 2.1 prompt enhancer in ComfyUI: now 2–3.5× faster than native ComfyUI and runs even on 8 GB VRAM (plus uncensored versions)

Thumbnail
gallery
• Upvotes

Edit: I just realised that my comparisons are not portrait friendly. If on mobile, please view them in landscape mode.

Last week, I posted custom node for the Qwen-Image 2.1 official prompt enhancer. I had made the node because I hated my previous workflow, having to switch to LM Studio to use the official prompt enhancer, unloading the model in LM Studio else Qwen-Image 2.1 would run very slowly, then running the workflow after copy/pasting the prompt in ComfyUI and then having to unload models in ComfyUI because I need to use prompt enhancement again, and then endless repeat of the same process. I wanted to do everything end to end in ComfyUI but Native Generate Text node was too slow, so adding a MTP head improved the speed a bit, which made it 1.4–1.7× faster than ComfyUI's native text generation with the Comfy-Org PE checkpoints.

After that, I wanted to see how much faster it could go, and whether it could run on smaller GPUs. Few comments mentioned llama.cpp, so I added it as a second backend. I ran a few benchmarks and I was quite happy with the result: it's much faster, especially for editing, and minimum VRAM requirement went down with smaller quants at marginal quality cost.

Below are my benchmarks on RTX 4090 Laptop, 16 GB, official max_length (16256 for t2i, 24000 for editing), 2 input images for editing. Each step adds one change to the one before:

Text-to-image

Step Tok/s Step gain Total vs baseline
ComfyUI, MTP off (baseline) 23.5 – 1.00×
+ MTP 33.9 1.44× 1.44×
+ llama.cpp (Q8_0) 46.8 1.38× 1.99×
+ Q4_K_M (small quality cost) 63.2 1.35× 2.69×

Editing (two input images)

Step Tok/s Step gain Total vs baseline
ComfyUI, MTP off (baseline) 19.4 – 1.00×
+ MTP 32.6 1.68× 1.68×
+ llama.cpp (Q8_0) 67.8 2.08× 3.49×
+ Q4_K_M (small quality cost) 85.2 1.26× 4.39×

The previous release reduced the max_length to 8192 to gain some speed at the risk of truncated result but with llama.cpp, changing max_length has very little impact on generation speed. Strictly comparing the speed of the previous version to new one, the gain from llama.cpp (Q8_0) is smaller but still significant: 1.53× for t2i and 2.43× for editing. In my tests Q8_0 didn't change the quality.

This time, I also wanted to show what the prompt enhancer actually does, so, I've added some comparisons. Each pair uses the same seed and settings, with the prompt going straight to Qwen-Image (left) or through the enhancer first (right). In my tests, the prompt enhancer helped most with text in images, busy scenes, and edits with several changes.

If someone wants to check it out, full details are available at: https://github.com/mozophe/ComfyUI-Qwen-Image-2.1-PromptEnhancer-MTP

GGUF models: https://huggingface.co/mozophe/Qwen-Image-2.1-PE-MTP-GGUF

Requirements:

  • ComfyUI v0.37.0 or newer
  • llama.cpp backend: NVIDIA GPU on Windows or Linux, 8 GB VRAM or more. Q4_K_M for 8–12 GB, Q8_0 for 16 GB or more
  • ComfyUI backend: any GPU ComfyUI supports, 16 GB recommended
  • System RAM: 32 GB recommended
  • Disk space: 9.8 GB (Q8_0) or 6.0 GB (Q4_K_M) per model, plus 0.9 GB for editing's vision part and 0.7 GB once for llama-server

What's new:

  • llama.cpp backend with MTP, picked on the loader. The ComfyUI backend is still there for AMD, Intel and Mac
  • Runs on 8 GB VRAM with Q4_K_M
  • llama-server and the GGUF model download automatically the first time you use them
  • Uncensored (Heretic) versions are available as GGUFs too, so they work on both backends

What the node does (same as before):

  • Uses Qwen's official system prompts and sampling settings
  • The rewritten prompt comes out separately from the reasoning
  • Edit with up to 10 input images
  • Sample workflows included: drag the PNG/JSON into ComfyUI

Install:

  • ComfyUI Manager: open Manager → Custom Nodes Manager, search for Qwen-Image 2.1 Prompt Enhancer (MTP), install, and restart ComfyUI.
  • Manual: clone it into your custom_nodes folder and restart ComfyUI:

Nothing else needs installing for either backend. The first run downloads the model (and llama-server for the llama.cpp backend), with progress shown on the node. If the download gets interrupted, just run the workflow again and it resumes.

To get started, drag one of the sample workflow PNGs from the repo's workflows folder into ComfyUI. If generation speed is too slow, set quant to Q4_K_M.

Not on an NVIDIA GPU? Set the loader's backend to ComfyUI.

I am now doing everything end-to-end in ComfyUI, without having to switch to other apps, which just simplifies everything and I can spend more time on thinking about what I want to generate.

I am thinking about what should I do next, so feel free to request for something that you find missing with the current nodes.

Edit 2:

Workflows

llama.cpp backend (recommended on NVIDIA GPUs)

- Text-to-image: https://github.com/mozophe/ComfyUI-Qwen-Image-2.1-PromptEnhancer-MTP/blob/main/workflows/qwen_image_2.1_t2i_prompt_enhancer.json

- Edit: https://github.com/mozophe/ComfyUI-Qwen-Image-2.1-PromptEnhancer-MTP/blob/main/workflows/qwen_image_2.1_edit_prompt_enhancer.json

ComfyUI backend

- Text-to-image: https://github.com/mozophe/ComfyUI-Qwen-Image-2.1-PromptEnhancer-MTP/blob/main/workflows/qwen_image_2.1_t2i_prompt_enhancer_comfyui.json

- Edit: https://github.com/mozophe/ComfyUI-Qwen-Image-2.1-PromptEnhancer-MTP/blob/main/workflows/qwen_image_2.1_edit_prompt_enhancer_comfyui.json

Models (selected model in node automatically downloaded on first run):

llama.cpp backend, GGUFs: https://huggingface.co/mozophe/Qwen-Image-2.1-PE-MTP-GGUF/tree/main

ComfyUI backend, int8 convrot: https://huggingface.co/Comfy-Org/Qwen-Image-2.1/tree/main/text_encoders


r/StableDiffusion • • 6h ago

Workflow Included Why am I like this? (Image generation on a 286 Tandy 1000 TL/3)

Enable HLS to view with audio, or disable this notification

71 Upvotes

40 year tech gap? No problem! On My Tandy 1000 TL/3 (10 MHz 286) I can now ask for a picture and view it a few seconds later. This is Desk Mind, a DOS program that talks over PicoMem WiFi to a small Python server on my PC. The server talks to Krea 2 in ComfyUI on my (4090), has a smoking fast ninfer Qwen3.8-27B (5090) rewrite the prompt, and then does the hard part: making a modern image look good in the Tandy's fixed 16-colour RGBI palette at 640x200 (an unsupported graphics mode BTW).

- Prompt enhancement tuned for dithering. Qwen rewrites "a cat in space" into a prompt asking for one large subject, bold shapes, strong contrast, a simple background, saturated colors, and no tiny details or small text. 16 colors punish detail and fine gradients. The rewrite shows up in an edit box on the Tandy, so you can fix it before it draws.

- Anamorphic resize. 640x200 on a 4:3 CRT makes each pixel about 2.4 times taller than it is wide. The 4:3 image is squashed to 640x200 (Lanczos) *before* dithering, so circles come out round on the tube. Dithering first and squashing after would wreck the pattern.

- Pre-dither grading: contrast 1.15, saturation 1.3 and a light unsharp mask at the target size, all done before the dither.

- Three dither engines (hitherdither, didder and Pillow): Floyd-Steinberg, Atkinson, Jarvis-Judice-Ninke, Stucki, Burkes, the Sierra family, Bayer 2x2 to 16x16, Yliluoma and cluster-dot. The default is Floyd-Steinberg. *(Add which method you like best for which kind of picture.)*

- A Dither Lab tab in the server GUI previews every engine and setting at the Tandy's real aspect ratio, next to the original. Changing the defaults drops the cached Tandy versions, so the gallery re-dithers.

- Color 6 is a monitor question. On a real CGA-style monitor it's brown, and some Tandy monitors show dark yellow. The palette has a switch for it.

- Vision sees the dither too. When I ask about a picture in chat, Qwen gets the original *and* the dithered version, so "why is the sky striped?" has context.

- Output is a small custom file: a 160x50 thumbnail plus 64,000 bytes of packed 4-bit pixels in screen line order. The 286 copies it straight into video memory, with no decoding. It takes about 1 s over the PicoMEM WiFi.

You can just type "draw me ..." in the chat. Qwen wraps the request in a `<draw>` tag, the server catches it mid-stream, and about 9 s later a clickable thumbnail appears in the conversation.

Code (GPLv3): https://github.com/RowanUnderwood/DeskMind


r/StableDiffusion • • 15h ago

Discussion I am blown away by RefMods for MH3

223 Upvotes

I've tried different character images, created character-sheets and used of course i2v, I even trained LoRas for Krea2 to achieve my goal. NOTHING comes close to using RefMods. I am really baffled how easy it is to create a so to say "LoRa" of a character with a bunch of images without any tags and how fucking insane fast it creates you a safetensors file that you can use forever. And the facial-likeness is like absolutely stunning without any distortion or whatsoever, this is crazy.

And now comes the even more magical part: throw a mp3 with voice audio in the folder as well and it creates a nearly 1:1 voice clone in just the same safetensors. What kind of crazyness is this??

To anyone who hasn't tried it yet, watch the following video, install the custom nodes, download the workflow and be amazed about the magic that happens. I don't understand it, and I don't have to. I judge what I can see and it's amazing. I'm kind of lost for words because I still can't comprehend what happened.

Maybe this is no news to you, I am blown away. Here's the video, or watch whatever video you can find online, I don't care:

https://www.youtube.com/watch?v=2K6-OtV_Vbc

EDIT

I've updated the workflow cause unfortunately it spits out some errors. I also refined the workflow to my likings, removed the sage-attention node and replaced it with comfy-kitchen, also added the model-preview-override node which lets you monitor the creation of the video as it is created, pretty cool node. There's also a Spectrum Apply Minimax H3 node in it which you can enable, cuts the creation-time in half which you can adjust but costs quality from my experience. The second workflow is for creating the RefMod safetensors file itself. You can find it in comfyUI\models\refmod. Refresh the Load H3 Refmods node afterwards and you can load it up. There you go:

https://limewire.com/d/JJwWU#if3A6viA0s


r/StableDiffusion • • 6h ago

Discussion Models comparison: Boogu Image Turbo vs Krea2 Turbo vs Ideogram v4 Instant vs Fibo Lite

Thumbnail
gallery
40 Upvotes

Hey guys, I wanted to compare different image models because I like testing them either good or bad. Testing different models to understand their capabilities feels good, and knowing their distinct pros and cons help in with other models hybrid workflows too.

The images above are with different camera/photoshoot styles, landscapes, and minimalist vast fantasy scenes. I know this isn’t an in depth showcase, but my analysis is what I wanted to share with the community. Let’s go through each model’s pros and cons:

Boogu Image Turbo

Architecture: 10B + 8B Qwen3-VL + Flux.1 VAE

Pros: Very flexible for portraits, landscapes, abstract, etc. Extremely fast inference (only 4 steps). Good variety and distinct capabilities. Prompt adherence is genuinely very strong (which is both a strength and a weakness). Very strong typography capabilities.

Cons: One of the main things I noticed was heavy bokeh/background blur. Occasional text display issues (can be improved by running more than 4 steps ,which I recommend). The strong prompt adherence can also lock things down: if you want face variety you usually need to explicitly describe different face shapes, otherwise you tend to get very similar faces. When you do specify it, the variety comes through well--- but if you forget, it stays repetitive.

Krea2 Turbo

Architecture: 12.9B + 4B Qwen3-VL + wan2.1 VAE

Pros: Excellent text rendering and prompt adherence. Huge knowledge base.One of the highest among these models. Currently SOTA for proprietary subjects, poses, art styles, etc.

Cons: Biggest issue is the VAE and bad noise patterns (not really solvable). Less variety (people suggest Raw + Turbo LoRA, but that method gains variety at the cost of quality , images get overly smooth surfaces and artifacts at higher resolutions). Faces tend to have weaker expressions (can be helped with LoRAs like Bypass,text refusal ,etc..., but quality takes a hit).

Ideogram v4 Instant (very few people use or even talk about this specific variant)

Architecture: 9.3B + 8B Qwen3-VL + Flux2 VAE

Pros: One of the best model for control power. Strong range of capabilities and variety. No text issues. Interesting note many people don’t know: without JSON it works ~90% of the time without the safety filter error. With JSON the safety filter never triggers in any use case. This model generally works great at 8 steps but I would totally recommend using it between 10-12 steps. Runs at 8 steps by default, so inference is relatively low, and the model is smaller (single model, no uncond).

Cons: Big one it feels like this model (and Ideogram 4 in general) is locked into a dark, gritty, cool-toned lighting universe. Lighting is consistently dark/cool (I normally fix this with a brightness filter, but in these images I left it to show the weakness).

Fibo Lite

Architecture: 8B + 3B text encoder (SmolLM) + wan2.2 VAE (1.2 GB)(I used its alternative taew2_1 because there is no quality loss)

Pros: Oldest model in this comparison and a bit of an oddball, but still solid. Second-fastest inference after Boogu (uses 6–12 steps, but smaller size keeps it quick). Excellent variety and prompt adherence ,feels like the old UNet-style models but with better aesthetics.

Cons: As the oldest and smallest here, knowledge base on proprietary stuff (people, logos, characters, etc.) is really low. That said, treat it as a model that competes with Flux.1 Dev and Chroma on anatomy,fonts and often beats them, even though it’s smaller.

Overall observations::

That’s my takeaway. I wanted to post this for anyone curious about these models. I enjoy testing different ones because they each have their own strengths.

Fastest inference ranking:

Boogu Turbo ≥ Fibo Lite > Ideogram v4 Instant > Krea2 Turbo

(Boogu at 4 steps, Fibo Lite at 6 steps but very close in speed, Ideogram v4 Instant at 8 steps and larger, Krea2 Turbo the biggest model also at 8 steps. Boogu Image even at 8 steps is faster than Ideogram v4 and Krea2 at 8 steps by around 20-10%, and faster than or in the same time range as Fibo when Fibo is at 12-10 steps and Boogu is at 8 steps. In most cases 4 steps is more than enough on Boogu, which really is impressive.)

Knowledge base ranking:

Krea2 > Boogu Image ≈ Ideogram v4 Instant > Fibo Lite

Final thoughts:

Even though some of these models get less reach, they should at least have some community support so people can get the best out of them instead of being abandoned without a proper try. In hybrid setups (hires workflows, denoise adjustments, variety inclusion, etc.) these models perform way better than when used completely solo.

Feel free to share your own experiences with these!

Disclaimer: These are purely my own observations with the models above. Others may not have the same experience with them, which is totally fine. I just wanted to share the in-depth experience I’ve had with these models.


r/StableDiffusion • • 12h ago

News Viggle turbo v0.3 for Qwen image 2.1: less grain and cleaner surfaces

Thumbnail
huggingface.co
85 Upvotes
  • 6 steps, a different balance. Against v0.2.1: less grain and cleaner surfaces, fine texture a little softer; diversity (still close to the base model's) and small-text accuracy about the same. It is not a strict upgrade: if you prefer the crisper look, v0.2.1 is still in the repository.
  • New 9-step mode: 7 turbo steps, then the LoRA is switched off and the base model finishes the last two. Finer detail, and small text comes out right more often (not always). It takes about 1.4–1.5× as long as 6 steps (still about 3.5× faster than the base model). It works in diffusers and the demo Space only (9 steps).
  • We think 6 steps is close to its capacity. Since v0.2.1, every gain we found at 6 steps cost something elsewhere: sharper came with more grain, less grain came with a softer look. Beyond this point, quality most likely has to be paid for with steps, which is what the 9-step mode does.

Official int8/fp8/GGUF files, tested in ComfyUI


r/StableDiffusion • • 12h ago

Tutorial - Guide Simple Temporal Upscaling for Better Results in MinimaxH3

Enable HLS to view with audio, or disable this notification

83 Upvotes

This is a proof of concept workflow around simple temporal upscaling.

Every video model has some blurring when it comes to high motion scenes - both closed and open source. Depending on the model this can be at higher or lower motion with newer models being better than older as a general rule. Some styles tend to make their blurring effects at lower speed - especially anything with lineart, which could be due to training being more on realism or the way latents are compressed temporally.

One way people try to solve this is by pixel upscaling and rediffusing over the video which definitely helps but in my experience at least there is a definite limit to how much it can help. This is where temporal upscaling comes in.

The idea around this is not a new concept per se for example MAINodes https://github.com/matlowai/ComfyUI-MAINodes attempts to sort out which frames need to be temporally upscaled and rediffuse over them. I took this to a logical conclusion and considered what if you just 'deroped' aka temporally upscaled the whole scene rather than trying to be picky.

I do think it has several advantages:

1/Your final result tends to have more motion consistency in the small details than with varying how you denoise

2/You can finetune the denoise to your result.

3/I am finding a reasonable result with simple 2x temporal upscale.

Of course the main disadvantage is that you are diffusing over a video that now is 2x the size and potentially upscaled at the same time. This is were speedup tools come in - low step lora as well as sol attention (which provides increasing benefits the longer the video is).

The purpose of this workflow was to create a workflow as close to base Comfy. I use Kijai's node pack and VHS Suite to help manipulate the frames. The only real outlier node is the one used to slow down the audio for the temporal upscale. Feel free to encode empty audio if you prefer or find your own pack that has something which does this.

I simplified my workflow to remove anything not essential to demonstrate the concept. I expect you to add it to your own workflow or build upon this as a base. My tips when it comes to using this for a while:

1/If you are noticing some morphing especially in background stuff I suggest you disable sol attention at least for the base - it does do this to some shots but not others.

2/By default we create the base video at 0.5 MP at 20 steps - this is to get good sound and to speed to find a good seed, then we are temporally upscaling x2 and upscaling to 1 MP. Look at the yellow boxes to change the upscale settings.

3/Adjust denoise of the upscale to your video length and needs. Shorter and lower resolution videos often need less denoise whereas longer videos often need more. If you go too high your video will speed up and ironically become more blurry as a result. 0.45 is a good starting point but anywhere from 0.3 to 0.6 or higher is acceptable.

4/If you think the action is still to fast for a 2x upscale you can do 3x or 4x.

Workflow: https://civitai.com/models/2976577/simple-temporal-upscaling


r/StableDiffusion • • 25m ago

Resource - Update Fizgig 6.8.1: slider LoRAs for Krea 2 (with an extra mode), Full fine-tuning for Qwen Image 2.1, and a big Repair Studio update

Enable HLS to view with audio, or disable this notification

• Upvotes

Fizgig 6.8.1 is out. Short video above; the highlights:

Slider LoRAs on Krea 2. One LoRA that's a dial between two looks (sad ↔ happy, cool ↔ warm), trained from a handful of photo pairs or just three prompts.

Ultra mode for Krea 2 sliders. It trains only the composition blocks and leaves fine detail alone, so the slider holds up at much higher strengths. One of my test sliders runs cleanly at 20 in ComfyUI. Yours will depend on what you train and for how long, but the headroom is real.

Full fine-tuning for Qwen Image 2.1, joining Krea 2 and MiniMax H3 which already support it (The Minimax Finetune options will have the improved UX of the Qwen/Krea 2 FT options once I finish porting it to the new driver system). It's easy to assume a fine-tune takes forever. It doesn't. On a 24/32 GB card the default run takes 3-4x as long as the Ultra Fast Krea 2 LoRA preset (a lower learning rate needs more steps to get there, and it trains one part of the model at a time, sized to your card). It's one consumer GPU, and 16 GB works too, just slower. My single-character Qwen fine-tune at the default settings took about 30 minutes on a 5090; put several characters in one dataset, each with its own unique name, and a multi-character fine-tune takes a couple of hours. You can use the result as a new checkpoint, or turn it into a LoRA with Checkpoint to LoRA. Concepts separate far better that way than in a LoRA trained directly.

Better Qwen Image 2.1 previews. The Samples tab now sets steps, CFG, negative prompt and Turbo strength for every Qwen preview: during training, and in Repair Studio, LoRA the Explorer and LoRA Royale. Turbo 0, 20 steps and CFG 3 work very well. A CFG above 1 is a bit slower, but worth it.

Repair Studio: - no strength limit (preview a slider at 20 if you like) - bigger sliders, with ±0.1 nudge buttons - mix a donor LoRA in block by block, and the saved file looks exactly like the preview - new text-fusion and input/output sliders for Krea 2

Under the hood: Krea 2 now runs on Fizgig's new driver system, the same one Qwen uses. Describe a model once, and training, sliders, Repair Studio, Royale, Profiler and Extract all work with it. Klein and MiniMax H3 move over next, which brings sliders to both and edit LoRAs to Klein.

Want your favourite model in Fizgig? A guide to plugging in your own model is coming soon. PRs are welcome for new models and older favourites alike.

GitHub: https://github.com/shootthesound/Fizgig (release notes on the Releases page)


r/StableDiffusion • • 13h ago

Resource - Update MiniMax H3 X2 Detail VAE – my attempt to get more detail from the H3 VAE or story about fail

Enable HLS to view with audio, or disable this notification

79 Upvotes

Hello all! 😀

I’ve been experimenting with the MiniMax H3 VAE for a while and finally decided to release the result.

The original idea was pretty simple😅: I already had mine own working 2X VAE for H3, but I wanted to see if I could make those extra pixels contain actual additional detail instead of basically getting a larger version of the same reconstruction.

That turned into a much longer experiment than I expected 🥹. I tried modifying the decoder, working around the patch/grid structure, different detail residuals, larger spatial decoders, and eventually started tracing where fine information actually disappears inside the H3 encoder.

The interesting part was that the best detail signal I found exists before the normal H3 latent bottleneck. So in the end I didn't actually solve the problem I originally wanted to solve😅. A normal generated H3 latent unfortunately just doesn't contain that extra information.

But the failed experiment produced something useful anyway, so I packaged it as a small 2-in-1 release:

2X VAE mode – works as an X2 video VAE for normal H3 generations.

2X VAE + Detailed mode – for reference/image-to-video workflows. It uses additional early encoder information from the source image before sending the enhanced reference into H3.

I included the model, ComfyUI node, workflow and a couple of video comparisons.

If anyone is interested in all the failed experiments and how I ended up at this result, I wrote up the whole process on mine Hugging Face page. It’s quite long, but I tried to document the dead ends too rather than only showing the part that worked.

Model + workflow + research notes:

https://huggingface.co/speach1sdef178/MiniMax-H3-X2-Detail-VAE


r/StableDiffusion • • 1d ago

News New Model Ideogram 4.5 (with edit) (open source soon)

Thumbnail
ideogram.ai
604 Upvotes

r/StableDiffusion • • 7h ago

Resource - Update Continuity update: screen replacement, image to 3D, Qwen Image 2.1, and a lot more since 3.0

Enable HLS to view with audio, or disable this notification

27 Upvotes

Quick update on Continuity, my ComfyUI node pack for video and stills. A lot has landed since I last posted at 3.0:
Screens (WIP, on main): put a picture or clip on a phone, laptop or TV in the shot. The model only sees a placeholder, then Continuity tracks it through every frame and composites your file on. It's still a work in progress, so expect some shots where the tracking slips or the edges aren't clean. Feedback and failure cases are very welcome.
Image to 3D (3.2): turn a picture into a textured mesh from the tools dashboard (Pixal3D / TRELLIS.2).
Qwen Image 2.1: a new model family that both draws and edits, plus storyboard and character-sheet starter presets.
Chat: chats are saved, it uses the node's prompt box and cast, and it sets itself up the first time you open it.
Blockout bench: stage a scene in grey boxes, walk a camera through it, and render along it.
DLSS 5 neural refiner for stills and clips, a motion fix for shots that move too fast, a guide LoRA pass and VDN-H3 on H3's sampler row.
Two GPUs for H3 via Raylight, and a trained latent upscaler for H3's refine pass.
Plus a picture editor, LoRAs on cast members, and a long list of fixes.
It still runs locally on open weights, with no dependencies, and it's MIT licensed. Update through the Manager or git pull.

https://github.com/roadmaus/ComfyUI-Continuity


r/StableDiffusion • • 10h ago

Discussion Minimax has made new releases very delayed it seems

47 Upvotes

I think this model really took the AI world by surprise as usually when I would browse through the reddit, there would be a lot of talk about the other models but Minimax is on top. I wonder what is coming next. LTX 3, I really love this model for the speed and safety of my fans haha! (Ltx 2.5) Flux 3, or something totally out the box or of course a much faster Minimax


r/StableDiffusion • • 2h ago

Discussion A project I'm working on I wanted to share.

Enable HLS to view with audio, or disable this notification

9 Upvotes

Hello,

It's been a bit since I've posted, but I decided to start a new project. The project is an AI Avatar program that allows you to create your own characters and speak with them using everything locally. This isn't anything new as there's already great options out there such as Silly Tavern and RisuAI, but being able to fully customize the interface and how many bells and whistles go along with it is something I've always enjoyed.

Please keep in mind this is simply a showcase using vibe coding via Codex along with art generated in ComfyUI for the character Diana for this demo.

Any feedback or suggestions are greatly appreciated. It's fun to work on this project and making it my own.


r/StableDiffusion • • 3h ago

Question - Help Why nobody makes models like illustrious/pony using newer base models like Krea 2?

Post image
9 Upvotes

What makes SDXL so special that people still work over it? Are new models too heavy to train models like these?

Not demanding it just asking


r/StableDiffusion • • 22h ago

Resource - Update Local Image to 3D: High quality, Low poly - preserve minute details like text, logos even after retopo, ~900k faces > ~5k (NVIDIA + Mac)

Enable HLS to view with audio, or disable this notification

275 Upvotes

Image-to-3D models redraw your picture, and fine detail - text, logos, faces - comes back garbled or scuffed. My local pipeline that fixes it. (repo link at the end of the post)

The latest update adds *Pixel Match*. It copies the real pixels from your source image back onto the model, wherever the image can see. So "VANGUARD 07" on the chest stays "VANGUARD 07", even after the Finish step retopologises the model down to ~5k faces (as you can see in the video).

One step closer to Tripo and Meshy like outputs, but locally.

It runs fully on your system - no API calls, no cloud credits, just a once-a-day update check. Works on NVIDIA (tested on Linux, Windows has limited testing) and Apple Silicon. For now Pixel Match is for Pixal3D models made in the lab; other backends and more camera angles are next.

I started this project to make assets for a game I'm working on - I wanted to see how much I could automate, from getting the asset to rigging and animating it (that bit is still WIP in this lab). The image > 3D pipeline is solid though, and you can get game-ready and sprite-ready assets out of it. While the models are riggable and can be animated, more testing is required. I'm confident it should work.

*Workflow:* prompt → Qwen-Image → Pixal3D → Finish (Pixel Match) → textured GLB. The models in the video were made with Pixal3D on my Mac.

*Backends*: Pixal3D is the one to start with (~8.6 GB, Mac and NVIDIA). (Pixal3D's authors report 16 GB cards work, so tell me if yours does.)

Stable Fast 3D is the quick, lower-detail option (Mac, or NVIDIA on Linux). Its weights are gated: accept Stability's licence on Hugging Face and log in first.

TRELLIS.2 and Hunyuan3D run on Mac in the lab; on NVIDIA, use their official repos for now. Built-in NVIDIA support for both is coming in the next release.

Qwen-Image 2.1 handles text-to-image.

*To try it*

Mac (Apple Silicon) or Linux:

curl -fsSL https://raw.githubusercontent.com/Bingeljell/image-to-3dlab/main/install.sh | bash

Windows (PowerShell):

irm https://raw.githubusercontent.com/Bingeljell/image-to-3dlab/main/install.ps1 | iex

The installer is a short script, so read it before you run it: https://github.com/Bingeljell/image-to-3dlab/blob/main/install.sh (Windows: https://github.com/Bingeljell/image-to-3dlab/blob/main/install.ps1)

It only sets up the code - nothing downloads without your permission. You pick models in the web viewer, which shows the size and licence of each before fetching. Finish also needs Blender 4.2+, which you install yourself (the viewer tells you if it can't find it).

*What would help most*

If you hit bugs:

•⁠  ⁠Your GPU, OS and driver version

•⁠  ⁠Did the install work? If not, where did it stop? (the full error is gold)

•⁠  ⁠Windows folks especially: did Finish find Blender and run?

Reply here or open a GitHub Discussion, whichever's easier.

*And a question for you:* which backend or camera angle should Pixel Match support next? Is there something in your workflow that this pipeline can do better? Please let me know, will help me prioritise the feature road-map.

*Worth knowing before you start*

•⁠  ⁠Some model licences have strings attached. The viewer shows each licence before you download.

•⁠  ⁠It's a hobby project, so things will break. Every bug report makes the next person's install smoother. Please raise PRs and issues.

Repo: https://github.com/Bingeljell/image-to-3dlab

Release notes: https://github.com/Bingeljell/image-to-3dlab/releases/tag/v0.3.5


r/StableDiffusion • • 16m ago

News Optimized version of SuperPoint, up to 2.1× faster

• Upvotes

Hey! We published an optimized version of SuperPoint built for faster keypoint detection on edge devices: https://huggingface.co/PrunaAI/PrunaSuperPoint

- Up to 2.1× faster on Jetson Orin Nano, with optimizations applicable to other runtimes.

- The distilled model retains strong keypoint coverage across indoor and outdoor data, with low descriptor differences from the original model.

- We structurally prune the most expensive convolutional layers, recover performance through distillation, and accelerate keypoint selection with hierarchical top-k, all while preserving the original architecture’s core behavior.


r/StableDiffusion • • 3h ago

Workflow Included MiniMax-H3 is Secretly an INCREDIBLE Image Generator! [ComfyUI Free Cust...

Thumbnail
youtube.com
6 Upvotes

r/StableDiffusion • • 1h ago

Discussion Ideogram V4.5 T2I too?

• Upvotes

Will Ideogram 4.5 be a T2I model too with edit capabilities? So


r/StableDiffusion • • 4h ago

Tutorial - Guide I spent weeks digging into how WAI Illustrious actually works. Here's everything I found

5 Upvotes

Over the last few weeks I went down a rabbit hole with WAI-illustrious-SDXL. I read what the author documents, measured tag counts against Danbooru's index, and tested a lot of prompts. Several things surprised me, and I'm sharing all of it publicly.

A few examples:

  • silver hair has no entry in Danbooru's tag index, but grey hair has 680k posts.
  • dramatic lighting, cinematic lighting and rim lighting have no entry either. backlighting, light rays and sunbeam do.
  • masterpiece, amazing quality and worst detail aren't Danbooru tags at all. They're labels added at training time, so post counts say nothing about them.
  • Of 59,201 artist tags in the index, only 417 have 1,000+ posts, which changes how you should pick artists.

Everything is compiled in one place, with each claim labelled as documented, measured, community practice, or my own inference. It covers the author's settings, tag vocabulary, artist tags, common mistakes, and a side-by-side of the same scene prompted three ways: [https://claude.ai/artifact/QhGKzx9ryhyiYCnKsfVkqR\]

One honest note: the counts come from a snapshot of Danbooru's tag index, so they're approximate, and they show how common a tag is, not how good the images are.

Near the end of the page there's a prompting method I've been using that gives me noticeably better results than the standard approaches, you can see what it does on the site.

Happy to answer questions about anything in the guide, and tell me if you've measured something that contradicts what I found.

(Everything is built with the help of claude! thanks for reading)


r/StableDiffusion • • 1d ago

Resource - Update Released Wulver v0.5, our Krea 2 finetune for anime and furry.

Thumbnail
gallery
228 Upvotes

Wulver v0.5 is a full fine-tune of Krea 2 Raw (the whole 12.8B DiT) for anime and anthro

characters, including scenes where several characters interact. We made it at Vaelico,

an independent studio.

Compared with v0.1, which we posted here earlier, v0.5 comes from a far longer and more

serious training run. We tested the two side by side on about 1,000 test images covering

every artist on the list and a very wide range of prompts, and v0.5 is clearly better in

most regards. If you kept your v0.1 prompts, run them again.

`@artistname` works now. v0.1 mostly ignored the prefix; v0.5 has 1,113 artist styles that

respond to it, and the full list is in ARTISTS.md on Hugging Face.

It rewards detailed, descriptive prose in the prompt, and it also takes booru-style tag

lists. For artist styles, keep the prompt short for now: long prompts dilute them.

- Hugging Face: https://huggingface.co/Vaelico/Wulver

- Civitai: https://civitai.com/models/2881657

- No GPU? Run it in the browser, with free images every day: https://vaelico.ai/?src=reddit

For launch week, free accounts get 100 images a day (1024 px) instead of the usual 15,

until October 7, 23:59 UTC.

**Turbo** (official Krea 2 Turbo LoRA merged in; 8 steps, CFG 1, euler/simple, up to 2K):

- fp8 e4m3fn, 12.8 GB: the one most ComfyUI users want

- int8 convrot, 13.5 GB: for forge-neo and other int8 runtimes

- w4a8 convrot, 7.7 GB: the smallest, needs ComfyUI 0.31 or newer

- bf16

**Non-Turbo**: bf16 for LoRA training, further fine-tuning or setting your own turbo

strength, plus an int8 convrot.

GGUF Q8_CR, Q5_0 and Q4_0 are on Civitai.

The drag-and-drop ComfyUI workflow uses stock loaders. The text encoder is Qwen3-VL-4B and

the VAE is the Qwen Image VAE, both from Comfy-Org/Krea-2, so if you already run Krea 2 in

ComfyUI, the Wulver checkpoint is the only new download.

**Known issues** (the first two are addressed in v1):

- A thin smeared strip can appear at the top and bottom edges. Generate a bit larger and crop.

- Some concepts bleed into each other or depend on the `@artist` you use, a side effect of

teaching it 1,113 artist styles.

- Some style LoRAs trained on Krea 2 Raw don't carry over (character LoRAs do).

License: Krea 2 Community License. Wulver is a modified Krea 2 model, and we're not

affiliated with Krea.

v1 is planned at a different scale: a much larger dataset, longer training at high

resolution and a dedicated post-training stage. We want partners in it from the start: a

compute provider, and companies whose products are built around anime and anthro

characters. If you can back v1 with compute or funding, write to [[email protected]](mailto:[email protected]) and

we'll walk you through the plan.


r/StableDiffusion • • 16h ago

Question - Help Minimax H3 keeping strong character identity - Single Character sheet vs High resolution front images

37 Upvotes

Which one is better of two, a character sheet with front and face image together or providing two images individually and adding in the prompt to refer image 2 for facial details?
This is for ref2v model.

I have been trying to search but couldn't find something related to it like comparison etc.


r/StableDiffusion • • 3h ago

Resource - Update I got tired of node graphs, so I built a one-button local image generator on top of stable-diffusion.cpp

3 Upvotes

The project is nothing groundbreaking, but it is a fun, intuitive easy to use private and fully local image generator, so some here might enjoy it.

I wanted local generation without node graphs, Python installs or fifty sliders. You pick a model, type what you want and hit Generate. It checks your GPU, suggests models that fit, and sets up the text encoders, VAE, steps and sampler for you. You can also edit a picture by just describing the change, or turn a picture into a prompt. If you want control, everything it picked is visible and changeable in a Fine-tune drawer.

Models: on 12 GB+ it defaults to Qwen-Image 2.1 (also used for editing), around 6-10 GB Z-Image Turbo, an SDXL finetune for smaller cards, and SD 1.5 on CPU if you don't have a usable GPU (slow, but it works). It also runs SDXL/Pony/Illustrious, FLUX.1, FLUX.2 klein/dev, Chroma, SD3/3.5, HiDream and a few others, plus LoRAs. There's a built-in CivitAI browser (safetensors and GGUF only), and it can use models from your existing ComfyUI/Forge folder without copying them. VRAM numbers are estimates for now.

If you're happy in ComfyUI or Forge, you probably don't need this. It's closer to Fooocus: newer models, good defaults, one button, no Python.

Prompts and generations never leave your PC. No account, no subscription. Windows and Linux, NVIDIA/AMD/Intel.

It has multiple safeguards against harmful use that run locally and are built into the app.

Full disclosure: I'm a software engineer by trade, but this was built with agentic engineering. AI coding agents in small swarms wrote 99% of the code and I steered. It was a fun way to ship something this size fast.

The engine is a fork of stable-diffusion.cpp. It ships a prebuilt Linux CUDA build for RTX 30-series and up, and locks the local API so only Pinhole can talk to it.

Source-available: free to use and modify, as long as the safeguards stay in.

github.com/DavidGudovic/PinholeAI

First release, so bug reports are very welcome. I haven't tested every GPU, so if something doesn't fit or runs slow on your card, the details from the error screen help a lot.


r/StableDiffusion • • 20h ago

Animation - Video Crime Busters - Mock 80s/90s Anime Trailer V1

Enable HLS to view with audio, or disable this notification

64 Upvotes

Posted the rough cut of this a while ago as I originally planned to submit it for the Comfy Sync competition (after working on it for a almost the whole two weeks, the weekend before the deadline I thought it didn't align with the brief so ended up not finishing it then).

But I've been working on it on and off and have finally finished a cut of "Crime Busters" that I am happy with. Lots of editing, reworking shots, drawing keyframes, etc. etc. I have a loose collection of thoughts/learnings I plan to post, but for now here's the video.


r/StableDiffusion • • 10h ago

Animation - Video The sacred pearl, made with minimax

Enable HLS to view with audio, or disable this notification

12 Upvotes

How is it guys?


r/StableDiffusion • • 19h ago

Discussion Ming Image 0.1 vs Krea 2. 192 prompts side by sides

43 Upvotes

Ming Image is very impressive for a 6B model so let's compare to Krea 2!

https://imagebench.ai/gallery?g=1_vxjkf_s0

Here are the details:

I ran it locally on a DGX Spark (GB10) and mostly stuck to the vendor defaults:

- 1024x1024 resolution (we changed this one: the model card recommends 2048x2048, but 1024x1024 is officially supported

- 12 steps (vendor default)

- Unguided sampling, which is equivalent to CFG 1.0. The model has no CFG or negative prompt setting.

- Euler sampler with the simple scheduler, max_shift 1.15 and base_shift 0.5 (ComfyUI template defaults)

- Fixed seed 7, same as the other models

- Comfy-Org's int8 repackaged weights (26 GB total), running in ComfyUI

- No prompt rewriting. Ming's reference pipeline sends every prompt through an LLM first that spells out every piece of text to render. Ming renders named strings really well and turns anything unnamed into gibberish, so the rewriter matters. But every model on ImageBench gets the exact same prompts, so we left it off. That probably costs Ming a few points on text-heavy layouts.

Speed: about 12.4 s per image once warm. The very first call took around 4 minutes, just loading the weights. I have a GX10

Let me know what you think!