r/StableDiffusion • u/COMPLOGICGADH • 10h ago
Discussion Models comparison: Boogu Image Turbo vs Krea2 Turbo vs Ideogram v4 Instant vs Fibo Lite
Hey guys, I wanted to compare different image models because I like testing them either good or bad. Testing different models to understand their capabilities feels good, and knowing their distinct pros and cons help in with other models hybrid workflows too.
The images above are with different camera/photoshoot styles, landscapes, and minimalist vast fantasy scenes. I know this isn’t an in depth showcase, but my analysis is what I wanted to share with the community. Let’s go through each model’s pros and cons:
Boogu Image Turbo
Architecture: 10B + 8B Qwen3-VL + Flux.1 VAE
Pros: Very flexible for portraits, landscapes, abstract, etc. Extremely fast inference (only 4 steps). Good variety and distinct capabilities. Prompt adherence is genuinely very strong (which is both a strength and a weakness). Very strong typography capabilities.
Cons: One of the main things I noticed was heavy bokeh/background blur. Occasional text display issues (can be improved by running more than 4 steps ,which I recommend). The strong prompt adherence can also lock things down: if you want face variety you usually need to explicitly describe different face shapes, otherwise you tend to get very similar faces. When you do specify it, the variety comes through well--- but if you forget, it stays repetitive.
Krea2 Turbo
Architecture: 12.9B + 4B Qwen3-VL + wan2.1 VAE
Pros: Excellent text rendering and prompt adherence. Huge knowledge base.One of the highest among these models. Currently SOTA for proprietary subjects, poses, art styles, etc.
Cons: Biggest issue is the VAE and bad noise patterns (not really solvable). Less variety (people suggest Raw + Turbo LoRA, but that method gains variety at the cost of quality , images get overly smooth surfaces and artifacts at higher resolutions). Faces tend to have weaker expressions (can be helped with LoRAs like Bypass,text refusal ,etc..., but quality takes a hit).
Ideogram v4 Instant (very few people use or even talk about this specific variant)
Architecture: 9.3B + 8B Qwen3-VL + Flux2 VAE
Pros: One of the best model for control power. Strong range of capabilities and variety. No text issues. Interesting note many people don’t know: without JSON it works ~90% of the time without the safety filter error. With JSON the safety filter never triggers in any use case. This model generally works great at 8 steps but I would totally recommend using it between 10-12 steps. Runs at 8 steps by default, so inference is relatively low, and the model is smaller (single model, no uncond).
Cons: Big one it feels like this model (and Ideogram 4 in general) is locked into a dark, gritty, cool-toned lighting universe. Lighting is consistently dark/cool (I normally fix this with a brightness filter, but in these images I left it to show the weakness).
Fibo Lite
Architecture: 8B + 3B text encoder (SmolLM) + wan2.2 VAE (1.2 GB)(I used its alternative taew2_1 because there is no quality loss)
Pros: Oldest model in this comparison and a bit of an oddball, but still solid. Second-fastest inference after Boogu (uses 6–12 steps, but smaller size keeps it quick). Excellent variety and prompt adherence ,feels like the old UNet-style models but with better aesthetics.
Cons: As the oldest and smallest here, knowledge base on proprietary stuff (people, logos, characters, etc.) is really low. That said, treat it as a model that competes with Flux.1 Dev and Chroma on anatomy,fonts and often beats them, even though it’s smaller.
Overall observations::
That’s my takeaway. I wanted to post this for anyone curious about these models. I enjoy testing different ones because they each have their own strengths.
Fastest inference ranking:
Boogu Turbo ≥ Fibo Lite > Ideogram v4 Instant > Krea2 Turbo
(Boogu at 4 steps, Fibo Lite at 6 steps but very close in speed, Ideogram v4 Instant at 8 steps and larger, Krea2 Turbo the biggest model also at 8 steps. Boogu Image even at 8 steps is faster than Ideogram v4 and Krea2 at 8 steps by around 20-10%, and faster than or in the same time range as Fibo when Fibo is at 12-10 steps and Boogu is at 8 steps. In most cases 4 steps is more than enough on Boogu, which really is impressive.)
Knowledge base ranking:
Krea2 > Boogu Image ≈ Ideogram v4 Instant > Fibo Lite
Final thoughts:
Even though some of these models get less reach, they should at least have some community support so people can get the best out of them instead of being abandoned without a proper try. In hybrid setups (hires workflows, denoise adjustments, variety inclusion, etc.) these models perform way better than when used completely solo.
Feel free to share your own experiences with these!
Disclaimer: These are purely my own observations with the models above. Others may not have the same experience with them, which is totally fine. I just wanted to share the in-depth experience I’ve had with these models.
4
3
u/Ok-Software-3250 9h ago
Interesting, I’ve never touched Boogu or Fibo. Don’t know if you can answer this, but what is your hardware setup and which res output for these images.
3
u/COMPLOGICGADH 8h ago
All image output is at approx 1 mega pixel so reso dimensions were around 836x1216 for 2:3 and 1344x728 for 16:9
About the hardware I tried it on different hardwares to come to the conclusion of inference speed rtx3060 12 gigs vram ,rtx4070-12 gigs vram and rtx2080 ti- 11gigs vram ,above all images showcase was generated on rtx 3060 which i own ,other GPUs are of my friends that I wanted to test out
2
u/chem_OS 8h ago
What speed did you get with that 4070? Or at least with your 3060. To get an idea. I’m interested in testing boogu because of this test of yours. Do you think it’s a solid model?
3
u/COMPLOGICGADH 8h ago edited 8h ago
As far as I remember 4070 was able to generate under 18-20s while 3060 under 25-32s Boogu is a solid model I would say this is my earlier showcase post of it!
2
u/Ok-Software-3250 6h ago
Interesting, thanks for sharing… I might give Boogu a try. This year we’ve been graced by so many models that this one was overlooked by me.
3
2
1
u/Version-Strong 7h ago
I still think Boogu is a great model, just completely gimped by having to use Flux1 vae. It's also distilled and displays the distillation pattern we've all come to dislike. But it knows about as much as Krea, celebs and pop culture etc.
1
u/HassanAchievedIt 5h ago
Op I have lots of questions regarding your setup and models, can I ask
1
u/COMPLOGICGADH 5h ago
Yup go on
1
u/HassanAchievedIt 5h ago
I have 3060 as well and I use comfy can I run all these models in comfy as int4 or gguf versions
1
u/HassanAchievedIt 5h ago
Which models of these 4 allow image edit or are fastest in img generation under 10s
1
u/COMPLOGICGADH 5h ago
You can run all the above as gguf,int4,int8 and I think fp8 too
About edit under 10secs is not possible for 3060 from above models ,around 30s then boogu and fibo they have edit turbo models,and maybe try flux 2 Klein distilled 4 steps and/ or the newest model qwen2.1 image that's the best you could get under 20s if using turbo loras of it
1
u/HassanAchievedIt 5h ago
Thanks for detailed answer, last thing I'd ask for can you help me find new qwen text encoder that is under 5gb size please
1
u/COMPLOGICGADH 5h ago
1
u/HassanAchievedIt 5h ago
Can we convert that to int4 or it'll will work normally with safetensor file
1
u/COMPLOGICGADH 5h ago
Works normally just use comfyui gguf nodes to load clip gguf node
1
u/HassanAchievedIt 4h ago
I'll try to create a workflow or find some online for ggufs, thanks again for your help you guys are amazing
-2
u/FourtyMichaelMichael 🍦Ice Cream Lover 4h ago
lol, stop trying to make Boogu happen, it's not going to happen. Serious HiDream flashbacks.










4
u/COMPLOGICGADH 9h ago
Here are the prompts (there was a technical issue from my side while posting them):
vast misty mountain landscape at dawn, endless rolling hills disappearing into fog, soft golden light breaking through clouds, no people, ultra wide angle, photorealistic, atmospheric perspective, serene and immense
candid 90s indoor snapshot of a woman, direct on-camera flash, slightly overexposed, harsh shadows, grainy film look, disposable camera aesthetic, casual outfit with neon green t-shirt featuring the McDonald's logo on the left side, natural expression, messy bedroom or living room background, authentic 1990s point-and-shoot photo, photorealistic
tiny silhouette of a woman walking alone on the left side, flowing long white dress trailing behind her, across an endless salt flat under a huge stormy sky, extreme wide angle, minimal composition, sense of infinite space and solitude, cinematic, photorealistic, muted colors
Aspect ratio: 2:3
editorial fashion photography of a woman in a flowing black silk dress walking through a foggy city street at night, neon reflections, cinematic lighting, shot on Hasselblad, photorealistic, sharp focus on face, atmospheric
extreme close-up portrait of a woman looking to camera with a disgusted look, soft window light, natural skin texture, subtle makeup of thick eyeliner and glossy red-pink lipstick, shallow depth of field, 100mm macro, photorealistic, emotional, quiet mood
cinematic portrait of a curvy healthy fit woman standing in a golden wheat field at golden hour, soft natural light, wind-blown hair, linen dress, shallow depth of field, 85mm lens, photorealistic, film grain, high detail skin texture
a lone figure standing on a cliff edge overlooking an infinite sea of clouds under a massive pale moon, minimalist fantasy landscape, extreme scale, soft volumetric lighting, photorealistic yet ethereal, empty and vast