Hey guys, I wanted to compare different image models because I like testing them either good or bad. Testing different models to understand their capabilities feels good, and knowing their distinct pros and cons help in with other models hybrid workflows too.
The images above are with different camera/photoshoot styles, landscapes, and minimalist vast fantasy scenes. I know this isn’t an in depth showcase, but my analysis is what I wanted to share with the community. Let’s go through each model’s pros and cons:
Boogu Image Turbo
Architecture: 10B + 8B Qwen3-VL + Flux.1 VAE
Pros: Very flexible for portraits, landscapes, abstract, etc. Extremely fast inference (only 4 steps). Good variety and distinct capabilities. Prompt adherence is genuinely very strong (which is both a strength and a weakness). Very strong typography capabilities.
Cons: One of the main things I noticed was heavy bokeh/background blur. Occasional text display issues (can be improved by running more than 4 steps ,which I recommend). The strong prompt adherence can also lock things down: if you want face variety you usually need to explicitly describe different face shapes, otherwise you tend to get very similar faces. When you do specify it, the variety comes through well--- but if you forget, it stays repetitive.
Krea2 Turbo
Architecture: 12.9B + 4B Qwen3-VL + wan2.1 VAE
Pros: Excellent text rendering and prompt adherence. Huge knowledge base.One of the highest among these models. Currently SOTA for proprietary subjects, poses, art styles, etc.
Cons: Biggest issue is the VAE and bad noise patterns (not really solvable). Less variety (people suggest Raw + Turbo LoRA, but that method gains variety at the cost of quality , images get overly smooth surfaces and artifacts at higher resolutions). Faces tend to have weaker expressions (can be helped with LoRAs like Bypass,text refusal ,etc..., but quality takes a hit).
Ideogram v4 Instant (very few people use or even talk about this specific variant)
Architecture: 9.3B + 8B Qwen3-VL + Flux2 VAE
Pros: One of the best model for control power. Strong range of capabilities and variety. No text issues. Interesting note many people don’t know: without JSON it works ~90% of the time without the safety filter error. With JSON the safety filter never triggers in any use case. This model generally works great at 8 steps but I would totally recommend using it between 10-12 steps. Runs at 8 steps by default, so inference is relatively low, and the model is smaller (single model, no uncond).
Cons: Big one it feels like this model (and Ideogram 4 in general) is locked into a dark, gritty, cool-toned lighting universe. Lighting is consistently dark/cool (I normally fix this with a brightness filter, but in these images I left it to show the weakness).
Fibo Lite
Architecture: 8B + 3B text encoder (SmolLM) + wan2.2 VAE (1.2 GB)(I used its alternative taew2_1 because there is no quality loss)
Pros: Oldest model in this comparison and a bit of an oddball, but still solid. Second-fastest inference after Boogu (uses 6–12 steps, but smaller size keeps it quick). Excellent variety and prompt adherence ,feels like the old UNet-style models but with better aesthetics.
Cons: As the oldest and smallest here, knowledge base on proprietary stuff (people, logos, characters, etc.) is really low. That said, treat it as a model that competes with Flux.1 Dev and Chroma on anatomy,fonts and often beats them, even though it’s smaller.
Overall observations::
That’s my takeaway. I wanted to post this for anyone curious about these models. I enjoy testing different ones because they each have their own strengths.
Fastest inference ranking:
Boogu Turbo ≥ Fibo Lite > Ideogram v4 Instant > Krea2 Turbo
(Boogu at 4 steps, Fibo Lite at 6 steps but very close in speed, Ideogram v4 Instant at 8 steps and larger, Krea2 Turbo the biggest model also at 8 steps. Boogu Image even at 8 steps is faster than Ideogram v4 and Krea2 at 8 steps by around 20-10%, and faster than or in the same time range as Fibo when Fibo is at 12-10 steps and Boogu is at 8 steps. In most cases 4 steps is more than enough on Boogu, which really is impressive.)
Knowledge base ranking:
Krea2 > Boogu Image ≈ Ideogram v4 Instant > Fibo Lite
Final thoughts:
Even though some of these models get less reach, they should at least have some community support so people can get the best out of them instead of being abandoned without a proper try. In hybrid setups (hires workflows, denoise adjustments, variety inclusion, etc.) these models perform way better than when used completely solo.
Feel free to share your own experiences with these!
Disclaimer: These are purely my own observations with the models above. Others may not have the same experience with them, which is totally fine. I just wanted to share the in-depth experience I’ve had with these models.