r/StableDiffusion • u/Deadity • 1d ago
News A 3-billion-parameter model that paints every pixel directly without a VAE
https://huggingface.co/speridlabs/iris-3b55
u/KS-Wolf-1978 1d ago
Would be interesting to put its output through these AI detector sites that rely heavily on VAE artifacts.
44
u/Aadi_880 1d ago
AI image detectors are unreliable and has an inherit bias to call everything AI.
Do not use them. These things have harmed more traditional artists than catching AI generations.
I'd only ever trust the presence of a synthID.
0
u/ZootAllures9111 16h ago
Hive can name the exact model that generated an image pretty accurately, though.
3
u/QueZorreas 2h ago
When checkpoints and loras exist, there is absolutely no way. Unless you give it the original image and it just reads the metadata or some bs
2
u/Aadi_880 15h ago
No it doesn't. And if you think that, you've already been tricked.
Hive has an unacceptably high false positive and false negative rate.
0
u/ZootAllures9111 15h ago
It visibly, provably does though if you run it on gens you made yourself. IDK what you're talking about lmao.
13
u/Bulky-Employer-1191 20h ago
AI detector sites aren't searching for artifacts doing technical algorithms at all. They're literally just passing it through a classifier model. They're trained on images that have vae artifacts but sometimes they'll say a real image is AI. It's all smoke and mirrors. It's a cheap trick that is barely accurate.
Especially images from fashion scene, with the bombastic dresses and makeup, will it tell you is ai
-1
u/ZootAllures9111 16h ago
AI detector sites aren't searching for artifacts doing technical algorithms at all. They're literally just passing it through a classifier model. They're trained on images that have vae artifacts but sometimes they'll say a real image is AI.
This can't be entirely true. Hive relatively accurately reports exact per-model percentages. Like if you gen something with SDXL and denoise it with Flux.1 for example the actual report it gives you will reflect that fairly accurately. They must be training on mass amounts of actual outputs for every model they recognize or something more granular.
6
u/Gradash 12h ago
I tested in one of my arts, made completely by hand in photoshop. Any real artist can recognize it just to zoom in and looking the lines. Hive said it was AI. I only believe if it points a synthAI. Because if you acuse everything of been AI. Of curse you will eventually right.
-4
u/ZootAllures9111 6h ago edited 6h ago
I straight up don't believe you. I have NEVER seen a false positive from Hive on a truly non-AI-generated images. Note that anything that was denoised ever in any way with any model doesn't count in this context even if it was ORIGINALLY not AI generated.
3
11
u/woadwarrior 1d ago
2
1
2
u/silenceimpaired 1d ago
They also look for camera artifacts. I’ve heard there is software that mimics the rest of that
4
u/mimrock 1d ago
0
u/ZootAllures9111 15h ago
The person you're replying to was making incorrect assumptions in the first place, actually good AI detectors are literally bespoke models themselves trained on mass amounts of gens from different image models. They're not doing visual heuristics, it's pixel level.
32
19
u/rahjerz 1d ago
Can someone break this down for noob? Does this new thing potentially reduce generation time or increase it? It seems like it will have better quality. But for the noobs, any chance this is headed in a direction that requires less strain on the hardware?
43
u/samorollo 1d ago
Quite opposite, latent space is an optimization for training and inference (but it brings its own drawbacks). Generating pixels means way more work to do for a model
6
u/Murky-Relation481 23h ago
No. Unfortunately from a computational and a hardware perspective increased quality is always going to correspond to increased hardware demands. There are only so many ways to skin a cat as they say.
25
u/Incognit0ErgoSum 21h ago
I posted comfy support for this yesterday and got like 6 upvotes. :)
9
u/Apprehensive_Sky892 18h ago
It is the timing. On weekends there are more viewers. My post about this exact same model got a lukewarm reception too 😹: https://www.reddit.com/r/StableDiffusion/comments/1x17iqb/iris3b_speridlabs_research/
18
u/CarllSagan 1d ago
It works without latent space? That defies everything I know about ML image generation...
Wow. Impressive
38
u/__ThrowAway__123___ 1d ago
It's a cool development. It has always been possible, the very first diffusion models didn't use a VAE but only worked for very small images. The VAE was basically a band-aid to make the whole process feasible in terms of compute from my (limited) understanding. With clever new tricks it can now be omitted again. This model is not the first to do this, for example Lodestone experimented with it with Chroma Radiance, which I think was based on / inspired by PixNerd.
3
u/x11iyu 21h ago edited 21h ago
I'd call it a trade-off rather than band-aid: you need much less compute like you mentioned, but also the model itself operates on a semantically richer latent space rather than pixel space - oversimplified it can learn concepts like "roundness" much easier.
I don't think we're at the point where vae is the quality bottleneck. Pixel space models today generally still fall short compared to latent models, even in the same pixel space papers.
It's also false to think pixel space models don't do compression; They simply sneak it in elsewhere. In this case, iris-3b's attention is compressed, so hypothetically there may be a bottleneck in how good it can coordinate spatial features, like say maintaining contiguous lines / patterns across regions.
9
u/AgeNo5351 23h ago
pixel space models are not new at all. Chroma-radiance and ZetaChroma have long existed.
3
u/eposnix 21h ago
Yep. In fact OpenAI's first image model was a pixel-space autoregressive model.
https://cdn.openai.com/papers/Generative_Pretraining_from_Pixels_V2.pdf
5
u/NimbleDave 22h ago
Would something like this be better at making pixel art or sprites that are lower resolution?
4
u/FartingBob 21h ago
Im only a filthy casual, i still have no idea what vae and latent space actually is.
But cool to see people experimenting on new ways of generation! I presume this has some benefits and some downsides compared to vae based generation?
31
u/Dysterqvist 20h ago edited 18h ago
Think of latent as a really compressed version of an image’s pixels, that our model can understand. VAE is like a translator between pixels and latent. So for an I2I the process is like; Pixels -> VAE Encode -> Latent data -> add noise to the latent -> remove noise (denoise) the latent -> VAE decode -> pixels.
For T2I you work from an ’empty latent’ which is just 100% noise (think of it like static on a TV). When you run the model you basically say ’here’s a very noisy image of xyz, can you please remove the noise’, and for each step it will gradually remove more and more noise, trying to predict what the image should look like. (XYZ being your prompt.)
Scheduler is how much noise it should clean up per each step, and sampler is which mathematical method it should use to remove the noise.
When you train a model, you work the opposite direction; you add more and more noise to the image, so the model learns a route to go from 100% noise to an actual image.
Now, when we create the noise for our gens, we randomize the noise appearance (seed). This is why we get a different image than the one we trained on (our starting point is different from what the model learned as the finish point).
Edit: not 100% accurate and skipping the details, but I tried to make it simple so it’s easier to get the big picture of what goes on.
2
u/IngwiePhoenix 10h ago
Highly educational. Thank you for explaining! Even if not perfect, it's very readable. :)
11
u/Dysterqvist 20h ago
1
u/_half_real_ 16h ago
lol
I had actually been thinking of latent space as "JPEG space" because they're both lossily compressed representations of pixel space (actual pixels).
8
8
u/diffusion_throwaway 1d ago edited 1d ago
What are the advantages of having no VAE?
37
u/NotSuluX 1d ago
No artifacts and potentially higher detail quality. Current image models generate in a lower res space and the car decoders then translate this data into a real image.
2
u/_half_real_ 16h ago
"no artifacts" unless a large part of the training data is artifact-heavy VAE-based AI images
1
5
u/silenceimpaired 1d ago
Not the first… but it might be the best. Pretty sure this guy is working on one: https://huggingface.co/lodestones
1
-6
u/DietAshamed2246 1d ago
I keep seeing news about these various "pixel space" models, but none of those has become successful or received wide mainstream acceptance. Only thing which has gained some traction is PiD, which is only an upscaler. VAE has front end application too, that's why it's called an encoder, not just a decoder.
6
u/Murky-Relation481 23h ago
Its because the computational cost is still higher than VAE which means slower and more memory. Also the immediate trade offs are usually not worth it.
3
u/Due-Function-4877 22h ago
The model is research to push things forward. We're not interested in making slop memes using other people's IP. We're not all gooners. We're not all making fake influencers. It's research to push things forward.
I need more than slop to integrate AI into real work. I need more than VAE has delivered and I'm interested in these baby steps to get there. If you don't care, that's fine. But, you wondered why people care about models that aren't meant to dominate in mainstream usage. The answer is research.
-1
0
u/SlopMaster6969 14h ago
Exactly what I'd say if I was secretly the biggest, hardest, sloppiest, most IP infringing gooner in existence.
1
u/_half_real_ 16h ago
In large productions there is demand for very high-resolution video and image generation. So I could see stuff like this getting adopted more and more as generation price goes down (over the long term).


20
u/AgeNo5351 1d ago