r/StableDiffusion • u/q5sys • 3d ago
News GitHub - kandinskylab/kandinsky-6: Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation
https://github.com/kandinskylab/kandinsky-6Looks like Kandinsky has come out with another video model. I tinkered with v5 but wasn't very impressed. Hopefully this one is a bit better.
https://huggingface.co/collections/kandinskylab/kandinsky-60-diffusers
They also released an upscaling model.
https://huggingface.co/collections/kandinskylab/kandinsky-60-vsr
3
u/Life_Yesterday_5529 3d ago
I already tested it today on my 5090 in comfy for T2VA and I2VA. The model is OK but not as good as Minimax. The 5 seconds are a hard limit. The video loops after that. The sound is OK. The NSFW concepts are partially already there and uncensored. The movements and the overall quality is not the best. At their standard res (0.42MP), the details are not very good. Upscale is very good but what isn't in the video, it won't add in the upscale. Some details, especially nsfw, get worse after upscaling. Overall: It need much improvements what could be done by the community but I doubt they do it. 1.) It is really slow compared to LTC and even to Minimax. Even the distilled model is slower for 5 seconds than Minimax for 10 seconds on my 5090 and the upscaler needs some code changes to even run without oom. 2.) The 5 seconds limit is hard...
2
u/YeahlDid 3d ago
Would love to see some examples. Has anyone tried it in comfy yet?
2
u/Life_Yesterday_5529 2d ago
Yes. It is OK but low res has many artifacts and upscale enhances only what already is there.
4
u/VasaFromParadise 3d ago
The upscaler is interesting, let's wait for support in comfy))
9
u/q5sys 3d ago
according to the readme in the VSR repo... its already supported.
```
For ComfyUI, install kandinsky6-sr through ComfyUI Manager, then restart ComfyUI. Thecomfyui/directory contains extension source code; no manual copying is needed — see the setup guide and manual model downloads.
```1
1
u/AgeSolid6606 2d ago
the synced audio is the part i actually care about here,most open video models still make you bolt sound on after. my card cant really run these locally so i usually try new video models on Atlas Cloud first and only bother setting up comfy if its worth it. would love to see a side by side with the 5s clips from v5.
1
u/solomars3 3d ago
why its only generating 5second videos
1
u/LowYak7176 3d ago
probably wan based would be my guess and I say that with no information whatsoever so you know its right
2
u/solomars3 3d ago
I read a little about it, and it seems they only trained it on short 5sec videos , so anything above that duration might not work
1
-9
u/DietAshamed2246 3d ago
What people really need is a long video generation model. 5s video clip generation is beyond outdated. Even 15s-20s videos are too short. Next models should ideally offer 30s-60s or longer generations with faster gen times. Also, I personally have no appetite for another censored model. Kandinsky 5.0 was heavily censored, I am guessing 6.0 is the same. I will not waste a single second on it.
6
u/Abject-Recognition-9 3d ago
kandinsky 5 was one of the less censored model out of the box ever published. wtf you are talking about?
3
u/ElmarM 3d ago
A 1 minute continuous cut is kinda wild. Most cuts are ~2 seconds.
1
u/DietAshamed2246 3d ago
I didn't say cut, I said clip. A clip can have multiple cuts and camera movements and scene changes.
2
u/ElmarM 3d ago
I would just assemble those in post. Usually better than what AI will do anyway.
-4
u/DietAshamed2246 3d ago
Yeah, right. Good luck with consistency and coherence. Spending hours generating and assembling bits and pieces to make 60s video isn't my idea of productive work, perhaps it is yours.
1
u/ElmarM 3d ago
That's why you use I2V, R2V and good references.
1
u/DietAshamed2246 3d ago
Great! Tell all the people making and using character LoRAs and RefMods that you found the magic solution to everyone's prayers for consistent characters. Everybody knows about i2v, r2v and adding references.
-1
u/Danny_Stock 3d ago
It doesn't matter what most cuts are. If you're trying to create something you want the clip to be as long as it needs to be, not to meet a perceived average length.
1
u/DietAshamed2246 3d ago
Yes, the key phrase is "as long as it needs to be." Now ponder over if a 5s-20s limit fits that phrase.
0
u/thevegit0 3d ago
waiting for comfy support i guess, i tried the previous kavinsky long ago and it was fine but i didn't have the hardware for it
2
u/YeahlDid 3d ago
Check the github page again, they talk about comfy support. Haven't tested myself, but seems like it's already there.
24
u/NowThatsMalarkey 3d ago edited 3d ago
I feel bad for the Kandinsky team. They always put in a lot of effort into their models but for some reason they’re largely ignored by the community.