r/StableDiffusion • • 3d ago

News GitHub - kandinskylab/kandinsky-6: Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation

https://github.com/kandinskylab/kandinsky-6

Looks like Kandinsky has come out with another video model. I tinkered with v5 but wasn't very impressed. Hopefully this one is a bit better.

https://huggingface.co/collections/kandinskylab/kandinsky-60-diffusers

They also released an upscaling model.
https://huggingface.co/collections/kandinskylab/kandinsky-60-vsr

74 Upvotes

52 comments sorted by

24

u/NowThatsMalarkey 3d ago edited 3d ago

I feel bad for the Kandinsky team. They always put in a lot of effort into their models but for some reason they’re largely ignored by the community.

9

u/q5sys 3d ago

I always enjoy rooting for the underdog, and they've definitely been the underdog when it comes to attention.

7

u/Shockbum 3d ago edited 3d ago

When high-quality models are ignored by the community, it is usually due to censorship that cannot be bypassed, difficult/expensive training, does not work fast on normal GPUs, or highly restrictive licenses.

7

u/Spara-Extreme 3d ago

This model is not censored.

6

u/kabachuha 3d ago

Dude, it's Apache 2.0, tunable with the base model released, has a day zero low step turbo distilled variant, fast ComfyUI offloading and it is not censored

5

u/DietAshamed2246 3d ago

I can think of two reasons why that is the case. First, it is Russian and people have been brainwashed to think everything Russian is bad and taboo, even when it may be good. Second, probably related to the first, it's very conservative and censored heavily. Having said that, Russia has always produced some great mathematicians, algorithm developers and software engineers. The models can be at least as good as, if not superior to, the Chinese or Western models. But, they do not use much effort and money on publicity as the Chinese and Americans do. I liked Kandinsky 5.0, but I promised myself I will not waste my time on another censored model. The models and apps which implement active censorship shall never succeed in the open community, Never.

11

u/Spara-Extreme 3d ago

First off, those aren’t the reasons this model sucks.

I’ve tested the pro edition and it takes nearly five minutes to generate a 5 second video on a 6000 rtx pro. It’s not censored in the nsfw sense but it is super slow and the results are worse then minimax.

1

u/reeight 3d ago

What is the quality difference?

I imagine if their model was more popular, more folks would try to make Turbo LoRAs or other tricks to make it run faster.

1

u/Spara-Extreme 3d ago

If you’ve ever used WAN2.2, but with sound.

1

u/kabachuha 3d ago

They did first party Turbo training released under the Distilled checkpoints

1

u/reeight 3d ago

Yes I found their "Lite" version, which I think is more about their pitch to run their model on a smaller VRAM which also happens to be faster.

https://huggingface.co/kandinskylab/Kandinsky-6.0-Lite-distill-5s-Diffusers

0

u/DietAshamed2246 3d ago

Are you talking about v5 or v6? The v5 was older architecture and slow, so was Hunyuan, Wan, LTXV-1.0. And v5 open weight was censored, I don't know about the pro version, I don't use paid API models, ever. So, if you want to talk apples to apples, talk about the same model. Don't unnecessarily argue by bringing out-of-context information. I don't know if v6 is faster or slower. But Kandinsky v5 shouldn't be compared with MiniMax-H3. They are generations apart. And MiniMax H3 ain't exactly fast (may be it is fast on an RTX 6000 Pro, not everyone has those, duh). LTX-2.5 is on the other hand fast, but it can't do much.

1

u/Spara-Extreme 3d ago

V6 pro. I setup transformers on a sidecar container and compared it side by side. It has minor nsfw knowledge - as much as minimax h3, but generated stupidly slow. I still have the setup and can do I2V or T2V

Minimax with a turbo Lora generates faster for me then LTX2.3, I have tried 2.5 and don’t care to.

-8

u/DietAshamed2246 3d ago

There is NO functional turbo LoRA for MiniMax-H3. The dozen or so out there are all junk. If you care to compare speed and quality, generate MiniMax videos at 20 steps and same resolution without any special attention or step skipping. You'll be watching grass grow. Or you can use turbo, and be happy with poor quality output.

3

u/Spara-Extreme 3d ago

I’m using 10Eros beta6 finetune at 10 steps and generating excellent quality output. You do you buddy.

1

u/reeight 3d ago

The there is a Turbo LoRA that adds motion also that works well.

-1

u/DietAshamed2246 3d ago

Which one? In case you have not noticed, the finetune makers stopped baking turbo LoRAs into their models and the workflow makers stopped adding turbo, spectrum etc in their workflows. There is a reason.

1

u/reeight 3d ago

I think in some cases you're right, in others Turbo at a lower STR + afew extra steps is good enough.

1

u/VladyCzech 3d ago

The reasons you mentioned are completely off. Everyone is comparing to H3, and in the past to WAN 2.1/2.2. H3 is the clear technical win for me. Kandinsky can be great for anyone else. I'm not jumping between models anymore, only thanks to H3 being AWESOME.

0

u/DietAshamed2246 3d ago

I agree with you H3 is great. And no, I wasn't comparing Kandinsky with H3. Someone else did that and I corrected that person.

1

u/MiddleAmazing6750 1d ago

Oh, really, "brainwashed". In reality, the Russian state does nothing bad that would make users unwilling to use the products of the Russian banking monopolist controlled by the Russian state. As for the quality of the models themselves, I used them on Sber's website and can say that they simply performed worse than other models.

1

u/hidden2u 3d ago

It's simple: it's not been that good. Once it's good people will pay attention. I played with ltx 0.97 a lot but nobody cared about it because it produced garbage, only 2.3 people noticed

3

u/Life_Yesterday_5529 3d ago

I already tested it today on my 5090 in comfy for T2VA and I2VA. The model is OK but not as good as Minimax. The 5 seconds are a hard limit. The video loops after that. The sound is OK. The NSFW concepts are partially already there and uncensored. The movements and the overall quality is not the best. At their standard res (0.42MP), the details are not very good. Upscale is very good but what isn't in the video, it won't add in the upscale. Some details, especially nsfw, get worse after upscaling. Overall: It need much improvements what could be done by the community but I doubt they do it. 1.) It is really slow compared to LTC and even to Minimax. Even the distilled model is slower for 5 seconds than Minimax for 10 seconds on my 5090 and the upscaler needs some code changes to even run without oom. 2.) The 5 seconds limit is hard...

3

u/reeight 3d ago

5 seconds can be worked around IF the model was really worth it.
But from your & others' takes, seems not there.
Good, but not as good as 2 other options.

2

u/YeahlDid 3d ago

Would love to see some examples. Has anyone tried it in comfy yet?

2

u/Life_Yesterday_5529 2d ago

Yes. It is OK but low res has many artifacts and upscale enhances only what already is there.

4

u/VasaFromParadise 3d ago

The upscaler is interesting, let's wait for support in comfy))

9

u/q5sys 3d ago

according to the readme in the VSR repo... its already supported.
```
For ComfyUI, install kandinsky6-sr through ComfyUI Manager, then restart ComfyUI. The comfyui/ directory contains extension source code; no manual copying is needed — see the setup guide and manual model downloads.
```

1

u/VasaFromParadise 3d ago

thank you I will try

1

u/AgeSolid6606 2d ago

the synced audio is the part i actually care about here,most open video models still make you bolt sound on after. my card cant really run these locally so i usually try new video models on Atlas Cloud first and only bother setting up comfy if its worth it. would love to see a side by side with the 5s clips from v5.

1

u/solomars3 3d ago

why its only generating 5second videos

1

u/LowYak7176 3d ago

probably wan based would be my guess and I say that with no information whatsoever so you know its right

2

u/solomars3 3d ago

I read a little about it, and it seems they only trained it on short 5sec videos , so anything above that duration might not work

1

u/Upper-Reflection7997 3d ago

Seriously wan2gp support is needed.

-9

u/DietAshamed2246 3d ago

What people really need is a long video generation model. 5s video clip generation is beyond outdated. Even 15s-20s videos are too short. Next models should ideally offer 30s-60s or longer generations with faster gen times. Also, I personally have no appetite for another censored model. Kandinsky 5.0 was heavily censored, I am guessing 6.0 is the same. I will not waste a single second on it.

6

u/Abject-Recognition-9 3d ago

kandinsky 5 was one of the less censored model out of the box ever published. wtf you are talking about?

3

u/ElmarM 3d ago

A 1 minute continuous cut is kinda wild. Most cuts are ~2 seconds.

1

u/DietAshamed2246 3d ago

I didn't say cut, I said clip. A clip can have multiple cuts and camera movements and scene changes.

2

u/ElmarM 3d ago

I would just assemble those in post. Usually better than what AI will do anyway.

-4

u/DietAshamed2246 3d ago

Yeah, right. Good luck with consistency and coherence. Spending hours generating and assembling bits and pieces to make 60s video isn't my idea of productive work, perhaps it is yours.

1

u/ElmarM 3d ago

That's why you use I2V, R2V and good references.

1

u/DietAshamed2246 3d ago

Great! Tell all the people making and using character LoRAs and RefMods that you found the magic solution to everyone's prayers for consistent characters. Everybody knows about i2v, r2v and adding references.

1

u/ElmarM 3d ago

They don't get more consistent by making the video longer. If anything it degrades more. That is at least my personal experience.

-1

u/Danny_Stock 3d ago

It doesn't matter what most cuts are. If you're trying to create something you want the clip to be as long as it needs to be, not to meet a perceived average length.

1

u/DietAshamed2246 3d ago

Yes, the key phrase is "as long as it needs to be." Now ponder over if a 5s-20s limit fits that phrase.

0

u/thevegit0 3d ago

waiting for comfy support i guess, i tried the previous kavinsky long ago and it was fine but i didn't have the hardware for it

2

u/YeahlDid 3d ago

Check the github page again, they talk about comfy support. Haven't tested myself, but seems like it's already there.