🔬Kandinsky 6.0 Video: MIT-Licensed Joint Audio-Video Models
TL;DR
Kandinsky 6.0 Video is a family of open diffusion models that generate video and synchronized audio together: Lite at 3B parameters and Pro at 29B. RL post-training cut Pro's speech word error rate by 47%.
Kandinsky 6.0 Video is a family of open diffusion models that generate video and synchronized audio together: Lite at 3B parameters and Pro at 29B. RL post-training cut Pro's speech word error rate by 47%. Code, checkpoints and Diffusers support are public.
Key Points
Outputs 5-second Full-HD (1920x1080) clips through built-in super-resolution, with 44 kHz audio and lip-sync
Dual-stream CrossDiT joins a pretrained video stream and a new audio stream by bidirectional cross-attention
Pipeline runs SFT, RL post-tuning, then distillation
Human evals show a significant speech-quality edge over LTX 2.5
Released under the MIT license
Why It Matters
Joint audio-video generation was a closed-model feature a year ago. An MIT-licensed 29B option puts it in reach of anyone with a few GPUs.
Quick Facts
Comments
Be the first to comment
Enjoyed this article?
Get it daily. 7am. Free. Reads in 5 minutes.