Sber's Kandinsky Lab released Kandinsky 6.0 Video, a family of diffusion models that generate video and synchronized audio in one pass. It ships as Lite (3B parameters) and Pro (29B parameters) and produces 5-second clips with 44 kHz audio, including lip-sync, in text-to-audio-video and image-to-audio-video modes. A separate super-resolution model raises output to Full-HD (1920x1080). Code, checkpoints and diffusers integration are released under the MIT license.
A workable dense video generator from Sber. Treat the modality benchmarks above as the leading indicator of fit — composite scoring across modalities is still maturing. Newly released, so production-readiness is still being shaken out.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
No benchmark data available for this model yet.
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | VRAM |
|---|---|---|
| ACEMAGIC M1A Pro (i9-13900HK + ARC A770)ACEMAGIC | SS | 0.5 GB |
| Acer Veriton GN100 AI MiniAcer | SS | 0.5 GB |
| AMD Instinct MI300XAMD | SS | 0.5 GB |
| AMD Instinct MI325XAMD | SS | 0.5 GB |
| AMD Instinct MI355XAMD | SS | 0.5 GB |
Cheapest current cloud rentals with at least 1 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA GeForce RTX 3060Vast.ai · Spot · 12 GB VRAM | $0.07 |
NVIDIA GeForce RTX 5060 TiVast.ai · Spot · 16 GB VRAM | $0.07 |
NVIDIA GeForce RTX 4070Vast.ai · Spot · 12 GB VRAM | $0.08 |
NVIDIA GeForce RTX 3080Vast.ai · Spot · 10 GB VRAM | $0.08 |
NVIDIA GeForce RTX 3060Vast.ai · On-Demand · 12 GB VRAM | $0.08 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
Sber's Kandinsky Lab released Kandinsky 6.0 Video, an open-weights foundation diffusion model family designed to generate temporally synchronized video and 44 kHz high-fidelity audio within a single unified pass. Standard open video generators produce mute video tracks that require separate Foley, ambient sound, or speech generation models in downstream post-production. Kandinsky 6.0 bypasses this pipeline bottleneck by coupling acoustic and visual generation within the diffusion process itself, enabling accurate environmental audio, sound effects, and lip-sync synchronization directly from text or image prompts.
The model is distributed under the permissive MIT license, granting unrestricted commercial and self-hosted deployment. The family ships across two main tiers: Kandinsky 6.0 Video Lite (3B parameters) and Kandinsky 6.0 Video Pro (29B parameters). Both models produce 5-second video clips at 24 frames per second (125 total frames) paired with uncompressed 44 kHz audio. Native generation handles Standard Definition (864x480) and High Definition (1280x720) outputs, which can be upscaled to 1080p Full HD using an accompanying plug-in super-resolution model.
Kandinsky 6.0 Video replaces conventional single-modality video transformers with a dual-stream CrossDiT (Cross-Diffusion Transformer) architecture. The backbone pairs a pretrained visual diffusion transformer stream with a dedicated audio diffusion transformer stream. These two pathways communicate continuously at every stage of the denoising trajectory through bidirectional cross-attention mechanisms. The cross-attention layers map temporal frame tokens to corresponding audio spectrogram patches, preserving semantic parity and millisecond-level alignment throughout the 125-frame diffusion process.
Training followed a multi-stage regimen designed to preserve unimodal generation fidelity while enforcing strict cross-modal cohesion:
lite-distill and pro-distill) compress the standard 50-step diffusion process down to low-step regimes for low-latency generation.To maintain compute efficiency across varying GPU generations, Kandinsky 6.0 integrates custom attention kernels:
Kandinsky 6.0 Video targets production media generation workflows where audio-visual synchronization is mandatory:
kandinsky6-sr latent model scales 480p and 720p generations up to clean 1920x1080 Full HD footage without corrupting fine textural details or temporal coherence.To run Kandinsky 6.0 Video locally, your hardware must handle both the dense visual transformer and the parallel audio generation stream. When running dense generative checkpoints, memory management during latent caching dictates whether a consumer workstation can process full sequences without out-of-memory (OOM) faults.
The non-distilled models run at 50 diffusion steps by default at 864x480 resolution (125 frames). VRAM consumption varies heavily based on whether you employ module offloading (enterprise cards) or sequential block offloading (consumer cards):
Because full 29B FP16 weights exceed standard consumer capacities, choosing the best quantization for Kandinsky 6.0 Video is essential for sub-enterprise deployments:
--enable-cpu-offload to swap transformer layers between system memory and GPU VRAM dynamically. This is required on any single 24 GB or 80 GB card running the full unquantized Pro pipeline.Unlike text LLMs measured by tokens per second, Kandinsky 6.0 Video performance is benchmarked by total clip generation latency for a complete 5-second, 125-frame render (excluding disk writes):
| Hardware Platform | Memory Config | Lite SD (864x480) | Lite Full HD (1080p SR) | Pro SD (864x480) | Pro Full HD (1080p SR) | Offload Strategy |
|---|---|---|---|---|---|---|
| NVIDIA RTX 5060 Ti | 16 GB VRAM | 1,310s | 1,774s | 3,080s | 3,530s | Block Offload |
| NVIDIA RTX 4090 | 24 GB VRAM | 437s | 578s | 936s | 1,247s | Block Offload |
| NVIDIA RTX 5080 | 16 GB VRAM | 577s | 770s | 1,336s | 1,546s | Block Offload |
| NVIDIA RTX 5090 | 32 GB VRAM | 309s | 406s | 754s | 854s | Block Offload |
| NVIDIA RTX PRO 6000 | 96 GB VRAM | 242s | 387s | 621s | 765s | Module Offload |
| NVIDIA A100 | 80 GB VRAM | 532s | 664s | 972s | 1,106s | Module Offload |
| NVIDIA H100 | 80 GB VRAM | 239s | 284s | 356s | 402s | Module Offload |
For typical local practitioner workstations, the best GPU for Kandinsky 6.0 Video Pro is the RTX 5090 or RTX 4090. The RTX 5090's 32 GB buffer and memory bandwidth yield a complete 5-second SD clip in roughly 12.5 minutes on the full Pro model, while the distilled weights cut that generation time significantly.
Practitioners have three main pathways to run the model locally:
kandinsky6 and kandinsky6-sr via ComfyUI Manager. Nodes handle cross-attention audio decoding and automatically route the output audio stream into the multiplexed MP4 container.1 git clone https://github.com/kandinskylab/kandinsky-6.git2 cd kandinsky-63 just setup4 just download pro-distill5 just generate "A sports car accelerating down an empty tunnel, engine roaring" --config kandinsky/configs/devices/rtx-4090.yaml --out clip.mp4
vllm-omni for headless multi-GPU inference and API hosting, utilizing HTTP multi-part requests to serve raw MP4 payloads directly to client applications.While tools like Ollama dominate local text-only LLM deployments, video-diffusion architectures require ComfyUI, Diffusers, or vLLM-Omni backends due to their specialized multi-latent decoders and CUDA attention kernels.
Tencent's HunyuanVideo is currently one of the leading open video foundation models, featuring a massive dense transformer architecture. In purely visual tasks, HunyuanVideo delivers exceptional spatial consistency, prompt comprehension, and complex physics handling.
However, HunyuanVideo is strictly a silent video generator. Deploying it requires orchestrating a secondary audio generation model (such as MMAudio or AudioLDM2) alongside a tertiary lip-sync pipeline if characters speak. Kandinsky 6.0 Video Pro eliminates this multi-model synchronization pipeline entirely. While HunyuanVideo holds a slight edge in complex multi-subject visual scenes, Kandinsky 6.0 delivers higher production utility for creators who need turnkey audiovisual generation without manual sound design.
THUDM's CogVideoX-5B operates in a similar weight class to Kandinsky 6.0 Video Lite (3B). CogVideoX-5B uses an expert transformer with 3D VAE compression that runs comfortably inside consumer VRAM limits.
Kandinsky 6.0 Video Lite offers two concrete operational advantages over CogVideoX-5B:

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every Sber AI model we track.

Explore the Family
The full Kandinsky family leaderboard with sizes, benchmark scores, and a release timeline.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.