Kandinsky 6.0 Video Pro is a 29B-parameter open-source diffusion model from Kandinsky Lab, released 2026-10-06. It generates 5-second clips at 24 fps with synchronized 44 kHz audio, including lip-sync, in text-to-audio-video and image-to-audio-video modes. A plug-in super-resolution model raises output to Full HD (1920x1080). Weights are published on Hugging Face as kandinskylab/Kandinsky-6.0-Pro-5s-Diffusers and code is on GitHub.
A situational 29B-parameter dense video generator from Kandinsky Lab. Treat the modality benchmarks above as the leading indicator of fit — composite scoring across modalities is still maturing. Newly released, so production-readiness is still being shaken out.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
Access model weights, configuration files, and documentation.
No benchmark data available for this model yet.
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | VRAM |
|---|---|---|
| Acer Veriton GN100 AI MiniAcer | SS | 18.3 GB |
| AMD Instinct MI300XAMD | SS | 18.3 GB |
| AMD Instinct MI325XAMD | SS | 18.3 GB |
| AMD Instinct MI355XAMD | SS | 18.3 GB |
| Apple M3 Ultra (32-core CPU, 80-core GPU)Apple | SS | 18.3 GB |
Cheapest current cloud rentals with at least 18 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA GeForce RTX 3090Vast.ai · Spot · 24 GB VRAM | $0.08 |
NVIDIA GeForce RTX 3090Vast.ai · On-Demand · 24 GB VRAM | $0.13 |
NVIDIA RTX A5000RunPod · Community · 24 GB VRAM | $0.16 |
NVIDIA RTX A5000RunPod · Spot · 24 GB VRAM | $0.16 |
NVIDIA RTX PRO 4000 BlackwellVast.ai · Spot · 24 GB VRAM | $0.17 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
Kandinsky 6.0 Video Pro is the high-capacity member of the Kandinsky 6.0 Video family, a set of foundation diffusion models for synchronized text-to-audio-video and image-to-audio-video generation. The Pro line has 29B parameters and delivers 5-second clips at 24 fps with 44 kHz synchronized audio and lip-sync; a plug-in super-resolution model upscales output to Full HD 1920x1080. SFT plus RL checkpoints, distilled checkpoints (pi-Flow trajectory distillation to 10 NFE with Sim-LADD refinement) and pretrain checkpoints are offered. A smaller 3B Lite model is in the same family.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every Kandinsky Lab model we track.

Explore the Family
The full Kandinsky family leaderboard with sizes, benchmark scores, and a release timeline.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.