Kandinsky Lab open-sourced Kandinsky 6.0 Video Lite on October 6, 2026. It is a 3B parameter diffusion model that generates 5-second clips at 24 fps with synchronized 44 kHz audio, including lip sync, in text-to-audio-video and image-to-audio-video modes. Native output is SD resolution and a built-in super-resolution model upscales it to Full HD (1920x1080). The weights and code are open source, with checkpoints for SFT plus RL, distilled and pretrain use.
A workable 3B-parameter dense video generator from Kandinsky Lab. Treat the modality benchmarks above as the leading indicator of fit — composite scoring across modalities is still maturing. Newly released, so production-readiness is still being shaken out.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
Access model weights, configuration files, and documentation.
No benchmark data available for this model yet.
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | VRAM |
|---|---|---|
| ACEMAGIC M1A Pro (i9-13900HK + ARC A770)ACEMAGIC | SS | 2.3 GB |
| Acer Veriton GN100 AI MiniAcer | SS | 2.3 GB |
| AMD Instinct MI300XAMD | SS | 2.3 GB |
| AMD Instinct MI325XAMD | SS | 2.3 GB |
| AMD Instinct MI355XAMD | SS | 2.3 GB |
Cheapest current cloud rentals with at least 2 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA GeForce RTX 3060Vast.ai · Spot · 12 GB VRAM | $0.05 |
NVIDIA GeForce RTX 3060Vast.ai · On-Demand · 12 GB VRAM | $0.05 |
NVIDIA GeForce RTX 4060Vast.ai · Spot · 8 GB VRAM | $0.07 |
NVIDIA GeForce RTX 4060Vast.ai · On-Demand · 8 GB VRAM | $0.08 |
NVIDIA GeForce RTX 3070Vast.ai · Spot · 8 GB VRAM | $0.08 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
Kandinsky 6.0 Video Lite is the lightweight line of the Kandinsky 6.0 family of diffusion models for synchronized video and audio generation. It produces 5-second clips at 24 fps from a text prompt or a reference image, with 44 kHz synchronized audio and lip sync. SFT plus RL variants use two-stage fine-tuning with model soup across 11 domains and OmniNFT reinforcement learning for speech intelligibility and audio-video sync; distilled variants use pi-Flow trajectory distillation to 10 NFE with Sim-LADD adversarial refinement. Pretrain checkpoints are published for fine-tuning and research. A plug-in super-resolution model raises output to Full HD. The higher-capacity sibling is Kandinsky 6.0 Video Pro at 29B parameters.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every Kandinsky Lab model we track.

Explore the Family
The full Kandinsky family leaderboard with sizes, benchmark scores, and a release timeline.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.