CogVideoX1.5-5B is an open-source text-to-video model from Zhipu AI, built with THUDM. It generates 5 or 10 second clips at 1360x768 resolution and 16 frames per second, with a prompt limit of 224 tokens. Weights are published on Hugging Face under an 'other' license and run through the diffusers library, needing about 10GB of VRAM in BF16. It is the upgraded version of the earlier CogVideoX-5B model.
A workable 5B-parameter dense video generator from 智谱AI. Treat the modality benchmarks above as the leading indicator of fit — composite scoring across modalities is still maturing.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
Access model weights, configuration files, and documentation.
No benchmark data available for this model yet.
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | VRAM |
|---|---|---|
| ACEMAGIC M1A Pro (i9-13900HK + ARC A770)ACEMAGIC | SS | 3.6 GB |
| Acer Veriton GN100 AI MiniAcer | SS | 3.6 GB |
| AMD Instinct MI300XAMD | SS | 3.6 GB |
| AMD Instinct MI325XAMD | SS | 3.6 GB |
| AMD Instinct MI355XAMD | SS | 3.6 GB |
Cheapest current cloud rentals with at least 4 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA GeForce RTX 4060 TiVast.ai · Spot · 8 GB VRAM | $0.06 |
NVIDIA GeForce RTX 5060 TiVast.ai · Spot · 16 GB VRAM | $0.07 |
NVIDIA Tesla V100 16GBVast.ai · On-Demand · 16 GB VRAM | $0.08 |
NVIDIA RTX A4000Vast.ai · Spot · 16 GB VRAM | $0.08 |
NVIDIA RTX A4000Vast.ai · On-Demand · 16 GB VRAM | $0.08 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
CogVideoX1.5-5B is an open-source text-to-video generation model developed by Zhipu AI in collaboration with Tsinghua University (THUDM). Designed to push high-resolution synthetic video generation onto accessible hardware, the model builds directly on the architecture of the earlier CogVideoX-5B. It upgrades visual fidelity to 1360x768 resolution at 16 frames per second, rendering either 5-second or 10-second continuous clips from standard text prompts.
Positioned in the open-weights generative media category, CogVideoX1.5-5B directly challenges both commercial text-to-video platforms and large community checkpoints. While proprietary models remain locked behind metered cloud APIs, Zhipu AI has made the model weights accessible on Hugging Face under a custom license. This gives developers complete local control over their generation pipelines, parameter tuning, and integration into custom visual workflows.
Running the model requires substantial compute, but modern memory management frameworks keep it within reach of workstation-class hardware. With a dense 5-billion parameter footprint, CogVideoX1.5-5B stands out as a practical local AI model 5B parameters 2026 deployment option for engineers building autonomous video pipelines, visual simulation environments, and automated production stacks.
CogVideoX1.5-5B utilizes a dense diffusion transformer backbone containing 5 billion parameters. Unlike Mixture-of-Experts (MoE) designs that activate a sparse subset of weights per token, every parameter in this dense model participates in every denoising step. This ensures consistent feature propagation across temporal and spatial dimensions, which is critical for preventing artifact flickering across video frames.
The model processes prompts using an integrated text encoder with a strict prompt token limit of 224 tokens. While text-only in its conditioning, the prompt parser is tailored for detailed English descriptive prompts that define lighting, subject geometry, camera trajectory, and temporal pacing.
Key architectural characteristics include:
16N + 1 (where N is less than or equal to 10). A standard 5-second clip renders 81 frames, while a 10-second clip scales to 161 frames.CogVideoX1.5-5B excels at rendering coherent motion dynamics and consistent spatial geometry across multi-second scenes. Where earlier open video generators struggled with morphing limbs, unnatural warping, and drifting backgrounds, the updated architecture in version 1.5 maintains structural stability over extended trajectories.
Primary real-world applications include:
Deploying CogVideoX1.5-5B locally differs from serving a language model. Instead of relying on conversational runtimes like Ollama (which focuses on text and vision-language LLMs), video generation models rely on the Hugging Face diffusers library, ComfyUI nodes, or Zhipu AI's native SwissArmyTransformer (SAT) framework.
Memory consumption depends heavily on the execution framework, precision, and offloading flags.
| Execution Mode | Precision | Offloading Enabled | Minimum VRAM | Recommended VRAM |
|---|---|---|---|---|
| diffusers Pipeline | INT8 (TorchAO) | Sequential CPU Offload | 7 GB | 12 GB |
| diffusers Pipeline | BF16 | Sequential CPU Offload | 10 GB | 16 GB |
| diffusers Pipeline | BF16 | No Offloading (Peak) | 26 GB | 32 GB |
| Multi-GPU diffusers | BF16 | Distributed Across GPUs | 24 GB combined | 32 GB combined |
| Native SAT Framework | BF16 | Enterprise Serving | 76 GB | 80 GB |
When operating through diffusers, enabling memory-saving features is critical for consumer hardware:
1pipe.enable_sequential_cpu_offload()2pipe.vae.enable_slicing()3pipe.vae.enable_tiling()
Disabling these flags increases VRAM requirements roughly threefold, though it yields a 3x to 4x improvement in generation speed.
The best GPU for CogVideoX1.5-5B in a local workstation setup is an NVIDIA GeForce RTX 4090 (24 GB VRAM) or an enterprise NVIDIA RTX 6000 Ada (48 GB VRAM).
pipe.enable_sequential_cpu_offload() and use TorchAO quantization.The best quantization for CogVideoX1.5-5B on consumer rigs is INT8 or FP8 via torchao. This drops active VRAM usage down to roughly 7 GB to 8 GB while preserving video sharpness.
Because video generation runs over 50 denoising steps across dozens of frames, CogVideoX1.5-5B performance cannot be measured in simple tokens per second like an autoregressive LLM. Instead, performance is measured in total generation time per clip:
Evaluating CogVideoX1.5-5B vs competitor models reveals distinct architectural trade-offs:

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every 智谱AI model we track.

Explore the Family
The full CogVideoX family leaderboard with sizes, benchmark scores, and a release timeline.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.