Prism is a joint video-audio generation model from Tencent Hunyuan and Fudan University, released under the MIT license. It generates video with synchronized audio natively at 720p, 1080p and 2K using a dynamic sparse attention framework that the authors report gives a 2.5x training speedup over full attention. Weights, training code and inference code are published on Hugging Face and GitHub.
A workable dense video generator from Tencent Hunyuan. Treat the modality benchmarks above as the leading indicator of fit — composite scoring across modalities is still maturing. Newly released, so production-readiness is still being shaken out.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
Access model weights, configuration files, and documentation.
No benchmark data available for this model yet.
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | VRAM |
|---|---|---|
| ACEMAGIC M1A Pro (i9-13900HK + ARC A770)ACEMAGIC | SS | 0.5 GB |
| Acer Veriton GN100 AI MiniAcer | SS | 0.5 GB |
| AMD Instinct MI300XAMD | SS | 0.5 GB |
| AMD Instinct MI325XAMD | SS | 0.5 GB |
| AMD Instinct MI355XAMD | SS | 0.5 GB |
Cheapest current cloud rentals with at least 1 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA Tesla V100 16GBVast.ai · Spot · 16 GB VRAM | $0.04 |
NVIDIA GeForce RTX 3060Vast.ai · Spot · 12 GB VRAM | $0.05 |
NVIDIA GeForce RTX 5060 TiVast.ai · Spot · 16 GB VRAM | $0.05 |
NVIDIA GeForce RTX 3060Vast.ai · On-Demand · 12 GB VRAM | $0.06 |
NVIDIA GeForce RTX 3090Vast.ai · Spot · 24 GB VRAM | $0.07 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
Prism is an open-source joint video-audio generation model developed by Tencent Hunyuan in collaboration with Fudan University and Zhejiang University. Released under the permissive MIT license, the model natively generates video with temporally synchronized audio at 720p, 1080p, and 2K resolutions. Unlike decoupled pipelines that generate silent video frames before running a separate audio synthesis stage, Prism treats video and audio as an interdependent generation task. It addresses the steep computational barrier of high-resolution video modeling through a dynamic sparse attention mechanism that cuts attention overhead while preserving visual detail and sound synchronization.
The underlying weights and architecture are organized around a dense backbone with undisclosed parameters. Rather than relying on proprietary cloud APIs, Tencent Hunyuan published the training code, inference scripts, and model weights to GitHub and Hugging Face. This release allows researchers and engineers to inspect the implementation, evaluate generation quality, and experiment with local deployments. In the broader ecosystem, Prism targets high-fidelity multimodal creation, challenging existing high-resolution generative models such as Wan2.1, LTX-Video, and proprietary tools like Kling.
For engineers seeking a local AI model undisclosed parameters 2026 deployment, Prism offers an accessible entry point to native 2K multimodal generation. The authors report that the dynamic sparse attention framework achieves a 2.5x training speedup over conventional full dense attention. This computational efficiency is critical for practitioners running local workloads, as native high-resolution diffusion models typically exhaust standard memory architectures during both training and fine-tuning.
Prism builds upon a dense diffusion transformer foundation with undisclosed parameters. In native high-resolution video generation, standard dense self-attention scales quadratically with token sequence length ($O(N^2)$). When scaling from 720p to 2K resolution, the explosion of spatial and temporal tokens causes full attention to assign disproportionate computational resources to redundant background patches. This dilutes the gradient updates intended for dynamic visual elements and audio-producing regions.
To resolve this bottleneck, Prism implements dynamic sparse attention structured around spatiotemporal macro-zones:
Because Prism is released with a dense architecture and undisclosed parameter count, context length is not specified in fixed token limits like standard text models. Instead, context is bounded by video duration, frame rate, audio sample rate, and target output resolution.
Prism is engineered for coordinated multimedia synthesis, accepting textual prompts to generate coherent audiovisual outputs. The framework supports native resolutions ranging from 720p up to 2K (2048x1080) without relying on secondary spatial upscalers.
Traditional video pipelines generate silent clips, requiring secondary automated dialogue replacement (ADR) or Foley audio generation models. Prism synthesizes audio and video synchronously. Key practical use cases include:
Because the model can render natively at 1080p and 2K, developers can generate production-grade video assets without the severe blur, boundary artifacts, and motion smearing commonly introduced by cascaded super-resolution passes.
To run Prism locally, practitioners must account for the computational overhead associated with joint high-resolution diffusion transformers. The model repository provides standard inference and Fully Sharded Data Parallel (FSDP) execution scripts (prism_infer.sh and prism_infer_fsdp.sh).
Because the model checkpoint features an undisclosed parameter count and handles multi-frame 2K token maps, memory consumption is dominated by latent activation caches and diffusion sampling passes.
The best GPU for Prism depends heavily on your target output resolution. For 720p generation, a consumer RTX 4090 delivers solid throughput. For native 2K workloads, enterprise hardware with 80 GB VRAM avoids out-of-memory errors caused by large spatiotemporal token sequences.
Selecting the best quantization for Prism is essential when balancing memory boundaries with output quality:
When evaluating Prism performance, diffusion generation speed is measured in diffusion steps per second and generated video frames per second rather than standard Prism tokens per second. On an RTX 4090 at 720p, generation speeds typically settle between 1.5 to 4 seconds per diffusion step depending on sequence length and offloading configurations.
Practitioners wondering how to run undisclosed model on consumer GPU setups should clone the official GitHub repository (Tencent-Hunyuan/Prism) and configure a dedicated PyTorch environment. While text generation frontends like Ollama serve as the quickest way to get started with local language models, multimodal diffusion frameworks like Prism require Python-based diffusion runtimes (such as Hugging Face Diffusers or native PyTorch execution with Triton and CUDA kernels). Ensure FlashAttention or Triton-compatible backends are installed to leverage the dynamic sparse attention kernels.
Evaluating Prism vs competing open-source models highlights clear tradeoffs in architecture, resource consumption, and generation targets:
Prism is the right choice for engineers and researchers who require native high-resolution output with direct audio synchronization and have the workstation or enterprise GPU hardware capable of handling high-resolution diffusion passes.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every Tencent Hunyuan model we track.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.