Wan2.2-S2V-14B is an open-source speech-to-video model from Alibaba's Wan team, released on 2025-08-27. It takes a single portrait image plus an audio clip and generates a speaking, singing or performing avatar, with portrait, bust and full-body framing options. Output resolutions are 480p and 720p, and the model uses the Wan2.2 mixture-of-experts video diffusion architecture at 14B parameters. It is licensed under Apache 2.0 and weights are published on Hugging Face.
A workable 14B-parameter MoE video generator from Alibaba. Treat the modality benchmarks above as the leading indicator of fit — composite scoring across modalities is still maturing.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
Access model weights, configuration files, and documentation.
No benchmark data available for this model yet.
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | VRAM |
|---|---|---|
| ACEMAGIC M1A Pro (i9-13900HK + ARC A770)ACEMAGIC | SS | 9.1 GB |
| Acer Veriton GN100 AI MiniAcer | SS | 9.1 GB |
| AMD Instinct MI300XAMD | SS | 9.1 GB |
| AMD Instinct MI325XAMD | SS | 9.1 GB |
| AMD Instinct MI355XAMD | SS | 9.1 GB |
Cheapest current cloud rentals with at least 9 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA RTX A5000RunPod · Community · 24 GB VRAM | $0.16 |
NVIDIA RTX A5000RunPod · Spot · 24 GB VRAM | $0.16 |
NVIDIA GeForce RTX 3080RunPod · Community · 10 GB VRAM | $0.17 |
NVIDIA GeForce RTX 3080RunPod · Spot · 10 GB VRAM | $0.17 |
NVIDIA RTX A4000RunPod · Community · 16 GB VRAM | $0.17 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
Wan2.2-S2V-14B is Alibaba's open-weight speech-to-video model, released August 27, 2025 by the Wan team. Give it a single portrait image and an audio clip, and it generates a talking, singing, or performing avatar with synchronized lip movement, natural facial expression, and body motion. It sits in the audio-driven character animation category, competing directly with closed systems like Hunyuan-Avatar and OmniHuman, and it is one of the few models in that category shipping under a permissive license with published weights.
The "14B" in the name is the full parameter count. The architecture is Wan2.2's mixture-of-experts video diffusion stack, which splits the denoising process across specialized expert sub-networks per timestep. That design raises total capacity without paying the full compute cost of a dense 14B model at every step. For anyone evaluating local hardware, the practical consequence is that active parameters per forward pass are lower than the headline number, but the full weight set still has to be resident in memory or streamed from disk.
Output is 480p or 720p, with portrait, bust, and full-body framing options. The model handles minute-level generation in a single pass and supports text prompts for scene and camera control on top of the audio-driven performance. Weights are on Hugging Face under Apache 2.0, which means commercial use without a licensing negotiation.
Wan2.2-S2V-14B is a diffusion transformer, not an autoregressive language model. It does not generate tokens and it does not have a context window in the LLM sense. The relevant inputs are a reference image, an audio track, and an optional text prompt; the output is a video latent decoded to frames. If you are coming from the LLM side of local inference, reset your mental model: throughput is measured in frames per second and seconds of video per minute of wall clock, not tokens per second.
The MoE design is the part that matters for hardware planning. Wan2.2 routes the denoising process across timestep-specialized experts, so a given sampling step activates a subset of the total 14B weights. This keeps per-step FLOPs closer to a smaller dense model while retaining the capacity of the larger one. In practice, you still need to hold or page in all 14B parameters, plus the VAE and audio encoder, so VRAM is governed by total weight size rather than active parameter count. The efficiency gain shows up as faster sampling, not as a smaller memory footprint.
The base Wan2.2 stack also brought cinematic-level aesthetic conditioning (lighting, composition, contrast, color tone) and a substantially larger training corpus than Wan2.1: roughly 65% more images and 83% more videos. That data expansion is what drives the improved motion and semantics generalization the Wan team reports.
This model does one job and does it at production quality: driving a still portrait from an audio track.
Where it is weak: it is not a general video model. It will not generate arbitrary scenes from text alone, and it is not a chat or reasoning model. Treat it as a specialized avatar renderer.
This is a large diffusion model, and the VRAM math is unforgiving. Plan around weight precision first.
There is no Ollama path here. Ollama targets LLMs; Wan2.2-S2V runs through the official Wan2.2 repository, the diffusers pipeline, or ComfyUI with the Wan S2V nodes. ComfyUI is the fastest way to get a working local setup, and the Hugging Face repo ships the requirements_s2v.txt needed for the reference implementation. Install the S2V-specific requirements, pull the weights, and drive it from the provided generate.py or a ComfyUI graph.
Budget your storage too: the full weight set plus VAE and encoders runs into tens of gigabytes, and you will want both an FP8 and a quantized copy on disk if you are testing quality tradeoffs.
The closest open-weight alternative is Hunyuan-Avatar from Tencent. Both target audio-driven human animation, and the Wan-S2V paper benchmarks directly against it, reporting stronger expressiveness and fidelity in cinematic contexts. Hunyuan-Avatar has its own ecosystem and tooling; pick it if you are already invested in Tencent's stack. Pick Wan2.2-S2V if you want the stronger reported benchmark results, Apache 2.0 licensing, and tighter integration with the broader Wan2.2 video toolchain (the same repo covers text-to-video and image-to-video, so one install serves multiple pipelines).
OmniHuman is the other direct comparison, and it is the one most practitioners will have seen in demos. It is not open-weight, so the comparison is really about whether you need local execution at all. If you need to run offline, keep audio and likeness data on-premises, or avoid per-clip API costs, Wan2.2-S2V is the answer among these three. If you only need occasional clips and want zero setup, a hosted service will beat a local install on time-to-first-output.
The deciding factor for most local deployments is VRAM. If you have a 24GB card, FP8 or a 5-bit quant gets you to 720p. If you have 16GB or less, quantize aggressively and accept 480p, or look at smaller avatar models.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every Alibaba model we track.

Explore the Family
The full Wan family leaderboard with sizes, benchmark scores, and a release timeline.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.