SGF+ (Self Gradient Forcing Plus) is an autoregressive text-to-video method from Tsinghua University, JD Joy Future Academy and the Chinese University of Hong Kong. It gives context writing and frame denoising their own parameter sets while keeping causal attention shared, because the authors found the two roles produce conflicting gradients on shared weights. Trained on 5 second rollouts, it generates one continuous video for up to 24 hours without long-video fine-tuning. Weights and code are released under Apache-2.0.
A workable dense video generator from Tsinghua University and JD Joy Future Academy. Treat the modality benchmarks above as the leading indicator of fit — composite scoring across modalities is still maturing. Newly released, so production-readiness is still being shaken out.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
No benchmark data available for this model yet.
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | VRAM |
|---|---|---|
| ACEMAGIC M1A Pro (i9-13900HK + ARC A770)ACEMAGIC | SS | 0.5 GB |
| Acer Veriton GN100 AI MiniAcer | SS | 0.5 GB |
| AMD Instinct MI300XAMD | SS | 0.5 GB |
| AMD Instinct MI325XAMD | SS | 0.5 GB |
| AMD Instinct MI355XAMD | SS | 0.5 GB |
Cheapest current cloud rentals with at least 1 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA GeForce RTX 3060Vast.ai · Spot · 12 GB VRAM | $0.05 |
NVIDIA GeForce RTX 3060Vast.ai · On-Demand · 12 GB VRAM | $0.05 |
NVIDIA GeForce RTX 4060Vast.ai · Spot · 8 GB VRAM | $0.07 |
NVIDIA GeForce RTX 4060Vast.ai · On-Demand · 8 GB VRAM | $0.08 |
NVIDIA GeForce RTX 3070Vast.ai · Spot · 8 GB VRAM | $0.08 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
SGF+ builds on Self Forcing (SF) and Self Gradient Forcing (SGF). SF blocks gradients from context writing, SGF restores them on shared parameters, and SGF+ separates the two roles into independent parameter sets that are trained jointly with the original generation objective and no auxiliary losses. Training uses a two-pass procedure: a no-gradient autoregressive rollout is recorded, then the forward computation is rebuilt from detached context and noisy target latents so future-generation losses supervise context writing through differentiable KV states. The project reports that across 128 prompts and four denoising timesteps, all 512 paired cosine similarities between context-writing and denoising gradients are negative in each module family. It is trained on the Wan2.1-T2V-1.3B base model with Wan2.1-T2V-14B as teacher; checkpoints are 11.4 GB. Evaluated baselines are SF and SGF, in framewise and chunkwise generation.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every Tsinghua University and JD Joy Future Academy model we track.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.