Advertising disclosure: we earn commissions when you shop through the links below.
T-Head, Alibaba's chip subsidiary, announced the Zhenwu V900 AI accelerator for training and inference at the 2026 Apsara Conference. It carries 216GB of memory, 1,200GB/s of inter-chip bandwidth and native FP8 and FP4 support. T-Head says more than 1,000 chips can operate as a single system and a cluster can scale to 500,000 accelerators. Mass production and commercial release are planned for Q1 2027, and no FLOPS, process node, foundry or power figures were disclosed.
Sized for production serving of 70B–200B class models at full or lightly-quantized precision. Overkill for a homelab; right call when the workload pays for itself in token volume.
Generated from this product’s spec sheet. Editor reviews refine it over time.
The Zhenwu V900 is a data center AI accelerator from T-Head, Alibaba’s chip subsidiary, announced at the 2026 Apsara Conference. It is not a consumer or prosumer part. It is a training and inference processor designed to sit in racks, wired into supernodes of more than 1,000 chips, and scaled to clusters of up to 500,000 accelerators. T-Head positions it as the most powerful AI chip in China, though the company has not disclosed FLOPS, process node, foundry, or power figures. Mass production and commercial release are planned for Q1 2027.
For anyone evaluating Zhenwu V900 for AI workloads, the headline numbers are 216GB of memory per chip, 1,200 GB/s of inter-chip bandwidth, and native FP8 and FP4 support. Those specs put it in the same tier as NVIDIA’s H200 and AMD’s MI300X, but with a crucial difference: T-Head has revealed enough to confirm memory capacity and scale-up ambitions, yet withheld the compute and memory bandwidth numbers that would let you estimate Zhenwu V900 AI inference performance or Zhenwu V900 tokens per second. That makes it a chip you track if you are building frontier models in China, but not one you can benchmark against alternatives today.
The V900 steps up from the Zhenwu M890, which had 144GB of memory and 800GB/s of inter-chip bandwidth. T-Head claims the V900 delivers three times the performance of its predecessor, but without a FLOPS figure, that claim is unverifiable. What is clear is the target: training and serving Qwen models in the 5T to 10T parameter range, a scale that requires the supernode architecture the V900 is built around.
The specs that matter for AI inference are memory capacity, memory bandwidth, compute throughput, and low-precision support. T-Head has disclosed some of these, and not others.
What is missing: FLOPS, process node, foundry, power consumption, and memory bandwidth (HBM). Without memory bandwidth, you cannot estimate token generation speed. Without FLOPS, you cannot estimate training throughput. Without power, you cannot calculate efficiency. T-Head has given you capacity and scale, but not speed.
That said, a 216GB GPU for AI is a serious amount of memory. For comparison, NVIDIA’s H200 has 141GB, and AMD’s MI300X has 192GB. The V900’s 216GB gives it the largest per-chip memory capacity among announced accelerators. The 1,200 GB/s inter-chip bandwidth is lower than the 900 GB/s per-GPU NVLink on NVIDIA’s H100 but is designed to scale to far more chips in a single supernode. Whether that trade-off pays off depends on workloads that are memory-capacity-bound rather than bandwidth-bound.
A single V900 with 216GB of VRAM can fit a wide range of models, but the largest models require quantization or multiple chips.
For hardware for running 10T parameter models, you need a supernode. A single V900 cannot hold a 10T parameter model at any practical quantization. But 1,000 V900s provide 216TB of aggregate memory, which is enough for a 10T parameter model at FP8 (10TB) or FP4 (5TB) with room for activations and KV cache. T-Head explicitly targets 5T to 10T parameter Qwen models with this architecture.
The sweet spot for quantization on this hardware is FP8 for models up to ~200B parameters and FP4 for larger models where memory capacity is the binding constraint. FP4 halves memory versus FP8, but quality degradation is model-dependent. For inference of 70B-class models, FP16 or FP8 will give better quality-to-speed tradeoffs than FP4.
Long-context tasks benefit from the 216GB capacity: you can allocate tens of gigabytes to KV cache without evicting model weights. Multimodal models (vision-language, audio) fit if their combined parameter count stays within the memory budget. Zhenwu V900 VRAM for large language models is its strongest asset.
The V900 is not for hobbyists, and it is not for local LLM deployment. It is a data center part, sold to cloud providers, AI labs, and large enterprises. If you are looking for the best hardware for running AI models locally or the best hardware for local AI agents 2026, this is not it. Look at consumer GPUs like the RTX 4090 or workstation cards like the RTX 6000 Ada.
Who should care about the V900?
Training vs. inference: T-Head designed the V900 for both. FP8 and FP4 support is relevant for training (lower precision reduces memory and increases throughput) and for inference (faster token generation). The 1,200 GB/s inter-chip bandwidth is more important for training, where gradients must be synchronized across chips.
The most direct comparison is Zhenwu V900 vs NVIDIA H200. The H200 has 141GB of HBM3e and 4.8TB/s of memory bandwidth. The V900 has 216GB but undisclosed memory bandwidth. The H200 has known FLOPS (FP8: 1,979 TFLOPS dense) and a mature software stack (CUDA). The V900 has neither disclosed. For a team outside China, the H200 is the safer choice because you can benchmark it and deploy it today. The V900 is announced, not shipping.
Zhenwu V900 vs Huawei Ascend 910B: Huawei’s chip is the other major domestic Chinese AI accelerator. The Ascend 910B offers 64GB of HBM per chip, far less than the V900’s 216GB. Huawei has its own supernode architecture (CloudMatrix) and a growing software ecosystem (CANN). If you are building in China and need maximum memory per chip, the V900 looks stronger on paper. But Huawei has shipped in volume, while the V900 is not yet in production.
Zhenwu V900 vs AMD MI300X: The MI300X has 192GB of HBM3 and 5.3TB/s of memory bandwidth. It runs ROCm, which is less mature than CUDA but improving. The V900 has more memory but no disclosed bandwidth or compute. For inference workloads that are memory-capacity-bound, the V900 could win. For anything else, the MI300X is a known quantity.
When would you pick the V900? If you are in China, need the largest possible per-chip memory, and are willing to bet on T-Head’s software stack and supernode architecture. If you need proven performance, available hardware, and a mature ecosystem, you pick NVIDIA or AMD. The V900 is a statement of intent from Alibaba. The specs that matter for best AI chip for local deployment are not here—this is a cluster chip, not a desktop one.
The top models this device can run at 4-bit, ranked by fit and speed.
Specs not available for scoring. This product is missing VRAM or memory bandwidth data.


Also in Our Store
Docks, fast storage, cables, and a UPS sized for a local AI machine.

Rent this class of GPU by the hour before you buy. Live prices from RunPod and Vast.ai.

Find the break-even between buying this hardware and paying for a cloud API.

Get a full budget-matched parts list for a local AI workstation.
Mac vs NVIDIA for local inference, if you are still choosing a platform.