Xing4.0-29B-A4B is a mixture-of-experts text model from China Telecom AI, the latest generation of the TeleChat series. It has 29B total parameters with 4B active per token, a 256K context window that extends to 512K, and is built for agent work such as multi-step planning, tool calling and long-context reasoning. Weights are published on Hugging Face under Apache 2.0, and with low-bit quantization it runs in about 15 GB of GPU memory on a single consumer card. It was trained on Ascend NPU hardware using the MindSpore framework.
A strong 29B-parameter MoE language model from China Telecom AI. High composite score across our benchmark mix — worth shortlisting when raw quality matters more than VRAM budget. Newly released, so production-readiness is still being shaken out.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
Access model weights, configuration files, and documentation.
No benchmark data available for this model yet.
See how different quantization levels affect VRAM requirements and quality for this model.
| Format | VRAM Required | Quality | |
|---|---|---|---|
| Q2_K | 10.4 GB | Low | |
| Q4_K_MRecommended | 11.2 GB | Good | |
| Q5_K_M | 11.6 GB | Very Good | |
| Q6_K | 12.1 GB | Excellent | |
| Q8_0 | 13.1 GB | Near Perfect | |
| FP16 | 16.9 GB | Full |
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | Speed | VRAM |
|---|---|---|---|
| AMD Radeon RX 7800 XTAMD | SS | 44.8 tok/s | 11.2 GB |
| AMD Radeon RX 7900 XTAMD | SS | 57.5 tok/s | 11.2 GB |
| AMD Radeon RX 9070AMD | SS | 46.0 tok/s | 11.2 GB |
| AMD Radeon RX 9070 XTAMD | SS | 46.0 tok/s | 11.2 GB |
| Google Cloud TPU v5eGoogle | SS | 58.8 tok/s | 11.2 GB |
Energy cost on Intel Arc B580 (~33 tok/s, Q4_K_M) vs flagship API pricing.
| Source | Cost per 1M tokens |
|---|---|
Local (energy only)Xing4.0-29B-A4B on Intel Arc B580 · ~33 tok/s · 190W | $0.193 |
GPT-6 SolOpenAI · in $2.00 · out $10.00 | $4.40 |
Claude Opus 5.5Anthropic · in $4.00 · out $20.00 | $8.80 |
Gemini 3.5 FlashGoogle · in $1.50 · out $9.00 | $3.75 |
Grok 4.5xAI · in $2.00 · out $6.00 | $3.20 |
API prices blended at 70% input / 30% output.
Hardware amortisation not included. Run the full ROI calculator for payback math.
Cheapest current cloud rentals with at least 11 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA GeForce RTX 3060Vast.ai · Spot · 12 GB VRAM | $0.05 |
NVIDIA GeForce RTX 3080 TiVast.ai · Spot · 12 GB VRAM | $0.05 |
NVIDIA GeForce RTX 3060Vast.ai · On-Demand · 12 GB VRAM | $0.05 |
NVIDIA GeForce RTX 3090Vast.ai · Spot · 24 GB VRAM | $0.08 |
NVIDIA GeForce RTX 5060 TiVast.ai · Spot · 16 GB VRAM | $0.08 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
Cumulative cost of buying the smallest device in our directory that runs this model at 4-bit, compared with paying flagship APIs for the same token volume.
Assumes 1 million tokens a day, a 70/30 input and output mix and $0.12 per kWh. The device only runs as long as the workload needs. Use the ROI calculator for your own numbers.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every China Telecom AI model we track.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.