GLM-TTS is a text-to-speech system from Zhipu AI that clones a speaker's voice from 3 to 10 seconds of prompt audio. It uses a two-stage design: a Llama-based language model generates speech tokens, then a flow matching model converts them into audio waveforms. A GRPO multi-reward reinforcement learning stage tunes pronunciation, speaker similarity and emotional prosody. The weights are open source under Apache 2.0, and it mainly supports Chinese with mixed English text.
A workable dense audio model from Zhipu AI. Treat the modality benchmarks above as the leading indicator of fit — composite scoring across modalities is still maturing.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
No benchmark data available for this model yet.
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | VRAM |
|---|---|---|
| ACEMAGIC M1A Pro (i9-13900HK + ARC A770)ACEMAGIC | SS | 0.5 GB |
| Acer Veriton GN100 AI MiniAcer | SS | 0.5 GB |
| AMD Instinct MI300XAMD | SS | 0.5 GB |
| AMD Instinct MI325XAMD | SS | 0.5 GB |
| AMD Instinct MI355XAMD | SS | 0.5 GB |
Cheapest current cloud rentals with at least 1 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA GeForce RTX 3060Vast.ai · Spot · 12 GB VRAM | $0.05 |
NVIDIA GeForce RTX 3060Vast.ai · On-Demand · 12 GB VRAM | $0.06 |
NVIDIA GeForce RTX 5060 TiVast.ai · Spot · 16 GB VRAM | $0.07 |
NVIDIA Tesla V100 16GBVast.ai · Spot · 16 GB VRAM | $0.09 |
NVIDIA RTX A4000Vast.ai · Spot · 16 GB VRAM | $0.09 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
GLM-TTS is an open-source, zero-shot text-to-speech synthesis system developed by Zhipu AI. Built specifically to handle expressive, highly controllable speech generation, the system reproduces a target voice from an audio prompt lasting only 3 to 10 seconds. Rather than relying solely on traditional acoustic regression, GLM-TTS couples an autoregressive language model with a flow matching diffusion framework, using reinforcement learning to close the naturalness gap that often plagues local voice synthesis.
Released under the permissive Apache-2.0 license, GLM-TTS is targeted at developers building conversational agents, voiceover pipelines, and interactive speech systems. While Zhipu AI has kept the exact parameter count undisclosed, the dense architecture combines a Llama-derived speech language model with a conditional flow matching model and a high-fidelity neural vocoder. It operates primarily on Mandarin Chinese with native handling of mixed English phrases, delivering fine-grained prosodic variance rarely seen in standard open-source TTS engines.
For engineers seeking to run GLM-TTS locally, this architecture presents distinct deployment characteristics. Unlike monolithic text models, voice generation requires managing two separate computational stages: discrete acoustic token prediction followed by continuous waveform synthesis. By incorporating post-training alignment techniques typically reserved for text generation, GLM-TTS provides production-grade voice cloning and streaming capabilities directly on local hardware.
GLM-TTS abandons end-to-end regression in favor of a modular, two-stage pipeline designed for stability and expressiveness:
A technical differentiator in GLM-TTS is its alignment stage. Standard autoregressive speech models frequently suffer from pronunciation drift, hallucinations, and flat emotional affect. Zhipu AI mitigates this by applying Group Relative Policy Optimization (GRPO) directly to the speech LLM. The multi-reward GRPO framework evaluates generated candidate audio sequences across several concurrent reward functions:
The underlying architecture is dense with undisclosed parameters. Because the context length is not explicitly capped in standard text-token terms, context capacity is bounded primarily by the duration of the audio prompt (typically 3 to 10 seconds) and the length of the target utterance. To guarantee accurate pronunciation of polyphones and domain-specific terminology, the architecture integrates a hybrid phoneme-text frontend that accepts International Phonetic Alphabet (IPA) or Pinyin alongside raw text.
GLM-TTS delivers specialized audio capabilities tailored for production interactive systems:
Running GLM-TTS locally requires provisioning enough compute for both the autoregressive token generator and the flow matching model. Because this is a local AI model undisclosed parameters 2026 release, inference memory depends on batching, prompt audio length, and precision mode.
Unlike single-binary LLMs that run directly through Ollama out of the box, GLM-TTS relies on a Python pipeline executing PyTorch, CUDA, or Ascend NPU kernels:
For practitioners wondering how to run undisclosed model on consumer GPU setups with constrained memory, weight quantization is the primary lever:
On a dedicated desktop GPU, GLM-TTS performance is measured by Real-Time Factor (RTF) alongside generation rate:
Evaluating GLM-TTS against open-source alternatives highlights significant trade-offs in architecture and voice quality:
Alibaba's CosyVoice series also uses an autoregressive model paired with flow matching for zero-shot cloning. However, GLM-TTS differentiates itself through its GRPO multi-reward reinforcement learning alignment. While CosyVoice relies primarily on supervised pre-training and supervised instruction tuning, GLM-TTS directly optimizes against Character Error Rate and emotional prosody during post-training.
CosyVoice offers broader pre-packaged language variants and established inference engines, whereas GLM-TTS achieves lower character error rates and cleaner speaker similarity in mixed Mandarin-English code-switching workloads.
F5-TTS is a non-autoregressive system based entirely on flow matching without a separate autoregressive LLM stage.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every Zhipu AI model we track.

Explore the Family
The full GLM family leaderboard with sizes, benchmark scores, and a release timeline.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.