Qwen3-TTS-12Hz-1.7B-Base is an open source speech generation model from Alibaba's Qwen team. It is the base model of the Qwen3-TTS family and can clone a voice from about 3 seconds of user audio, and it can be fine-tuned into other models. It covers 10 languages and supports streaming generation with end to end synthesis latency as low as 97ms. It is released under Apache 2.0.
A workable 1.7B-parameter dense audio model from Alibaba Cloud. Treat the modality benchmarks above as the leading indicator of fit — composite scoring across modalities is still maturing.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
Access model weights, configuration files, and documentation.
No benchmark data available for this model yet.
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | VRAM |
|---|---|---|
| ACEMAGIC M1A Pro (i9-13900HK + ARC A770)ACEMAGIC | SS | 1.5 GB |
| Acer Veriton GN100 AI MiniAcer | SS | 1.5 GB |
| AMD Instinct MI300XAMD | SS | 1.5 GB |
| AMD Instinct MI325XAMD | SS | 1.5 GB |
| AMD Instinct MI355XAMD | SS | 1.5 GB |
Cheapest current cloud rentals with at least 2 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA GeForce RTX 3060Vast.ai · Spot · 12 GB VRAM | $0.04 |
NVIDIA GeForce RTX 2080 TiVast.ai · Spot · 11 GB VRAM | $0.04 |
NVIDIA GeForce RTX 4060Vast.ai · Spot · 8 GB VRAM | $0.04 |
NVIDIA GeForce RTX 3060Vast.ai · On-Demand · 12 GB VRAM | $0.05 |
NVIDIA GeForce RTX 5060Vast.ai · Spot · 8 GB VRAM | $0.05 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
Qwen3-TTS-12Hz-1.7B-Base is Alibaba Cloud's open-weight text-to-speech base model, released under Apache 2.0. It is the foundation of the Qwen3-TTS family: a dense 1.7B-parameter model that takes text in and produces speech, clones a voice from roughly 3 seconds of reference audio, and serves as the starting point for fine-tuning the instruction-controlled variants in the same line. If you need a TTS model you can run on your own hardware, modify, and ship commercially without a license negotiation, this is one of the few credible options at this size.
The "12Hz" in the name is the important part. Qwen3-TTS uses a self-developed tokenizer that compresses speech into discrete codes at 12 frames per second, an unusually low frame rate for an audio codec. That compression is what lets a 1.7B dense model handle full speech synthesis end to end, without the separate diffusion decoder that most LM+DiT TTS pipelines bolt on. Fewer frames per second of audio means fewer autoregressive steps per second of output, which is why the vendor quotes end-to-end synthesis latency as low as 97ms for the first audio packet.
The model covers 10 languages: Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian, plus dialectal voice profiles. It supports both streaming and non-streaming generation from the same weights. That combination of small size, permissive license, and streaming support is what makes it interesting for production voice agents rather than just demos.
The architecture is a discrete multi-codebook language model. Text and audio are both represented as token sequences, and the model predicts audio codes directly. There is no cascaded pipeline where a text model hands off to a separate acoustic model, so there is no information bottleneck between stages and no error accumulation across them.
Key components:
speech_tokenizer/) and is required for voice cloning and for any audio output path.Context length is not published in the model card or config. Treat long-form input as something to chunk yourself rather than something to rely on the model to handle. The released weights are roughly 3.86 GB in model.safetensors (bf16), with the full repo around 4.5 GB including the tokenizer and vocabulary files.
This is the base model, not the instruction-controlled one. That distinction matters when you are picking a checkpoint. Qwen3-TTS-12Hz-1.7B-VoiceDesign and Qwen3-TTS-12Hz-1.7B-CustomVoice accept natural-language instructions for timbre, emotion, and prosody. The Base model does not. What it does instead is clone a voice from about 3 seconds of user audio and provide clean weights for fine-tuning.
Concrete work it fits:
Where it is not the right tool: expressive, instruction-driven performance direction (use VoiceDesign), or a curated set of ready-made premium voices (use CustomVoice).
The reference path is the qwen-tts Python package or vLLM, both of which download weights by model name on first load. If your environment cannot fetch weights at runtime, pre-download the repo and point the loader at the local directory. Note that Ollama, the usual fastest on-ramp for local LLMs, is a text-generation runtime; check its model library for TTS support before assuming a one-line ollama run works here. For most people, a Python environment with qwen-tts plus a CUDA GPU is the shortest path.
VRAM requirements (estimates, since the card does not publish them):
Hardware that realistically runs it:
Quantization guidance. The standard advice to reach for Q4_K_M comes from LLM practice and does not transfer cleanly. TTS quality is judged on prosody, timbre fidelity, and artifact absence, and those are the first things aggressive quantization damages. Start at bf16 or int8. Drop to int4 only if VRAM forces it, and A/B the output against the unquantized model on your actual reference voices before shipping. A quantized TTS model that is intelligible but subtly wrong in prosody is worse than useless for a voice product.
Performance. Tokens per second is the wrong metric for a TTS model. Measure two things instead: real-time factor (seconds of audio generated per second of wall clock) and first-packet latency. On a modern discrete GPU, expect RTF at or below 1.0 in bf16, meaning faster than real time for batch synthesis, and first-packet latency in the low hundreds of milliseconds or better depending on hardware and streaming configuration. The 97ms figure is the vendor's best case, not a floor you should assume.
Qwen3-TTS-12Hz-0.6B-Base is the obvious sibling. Same tokenizer, same 3-second cloning, same 10 languages, roughly a third of the parameters. The 0.6B checkpoint uses less VRAM and runs on weaker hardware, at some cost in fidelity and robustness. If you are deploying many concurrent streams on one GPU, the 0.6B is the better density play. If you are cloning a voice and shipping the output to customers, the 1.7B is the safer default.
Chatterbox and Orpheus-3B are the closest external points of comparison in open TTS. Orpheus-3B is larger and heavier to run, with expressive speech as its selling point. Chatterbox targets zero-shot cloning with emotion control. Both are credible, and both carry different license terms than Apache 2.0, which is often the deciding factor for commercial deployment. Qwen3-TTS-12Hz-1.7B-Base wins on license clarity, streaming latency, and the explicit fine-tuning story. It loses on instruction-driven expressiveness, which lives in the VoiceDesign and CustomVoice checkpoints rather than here.
Choose this model when you want a permissive license, a small enough footprint to run on a single consumer GPU, streaming latency that supports live conversation, and weights you intend to fine-tune. Choose something else when you need out-of-the-box emotional direction or a curated voice library without training.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every Alibaba model we track.

Explore the Family
The full Qwen family leaderboard with sizes, benchmark scores, and a release timeline.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.