Fish Speech is Fish Audio’s open-weight multilingual text-to-speech family with low-latency streaming and voice cloning, widely self-hosted for real-time speech.
See all models from Fish AudioModels in family
2
Open weight
2
API only
0
Avg score
77.0
Top benchmark
78.2
TTS Arena
Total HF downloads
3.9K
Primary modality
Audio
First release
Sep 2024
Latest release
Dec 2024
Every release in the Fish Speech family, ranked by composite score across benchmarks, popularity, efficiency, and versatility.
| # | Model | Modality | Score | Params | Released |
|---|---|---|---|---|---|
| 1 | audio | AA78.0 | — | Sep 2024 | |
| 2 |
When each release shipped, newest first. Useful for tracking version cadence.
Dec 4
Sep 10
Composite grades across this family. Higher is better, blending benchmarks, popularity, and efficiency.
Models with downloadable weights, ranked by composite score.
| # | Model | Modality | Score | Params | Released |
|---|---|---|---|---|---|
| 1 | audio | AA78.0 | — | Sep 2024 | |
| 2 |
The Fish Speech family is a series of AI models from Fish Audio. This page lists every release in the family with its benchmark scores, parameter count, and hardware requirements.
By composite score, Fish Speech v1.4 is currently the top model in the family. For local inference, match the parameter count to your VRAM budget. For quality, pick the highest scorer that fits.
See the open-weight section above for models you can run locally. The API-only section lists closed releases that must be accessed through the provider’s API.
Spin up an instance in the cloud, or pick local hardware that fits.
Advertising disclosure: we earn commissions when you shop through the links below.
AA76.1 |
| — |
| Dec 2024 |
AA76.1 |
| — |
| Dec 2024 |
Vultr
GPU cloud with hourly and monthly plans, starting at $0.50/hr.