
An encoder-only, CTC-based open speech foundation model from ESPnet/CMU that reproduces Whisper-style multilingual ASR, speech translation and language identification using fully public data. Trained on 320k hours of cleaned YODAS + prior OWSM data across 75 languages.
A solid 1B-parameter dense audio model from ESPnet. Treat the modality benchmarks above as the leading indicator of fit — composite scoring across modalities is still maturing.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
Access model weights, configuration files, and documentation.
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | VRAM |
|---|---|---|
| ACEMAGIC M1A Pro (i9-13900HK + ARC A770)ACEMAGIC | SS | 1.1 GB |
| Acer Veriton GN100 AI MiniAcer | SS | 1.1 GB |
| AMD Instinct MI300XAMD | SS | 1.1 GB |
| AMD Instinct MI325XAMD | SS | 1.1 GB |
| AMD Instinct MI355XAMD | SS | 1.1 GB |

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every ESPnet model we track.
OWSM (Open Whisper-style Speech Model) is a community effort led by CMU's WAVLab and the ESPnet team to build fully reproducible, openly trained alternatives to OpenAI Whisper. The v4 CTC variant is an encoder-only model using hierarchical multi-task self-conditioned CTC with an E-Branchformer encoder.
Reproducible research, multilingual transcription/translation across 75 languages, forced alignment (CTC segmentation), and as a base for further fine-tuning where training transparency is required.
Cheapest current cloud rentals with at least 1 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA GeForce RTX 4060Vast.ai · Spot · 8 GB VRAM | $0.03 |
NVIDIA GeForce GTX 1660 TiVast.ai · Spot · 6 GB VRAM | $0.03 |
NVIDIA GeForce GTX 1660 TiVast.ai · On-Demand · 6 GB VRAM | $0.03 |
NVIDIA Tesla V100 16GBVast.ai · Spot · 16 GB VRAM | $0.04 |
NVIDIA GeForce RTX 3060Vast.ai · Spot · 12 GB VRAM | $0.04 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.