A compressed version of OpenAI Whisper Large v3 produced via LiteASR low-rank approximation, keeping accuracy close to the original while reducing encoder size. The 'acc' variant optimizes for accuracy over speed.
A solid 1B-parameter dense audio model from Efficient Speech. Treat the modality benchmarks above as the leading indicator of fit — composite scoring across modalities is still maturing.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
Access model weights, configuration files, and documentation.
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | VRAM |
|---|---|---|
| ACEMAGIC M1A Pro (i9-13900HK + ARC A770)ACEMAGIC | SS | 1.1 GB |
| Acer Veriton GN100 AI MiniAcer | SS | 1.1 GB |
| AMD Instinct MI300XAMD | SS | 1.1 GB |
| AMD Instinct MI325XAMD | SS | 1.1 GB |
| AMD Instinct MI355XAMD | SS | 1.1 GB |
Cheapest current cloud rentals with at least 1 GB VRAM, refreshed hourly.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA GeForce RTX 5060 TiVast.ai · Spot · 16 GB VRAM | $0.07 |
NVIDIA GeForce RTX 5060 TiVast.ai · On-Demand · 16 GB VRAM | $0.08 |
NVIDIA GeForce RTX 5070 TiVast.ai · Spot · 16 GB VRAM | $0.09 |
NVIDIA GeForce RTX 3090Vast.ai · Spot · 24 GB VRAM | $0.09 |
NVIDIA GeForce RTX 5080Vast.ai · Spot · 16 GB VRAM | $0.10 |
Per-GPU rate across RunPod and the Vast.ai marketplace.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
Lite-Whisper is a family of compressed Whisper checkpoints produced using LiteASR (Kamahori et al., arXiv:2502.20583), a post-training compression technique that replaces dense linear layers in the Whisper encoder with low-rank approximations derived from activation statistics.
transformers-acc checkpoint targets the highest accuracy in the Lite-Whisper family, matching or outperforming Whisper Large v3 while being meaningfully smaller and fasterHigh-quality multilingual ASR (99 languages) in latency- or memory-constrained deployments; a cheaper alternative to running vanilla Whisper Large v3.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every Efficient Speech model we track.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.