Whisper (open source).
Self-hosted open-weight model; no per-minute fee, you run the compute.
Directory noteThe 'open source speech to text' facet target.
Whisper (open source) speech-to-text API overview
The four signals that usually narrow a speech-to-text shortlist first. Prices are the cheapest documented async tier; native billing units differ, so read the notes below.
Whisper (open source) pricing, languages, and features
Everything else the directory currently tracks, kept in one clear spec sheet for comparison.
- Company
- OpenAI (open weights)
- Listed price
- Free / self-hostedOpen source (MIT), self-hosted. No vendor fee; you pay only for the compute you run it on.
- Languages
- 9999 is the tokenizer's language count (98 non-English plus English); large-v3-turbo covers somewhat fewer.
- Real-time streaming
- No
- Speaker diarization
- No
- Free tier or trial
- YesFree and open source; no vendor tier.
- Models
- large-v3, large-v3-turbo, medium, base
Pricing, languages, streaming, and diarization verified against Whisper (open source)’s primary sources on 2026-08-27. Word error rate and latency are not yet measured.
Official Whisper (open source) website and API documentation
Confirm current pricing and capabilities in the provider’s own materials before making a purchasing decision.
Primary sources used to verify this profile
Alternatives to Whisper (open source)
Switch between profiles to narrow the set that fits your product, workflow, and deployment constraints.
Building with speech?
Explore speech-to-text alongside voice and messaging APIs.
Build on Telnyx