Listed price
The cheapest documented pre-recorded (async) rate for a provider’s current flagship model, expressed per minute so the column can sort. Each profile footnotes the native billing unit and the exact model or tier the number refers to.
A practical framework for narrowing a speech-to-text API shortlist before you spend time in the documentation, and a plain account of where every figure comes from.
Every provider profile uses the same six core fields, so the first comparison is consistent from one vendor to the next.
These fields describe the integration details that usually decide a shortlist. They are not a single verdict on which provider is best, because the right choice depends on your audio, your languages, and your budget.
The cheapest documented pre-recorded (async) rate for a provider’s current flagship model, expressed per minute so the column can sort. Each profile footnotes the native billing unit and the exact model or tier the number refers to.
The supported-language count a provider states in its own documentation, with a note on whether that figure is a documented total, a marketing number, or a count of locales.
Whether the API can return transcription while audio is still arriving, over a streaming connection, rather than only after a completed file is uploaded.
Whether the service separates and labels who is speaking, not only that a speaker changed. Where a provider documents speaker-change detection but not full labeling, the profile says so.
Whether free access exists, with a note on whether it is a perpetual free allotment, a time-limited trial, or a one-time credit.
The current speech-to-text models a provider offers, named as they appear in the provider’s own pricing and documentation.
Every pricing, language, streaming, diarization, and free-tier value is checked against the provider’s own pricing page and API documentation, not third-party summaries.
Each provider profile lists the primary sources used, so any figure can be traced back to where it came from. When a provider’s live pricing is rendered in a way that resists a clean read, the profile says the number is not fully documented rather than publishing a guess.
The listed price is the cheapest documented pre-recorded rate for a provider’s current flagship model, converted to a per-minute figure so the board can sort. It is not an all-in quote.
Streaming usually costs more than batch, volume tiers cost less, and providers bill in different native units: per minute, per hour, per second, or per token. The same workload can land at a very different real cost, so each profile footnotes the native unit and the exact tier behind the number. For the vocabulary behind these fields, see the speech-to-text glossary.
Accuracy, measured as word error rate, and latency are not published yet, because a fair comparison requires measuring every provider on the same audio and the same benchmark.
Rather than repeat vendor-reported accuracy numbers drawn from different test sets, the directory leaves those fields empty until they can be measured consistently. Treat every figure here as a starting point and confirm current details in the provider’s own docs before you commit.
Providers that support real-time streaming and speaker diarization appear first. Within the same capability set, the board uses the listed per-minute price where one is available.
This is an integration-first working sort, not a claim of overall quality, accuracy, or suitability for every use case.
Start by deciding whether your product needs real-time streaming, speaker separation, broad language coverage, or a particular pricing model.
Then use the provider profiles to build a shortlist and follow the direct website and API documentation links to validate the details that matter for your workload.
Provider pricing, models, and language support change often. Each profile carries the date its fields were last verified and links to the provider’s primary sources so a reader can check for changes.
If a figure looks wrong, the provider’s own pricing or documentation page is the source of truth. Board order follows the rule-based sort described above; no provider buys placement.
Use the provider board to start your shortlist.
Explore the provider board