AssemblyAI.
Speech AI models plus audio-intelligence features on top of transcription.
Directory noteSTT-first vendor; broad audio-intelligence layer.
AssemblyAI speech-to-text API overview
The four signals that usually narrow a speech-to-text shortlist first. Prices are the cheapest documented async tier; native billing units differ, so read the notes below.
AssemblyAI pricing, languages, and features
Everything else the directory currently tracks, kept in one clear spec sheet for comparison.
- Company
- AssemblyAI
- Listed price
- $0.0035 / minBilled per hour. Universal-3.5 Pro async $0.21/hr (~$0.0035/min); Universal-2 async $0.15/hr (~$0.0025/min). Streaming is billed on WebSocket session duration, not audio length.
- Languages
- 9999 applies to Universal-2; the Universal-3.5 Pro flagship supports 18 languages with code-switching.
- Real-time streaming
- Yes
- Speaker diarization
- Yes
- Free tier or trial
- Yes$50 signup credit (marketed as up to 185h prerecorded / 333h streaming). Credit-based, not a perpetual free tier.
- Models
- Universal-3.5 Pro, Universal-2, Universal-Streaming
Pricing, languages, streaming, and diarization verified against AssemblyAI’s primary sources on 2026-08-27. Word error rate and latency are not yet measured.
Official AssemblyAI website and API documentation
Confirm current pricing and capabilities in the provider’s own materials before making a purchasing decision.
Primary sources used to verify this profile
Alternatives to AssemblyAI
Switch between profiles to narrow the set that fits your product, workflow, and deployment constraints.
Building with speech?
Explore speech-to-text alongside voice and messaging APIs.
Build on Telnyx