Google Gemini (audio understanding).
LLM-native audio transcription via the Gemini API, priced on tokens not minutes.
Directory noteLLM-native transcription; token-priced, not per-minute.
Google Gemini (audio understanding) speech-to-text API overview
The four signals that usually narrow a speech-to-text shortlist first. Prices are the cheapest documented async tier; native billing units differ, so read the notes below.
Google Gemini (audio understanding) pricing, languages, and features
Everything else the directory currently tracks, kept in one clear spec sheet for comparison.
- Company
- Google DeepMind
- Listed price
- See pricing notesNot a per-minute STT product. Gemini does audio understanding on an LLM, billed per token: 2.5 Flash $1.00/1M audio-input tokens, 2.5 Pro $1.25/1M.
- Languages
- Not documentedNo documented transcription language total; language coverage is not a certified STT figure.
- Real-time streaming
- Yes
- Speaker diarization
- No
- Free tier or trial
- YesFree developer tier (rate-limited); not a production free tier.
- Models
- Gemini 2.5 Flash, Gemini 2.5 Pro
Pricing, languages, streaming, and diarization verified against Google Gemini (audio understanding)’s primary sources on 2026-08-27. Word error rate and latency are not yet measured.
Official Google Gemini (audio understanding) website and API documentation
Confirm current pricing and capabilities in the provider’s own materials before making a purchasing decision.
Primary sources used to verify this profile
Alternatives to Google Gemini (audio understanding)
Switch between profiles to narrow the set that fits your product, workflow, and deployment constraints.
Building with speech?
Explore speech-to-text alongside voice and messaging APIs.
Build on Telnyx