← Back to the directory

Google Gemini (audio understanding).

LLM-native audio transcription via the Gemini API, priced on tokens not minutes.

Directory noteLLM-native transcription; token-priced, not per-minute.

Google Gemini (audio understanding) speech-to-text API overview

The four signals that usually narrow a speech-to-text shortlist first. Prices are the cheapest documented async tier; native billing units differ, so read the notes below.

Listed priceSee pricing notesPay-as-you-go information
LanguagesApproximate coverage
StreamingYesReal-time transcription
DiarizationNoSpeaker separation

Google Gemini (audio understanding) pricing, languages, and features

Everything else the directory currently tracks, kept in one clear spec sheet for comparison.

Company
Google DeepMind
Listed price
See pricing notesNot a per-minute STT product. Gemini does audio understanding on an LLM, billed per token: 2.5 Flash $1.00/1M audio-input tokens, 2.5 Pro $1.25/1M.
Languages
Not documentedNo documented transcription language total; language coverage is not a certified STT figure.
Real-time streaming
Yes
Speaker diarization
No
Free tier or trial
YesFree developer tier (rate-limited); not a production free tier.
Models
Gemini 2.5 Flash, Gemini 2.5 Pro

Pricing, languages, streaming, and diarization verified against Google Gemini (audio understanding)’s primary sources on 2026-08-27. Word error rate and latency are not yet measured.

Official Google Gemini (audio understanding) website and API documentation

Confirm current pricing and capabilities in the provider’s own materials before making a purchasing decision.

Primary sources used to verify this profile

Alternatives to Google Gemini (audio understanding)

Switch between profiles to narrow the set that fits your product, workflow, and deployment constraints.

Building with speech?

Explore speech-to-text alongside voice and messaging APIs.

Build on Telnyx