Speech-to-text glossary
The vocabulary of speech-to-text, in plain English. Speech-to-text APIs turn spoken audio into written text, and choosing between them means reading a lot of jargon: accuracy metrics like word error rate, real-time features like streaming and endpointing, audio fundamentals like sample rate and codecs, and model architectures from acoustic models to transformers. Every term below links to the ones next to it, so you can follow a concept as far as you need.
Core concepts
The foundational terms: what speech-to-text is and the words used for it.
- Audio-to-text
- Automatic speech recognition (ASR)
- Batch transcription
- Dictation
- Natural language processing (NLP)
- Pre-recorded transcription
- Real-time transcription
- Speech analytics
- Speech recognition
- Speech-to-text (STT)
- Streaming transcription
- Text-to-speech (TTS)
- Transcription
- Voice recognition
- Voice typing
- Voice-to-text
Accuracy & evaluation metrics
How speech-to-text accuracy is measured, scored, and compared.
Latency & streaming behavior
Speed, streaming, and how a system decides you have finished speaking.
Audio & signal processing
The audio fundamentals underneath every transcript.
Model architecture & ML
The models and machine-learning ideas that turn audio into words.
Speaker modeling & diarization
Who spoke, and telling voices apart.
Features & post-processing
What providers layer on top of a raw transcript.
Use cases & domains
Where speech-to-text gets deployed.
Linguistics & phonetics
The language and sound concepts recognition rests on.
Telephony & voice-agent adjacent
The telephony and real-time voice terms that surround speech-to-text.