Speech-to-text glossary

The vocabulary of speech-to-text, in plain English. Speech-to-text APIs turn spoken audio into written text, and choosing between them means reading a lot of jargon: accuracy metrics like word error rate, real-time features like streaming and endpointing, audio fundamentals like sample rate and codecs, and model architectures from acoustic models to transformers. Every term below links to the ones next to it, so you can follow a concept as far as you need.

Core concepts

The foundational terms: what speech-to-text is and the words used for it.

Accuracy & evaluation metrics

How speech-to-text accuracy is measured, scored, and compared.

Latency & streaming behavior

Speed, streaming, and how a system decides you have finished speaking.

Audio & signal processing

The audio fundamentals underneath every transcript.

Model architecture & ML

The models and machine-learning ideas that turn audio into words.

Speaker modeling & diarization

Who spoke, and telling voices apart.

Features & post-processing

What providers layer on top of a raw transcript.

Use cases & domains

Where speech-to-text gets deployed.

Linguistics & phonetics

The language and sound concepts recognition rests on.

Telephony & voice-agent adjacent

The telephony and real-time voice terms that surround speech-to-text.