STT (speech to text)
Speech to text turns spoken audio into written words, in real time for a live phone call.
Also called ASR, automatic speech recognition. For a phone agent it runs continuously on a stream rather than on a finished recording, emitting partial results that get revised as more audio arrives.
Telephone audio is the hard case: it is narrowband, compressed, and often carries background noise, accents and names the model has never seen. This is why proper nouns, postcodes and order numbers are where recognition most often fails, and why agents that matter usually read those back to confirm.
Related
- TTS (text to speech) — Text to speech turns written words into spoken audio, and is what gives a phone agent its voice.
- Endpointing — Endpointing is detecting the start and end boundaries of speech in an audio stream.
- Latency budget — A latency budget is the total time allowed between a caller finishing speaking and hearing a reply, divided across every step in between.