TTS (text to speech)
Text to speech turns written words into spoken audio, and is what gives a phone agent its voice.
Quality is now less about whether it sounds human and more about whether it sounds right for the job: pace, warmth, how it handles a number or an address, and whether it can start speaking before the whole sentence has been generated.
That last property matters more than it sounds. Streaming synthesis lets audio begin while the rest of the reply is still being produced, which removes a meaningful chunk of the gap a caller would otherwise hear as silence.
Related
- STT (speech to text) — Speech to text turns spoken audio into written words, in real time for a live phone call.
- Latency budget — A latency budget is the total time allowed between a caller finishing speaking and hearing a reply, divided across every step in between.
- Barge-in — Barge-in is when a caller interrupts an automated voice while it is still speaking, and the system stops talking and starts listening.