Turn detection
Turn detection is deciding when a caller has finished speaking and it is the agent’s turn to reply.
Humans do this with grammar, intonation and context, and we are fast at it — the usual gap between two speakers in a conversation is around 200 milliseconds. An agent that waits noticeably longer feels slow, and one that jumps in early talks over people.
The naive approach is a silence timer: wait some fixed number of milliseconds of quiet and assume the turn is over. It fails on anyone who pauses mid-thought, which is most people when reading out a postcode, an order number or a date. Better approaches score whether the utterance actually sounds finished, using the audio and the words together, so "my number is four, four..." is treated as unfinished and "my number is four four seven" is treated as complete.
Getting it wrong is expensive in both directions. Cut too early and you truncate the caller mid-sentence and act on half an answer; wait too long and every exchange gains dead air that makes the agent feel unresponsive.
Related
- Endpointing — Endpointing is detecting the start and end boundaries of speech in an audio stream.
- Barge-in — Barge-in is when a caller interrupts an automated voice while it is still speaking, and the system stops talking and starts listening.
- Latency budget — A latency budget is the total time allowed between a caller finishing speaking and hearing a reply, divided across every step in between.