Endpointing
Endpointing is detecting the start and end boundaries of speech in an audio stream.
It is the lower-level mechanism most turn detection is built on: given a continuous stream of audio, mark where an utterance begins and where it ends, so the parts in between can be transcribed and the silence can be ignored.
The classic implementation is voice activity detection, which asks a much simpler question — is this frame of audio speech or not — and is cheap enough to run continuously. Its weakness is that it hears sound, not meaning: it cannot tell a thoughtful pause from a finished sentence, which is why endpointing alone makes for an agent that interrupts.
Related
- Turn detection — Turn detection is deciding when a caller has finished speaking and it is the agent’s turn to reply.
- Barge-in — Barge-in is when a caller interrupts an automated voice while it is still speaking, and the system stops talking and starts listening.
- STT (speech to text) — Speech to text turns spoken audio into written words, in real time for a live phone call.