· 8 min read
The latency budget of a phone call
A caller notices delay long before they can name it. Here is where the time actually goes between someone finishing a sentence and an agent starting one, and which parts of that budget you actually control.
by Woise Team
People are unforgiving about delay on the phone in a way they are not anywhere else. A web page that takes a second to load is fine. A person who takes a second to answer a question has, socially, done something. The gap reads as confusion, or as not listening, and callers start talking over it to fill the silence.
So the useful way to think about a voice agent is as a budget. There is a total amount of delay a caller will accept before the conversation stops feeling like a conversation, and every stage of the pipeline spends some of it. Knowing which stages you control is most of the work.
Where the time actually goes
Between a caller finishing a sentence and an agent starting one, roughly five things happen. The audio has to reach the platform over the phone network. Something has to decide the caller has actually stopped speaking rather than paused. The speech has to be turned into text. A model has to decide what to say. And that reply has to be turned back into audio and sent down the line.
They are not equal, and they are not equally yours. Network transit is largely fixed by geography and the carrier. Turn detection is a tuning decision. The model step varies enormously by which model you picked and how much you are asking it to do in one go.
The part everyone optimises first is rarely the biggest
The instinct is to reach for a faster model, because that is the stage with the most visible knobs. It is often not where the budget is going. A model that responds quickly but is being asked to re-read a long prompt on every turn can easily be slower in practice than a heavier model with a tighter context.
Turn detection is the stage most worth looking at first, because it is pure waiting. If the agent holds on for a fixed period of silence before deciding the caller has finished, that period is added to every single turn, whether or not the caller had finished. Shorten it and the agent interrupts. Lengthen it and the agent feels slow. There is no setting that is right for every conversation, which is why it is worth setting per flow rather than accepting a default.
Perceived latency is not measured latency
Two agents with identical timings can feel very different. An agent that says nothing at all until its full reply is ready feels slower than one that starts speaking as soon as it has the first clause, even when both finish at the same moment.
The same is true of acknowledgement. A short filler while a lookup runs changes how a two-second wait feels, because the caller knows something is happening. Used carelessly it becomes a verbal tic. Used on the branches that genuinely wait on an external system, it buys you real time.
What to do about it
Measure before you tune, and measure the whole turn rather than one stage. The stage you assume is the problem usually is not, and a change that improves one stage while pushing work into another can leave the caller worse off.
Then decide what the flow actually needs. A booking agent that reads a live calendar has a genuine wait in it and should be designed around that. An agent answering a fixed question does not, and should not be carrying the latency profile of one that does.