Skip to content
Share one workflow. DRING calls in two minutes and qualifies the need. Get a call in two minutes
Get a call in 2 minutes See Agent Factory
Voice AI operations · Latency measurement

Why your AI agent makes callers wait

The pause after a caller stops talking is the sum of six parts. Measure every turn, read the slowest ones, and fix the part that takes the most time.

The short answer

When a caller stops talking and hears silence, the pause is rarely one slow thing. It is the sum of six parts: End of turn, Transcription, Model reply, Lookup, Voice start and Line. Lookup happens only on lookup turns, where the agent checks a system, such as an order status. All other turns are plain turns.

This guide calls the pause the response gap: from the moment the caller stops speaking to the agent's first sound, as the caller hears it. An average response gap hides the turns callers remember. So measure every turn, split plain turns from lookup turns, and read the median and the 90th percentile.

Then find the largest part in each of the slowest turns. The part that is largest most often is your first fix. Change one part at a time and watch the opposite signal: an agent that waits less before it replies can start cutting callers off. Where a wait cannot be avoided, a spoken update gives the caller a first sound sooner. It does not make the answer come sooner.

Do not copy a target number from anywhere. Find the gap length at which your own callers start saying "Hello?" into the silence. That is your threshold.

Our article on where callers hang up on AI calls finds the stage where callers leave. This guide goes one level down: why the silence is long, and which part to fix. Every number below is an example, not DRING data and not industry data.

The six parts of a response gap

Between the caller's last word and the agent's first sound, the time splits into six parts. This guide always lists them in this order:

PartWhat happensWhat usually makes it longHow you measure it
End of turnThe system decides the caller has finished.It waits for a long silence before it decides.From the end of the caller's speech on your recording to the system's decision
TranscriptionThe speech recognition text is finalized.The final text arrives only after the whole sentence is processed.From the decision to the final text
Model replyThe model decides what to do and writes the reply.A long first sentence, a large prompt, a long chain of reasoning before it speaks.From the final text to the reply. On a lookup turn, two pieces: up to the Lookup request, and from the Lookup response to the reply
LookupA system check, such as an order status. Only on lookup turns.A slow system, checks that run one after another, no time limit.From the request to the system's response
Voice startThe voice produces the first audio of the reply.The voice waits for the whole reply text before it starts.From the reply text reaching the voice to the first audio leaving your system
LineTelephony and network carry the audio in both directions.Long call routes, distant regions, weak mobile networks.Test calls from ordinary phones, recorded on the caller's side

Two parts come in two pieces. On a lookup turn, the model works before the Lookup, to decide what to check, and after it, to phrase the answer. Both pieces count as Model reply. The Line comes in two pieces on every turn: the caller's voice reaching your system, and the agent's voice reaching the caller. Both pieces count as Line.

How to measure every turn

Measure turns, not calls: in a call average, one slow lookup turn can hide among seven quick ones. Give every turn one row with the same fields. The example values belong to the slow turn shown later in this guide. Example, not real data.

FieldWhat it holdsExample value
Call IDOne ID shared by the call recording and the system logsC1047
Turn numberThe position of the agent's reply in the call4
Turn typePlain turn or lookup turnLookup turn
Response gapFrom the caller's last word to the agent's first sound, including the Line3.9 s
End of turnYour recording and system timestamps0.7 s
TranscriptionSystem timestamps0.2 s
Model replySystem timestamps, both pieces on a lookup turn0.6 s (0.3 s before and 0.3 s after the Lookup)
LookupSystem timestamps, empty on a plain turn1.9 s
Voice startSystem timestamps0.3 s
LineTest calls on the networks your callers use0.2 s
Unexplained timeThe response gap on your system's recording, minus the parts you log0.0 s
Caller spoke into the gapYes if the caller said "Hello?" or repeated themselves during the gapYes
Cut-offYes if the agent started its reply while the caller was still speaking or had only pausedNo
Task completedPer call: yes if the caller got what they called forYes

Measure the response gap first on your system's call recording. The shared call ID matches each turn in the recording to its system timestamps, which give you End of turn, Transcription, Model reply, Lookup and Voice start. The Line is added in the next step.

The Line is not on your recording. On your system, the clock starts when the caller's audio reaches you and stops when your audio leaves. So place test calls from ordinary phones on the networks your callers use, and record them on the caller's side. The gap on that recording, minus the gap on your system's recording of the same turn, is the Line. Add it to the gaps of your real calls.

Then check the sum. Unexplained time is the response gap on your system's recording minus the parts you log. If it is large, a timestamp is missing. In the example turn, your system's recording shows 3.7 seconds and the logged parts add up to 3.7 seconds, so unexplained time is zero. Add the Line of 0.2 seconds, and the caller hears 3.9 seconds.

Why the average hides the problem

Example, not real data: 400 agent turns from 50 recorded calls on one line, 300 plain turns and 100 lookup turns. Every gap includes the Line, 0.2 seconds from test calls.

Turn typeTurnsMedianAverage90th percentile
Plain turns3001.1 s1.2 s1.8 s
Lookup turns1002.9 s3.2 s4.6 s
All turns4001.3 s1.7 s3.4 s

The average for all turns, 1.7 seconds, sits above the median of 1.3 seconds because the slow turns pull it up. Neither number shows the turns callers remember. The 90th percentile (p90) does: it is the gap that one turn in ten reaches or exceeds, here 3.4 seconds. The turns at or above it are the slowest turns.

Split by turn type as well. Plain turns have a median of 1.1 seconds, lookup turns 2.9 seconds. A week with more lookup turns raises the numbers for all turns, even if no part got slower.

Now put the same 400 turns into four bands by response gap: 130 turns under 1.0 second, 180 from 1.0 to 2.0 seconds, 50 from 2.0 to 3.4 seconds and 40 at 3.4 seconds and up. Then mark the turns where the caller spoke into the gap: the caller said "Hello?" or repeated themselves. That happened in 31 of the 40 slowest turns, and in 7 of the other 360: 3 times in the 310 turns below 2.0 seconds, and 4 times in the 50 turns from 2.0 to 3.4 seconds.

The 400 turns by response gap. The dark part of each bar shows the turns where the caller spoke into the gap: 0, 3, 4 and 31. Example, not real data.

So on this line, speaking into the gap is rare below 2.0 seconds, still uncommon from 2.0 to 3.4 seconds, and common at 3.4 seconds and up. The threshold of this line is near 3.4 seconds, not a general target. Find yours the same way: band your own turns, and look for where speaking into the gap goes from rare to common. Finer bands around that point place it more exactly.

Find the largest part in the slowest turns

The 90th percentile shows how slow the slowest turns are, not why. For that, open each slow turn and find its largest part: the part that took the most time in that turn.

Example, not real data: one of the 40 slowest turns. The caller asks about an order status, so it is a lookup turn. The clock starts at the moment the caller stops speaking, as heard on the caller's side.

One slow lookup turn, part by part: a response gap of 3.9 seconds, with Lookup as the largest part. Example, not real data.

Per part: End of turn 0.7 seconds, Transcription 0.2 seconds, Model reply 0.6 seconds, Lookup 1.9 seconds, Voice start 0.3 seconds and Line 0.2 seconds. Model reply is two pieces of 0.3 seconds: one to decide to check the order, one to phrase the answer. The Line is 0.1 seconds in each direction. Together, the parts make a response gap of 3.9 seconds. Lookup is the largest part, almost half of the gap.

Part times add up like this only inside one turn. Medians and 90th percentiles do not: the medians of the six parts, added together, are not the median response gap. So find the largest part turn by turn, then count.

Example, not real data: the largest part in each of the 40 slowest turns.

Largest partSlowest turns
End of turn8
Transcription0
Model reply5
Lookup27
Voice start0
Line0
Total40

Lookup is the largest part most often, in 27 of the 40. That fits the turn types: 36 of the slowest turns are lookup turns and 4 are plain turns. End of turn is the largest part in 8. In most of them, the caller paused while reading out a number, and the system waited longer before it decided they had finished. Model reply is the largest part in 5.

The Line is never the largest part here, because it is a steady 0.2 seconds on this line. On another line, it could be. Our article on good demos and bad live calls gives a quick test. Gaps on every turn point to the line. Gaps only on lookup turns point to tools.

Fix one part at a time, and watch the opposite signal

Start with the part that is largest most often in your slowest turns, here Lookup. Change one thing, then measure the same turn types again. With two changes at once, you cannot tell which one moved the 90th percentile.

Every fix has a cost, so each row also names the signal to watch:

PartWhat to tryWhat to watch
End of turnA wait that depends on what the caller says: shorter after short answers, such as yes, no or a name, and longer while the caller reads out a number or spells a word.Cut-offs
TranscriptionStreaming speech recognition, so most of the text is ready when the turn ends.Misheard words and numbers
Model replyA short first sentence. A simple reply plan for common questions.Answer quality and completeness
LookupRun independent checks at the same time. Set a time limit with a fallback: hand over to a person, or offer a callback. Give a spoken update before the check.Stale or partial answers. Spoken updates repeated on every turn.
Voice startStart the voice on the first sentence while the rest is written.Unnatural pauses in the middle of a reply
LineTest calls from the networks your callers use. A review of the route with your telephony provider.Audio quality, not only delay

Each row lists options. Try one, measure, then decide on the next.

The cost of a shorter End of turn: cut-offs

A cut-off is the opposite failure of a long gap: the agent starts its reply while the caller is still speaking or has only paused. Example, not real data: the same 400 turns had 6 cut-offs, 5 of them while the caller was reading out an order number or a phone number. Numbers are behind both failures: End of turn was the largest part in 8 of the slowest turns, and in most of them the caller was reading out a number too. Shorten the wait on every turn, and expect more cut-offs. That is why the End of turn row has two waits. A longer wait during a number does not make those 8 slow turns faster, though. What helps there is a signal that the number is complete: when the system knows that an order number has five digits, it can end the turn after the fifth digit instead of waiting for silence.

When a wait cannot be avoided: a spoken update

Some lookups stay slow, and the system behind them is not always yours to change. Then give the caller a first sound before the Lookup: a spoken update, a short line before a slow step. In the slow turn above, it would come right after the first Model reply piece:

  • Caller: "It's order 4 8 1 7 2."
  • Agent, before the Lookup: "Thanks, let me check that order."
  • Agent, after the Lookup: "Your order is on its way, and it arrives on Thursday."

The response gap gets shorter, because the first sound now comes before the Lookup. The time to answer runs from the moment the caller stops speaking to the start of the actual answer. It stays about the same, because the answer still waits for the Lookup. On these turns, log the time to answer too, so an update cannot hide a Lookup that is getting slower.

Keep spoken updates short, and use them only before a slow step, never inside an apology, a number or the closing. An update on every turn has a cost as well: callers hear the same line again and again.

What to track every week

Every week, read three signals separately for plain turns and lookup turns:

  • Response gap: the median and the 90th percentile, and on turns with a spoken update, the time to answer as well.
  • Caller spoke into the gap: the share of turns where it happened.
  • Cut-offs: the count, and what the caller was saying at the time.

Read a fourth signal per call: whether the task was completed. A shorter gap is worth little if fewer calls complete their task. If you run a pilot, count a task as completed the way your pilot card defines a resolved call.

Compare the same turn types before and after each change. A share measured on a few hundred turns can move by chance. Our article on voice AI pilot success criteria shows how to compute the margin of a rate.

Where DRING fits

Selected speech to text, text to speech and reasoning models, telephony, and testing and monitoring are part of the shared operating foundation of every DRING package. On DRING, three of the six parts and the spoken update work like this:

  • End of turn: a dedicated model runs barge-in and end-of-turn detection, so the agent does not talk over a pause or sit through a finished answer.
  • Transcription: the call is transcribed in real time, with streaming speech to text.
  • Line: audio quality and latency signals are tracked per number, so a degrading line is caught early.
  • Spoken update: when a tool or model runs long, a short, natural filler from DRING's approved set keeps the turn moving. The filler is DRING's form of the spoken update, and it never appears in an apology, a number or the closing.

Our speech core page shows how speech to text, turn taking and text to speech run as one pipeline, and our telephony page covers line health.

Every DRING agent runs 1,000 to 10,000 simulated conversations, built for your company, before the first real call. Simulations test turn taking and spoken updates before launch. They do not measure the Line of live calls. The Line comes from test calls on your callers' networks.

If callers on your line say "Hello?" into silences, request a callback. Have a few of those calls ready, with recordings where your policy allows. Our team can go through the slowest turns with you, part by part.

Find the part that makes callers wait

Leave your number and DRING calls you in two minutes. Tell it which line has long silences, and our team follows up to read the slowest turns with you.