Good demo, bad live calls: which layer broke?
Your voice agent passed every demo but struggles on real calls. The failed calls themselves show which layer is breaking and what to fix first.
The short answer
A voice agent that passes the demo and fails on live calls has not suddenly become worse. Live calls cross the same seven layers as the demo, from line audio and speech recognition to tools, handover and measurement. The demo never tested the layer that is now breaking under real conditions.
So do not start by rewriting the prompt. Take the real failed calls, walk each one through the layers in order and assign it to the first layer that broke. The layer with the most calls is the one that is failing, and it is your first fix. Make that one fix, run the same calls again and compare.
This article is for a line that is already live. If you have not launched yet, start with our guide to testing a voice AI agent before production.
Why the demo passed
A demo is a friendly test. The caller knows the script, speaks clearly in a quiet room and waits for the agent to finish. The test record is clean: the order exists and the slot is free.
Real callers interrupt. They call from a car or a warehouse, use local words and abbreviations, and ask about related topics the flow was not built for. Sometimes they are not the person the flow expected: an accountant answers for the owner, a relative picks up, or the call reaches another department.
Each of these conditions tests a different layer. A good demo shows that the layers hold on clean input, not which one breaks first on real traffic.
The seven layers a call passes through
Every call crosses the same layers in the same order. A break in an early layer produces symptoms in the later ones, so symptoms alone often point to the wrong place.
- Line audio and telephony: the call connects, audio flows both ways without delay or echo, and the call ends cleanly.
- Speech recognition: what the caller said becomes text.
- Knowledge: the facts the agent can use, such as policies, product details and answers to common questions.
- Decision: the conversation logic, meaning what to say or ask next, when to confirm, when to hand over and how to end.
- Tools and system actions: lookups and writes in the CRM, booking or ticketing system.
- Handover to people: the transfer or callback, and the context your teammate receives.
- Measurement: the recorded outcome matches what actually happened.
One more gate is not an AI layer. Some calls fail on the business side: no stock, no free slot, or outdated data in your own system. The agent did its job and the business could not deliver. Count these calls separately and send them to the owner of that process.
Start with the real callers, not the test sheet
Take all real calls from a recent period, not the scenarios from your test sheet. Then remove internal test calls: the same number and the same scripted lines, repeated a few minutes apart. A tester does not talk like a real caller.
Next, check the outcome labels on a sample. If voicemails count as conversations or completed bookings count as failures, the failed set itself is wrong. Measurement is the last layer, but the first to verify.
Finally, listen to the audio. The transcript shows what recognition heard, and only the recording shows what the caller said.
Diagnosis table: from symptom to layer
Use the table on one failed call at a time, and find the evidence before you assign the layer.
| Symptom | Likely layer | What to check | Evidence that confirms it |
|---|---|---|---|
| Outbound calls end in the first seconds | Line audio and telephony | Time from pickup to the agent's first word | People hang up after a silent gap, before the agent speaks. Hang-ups during the opening line point to decision. |
| Long silence after the caller stops talking | Line audio and telephony, or tools | The gap before each agent reply | Gaps on every turn point to the line. Gaps only on turns with a lookup point to tools. |
| The agent stops mid-sentence when nobody spoke | Line audio and telephony | Each stop, against the recording | It stops on traffic, machinery or its own echo. |
| Callers repeat themselves | Speech recognition | The transcript of the first attempt, against the audio | The speech is clear in the recording, but the transcript is wrong or empty. |
| The agent answers a question nobody asked | Speech recognition | The heard text, against the question the agent had just asked | The transcript holds a sentence the caller never said, and the agent answered it. A correct transcript points to decision. |
| Names, product codes or addresses are captured wrong | Speech recognition | The same terms across many calls | The same wrong forms repeat across different callers. |
| A confident answer that is wrong or out of date | Knowledge | The source behind the answer, and its date | The agent's material itself is wrong or out of date. Outdated data in your own system belongs to the business-side gate. |
| "I don't have that information" on a common question | Knowledge | Whether the topic exists in the agent's material | The question is frequent in real calls and missing from the material. |
| The agent asks again for details the caller already gave | Decision | The turn where the caller answered | The transcript is correct, and the next agent turn ignores it. |
| The wrong person is on the line and the script carries on | Decision | How the flow handles an accountant, a relative or another department | The agent keeps asking questions this person cannot answer, instead of asking who to speak to. |
| The agent says it booked, but nothing appears in the CRM | Tools and system actions | The tool log for that turn: request, response, record ID | An error or no record ID came back, and the agent confirmed anyway. No request at all points to decision. |
| After a transfer, your teammate asks the caller to start again | Handover to people | What the teammate saw at the moment of transfer | The summary was missing, or it arrived after the teammate answered. |
| Callers are cut off at the end of the call | Decision | Who ended the call, and the caller's last sentence | The agent hung up before the caller said goodbye. A call that neither side ended points to the line. |
| The dashboard says resolved, but complaints say otherwise | Measurement | A sample of calls marked resolved, against the outcome definition | Those calls did not reach the agreed outcome, or they were voicemails. |
Count each failed call once, at its first break
This is how we read failed calls at DRING. Each failed call passes the layers in order and gets one label: the first layer that broke. A misheard sentence that leads to a wrong reply and then a failed booking is one recognition break, not three problems.
With one value per call, the counts add up to the total number of failed calls, and the largest count shows where one fix can recover the most calls.
The common alternative is to tag every symptom wherever it appears. Counted this way, the 200 failed calls in the example below received 355 tags. Decision collected the most tags, 104, because a misheard sentence usually makes the next reply wrong as well. Yet decision was the first break in only 40 calls. Symptom tags overlap, their total does not match the number of failed calls, and they make every layer look urgent. Show them separately if useful, but rank fixes on first breaks.
The figures in this section are an example for this article, not data from a real deployment.
| First layer that broke | Failed calls | Share |
|---|---|---|
| Speech recognition | 62 | 31% |
| Decision | 40 | 20% |
| Tools and system actions | 24 | 12% |
| Business side, not an AI defect | 22 | 11% |
| Line audio and telephony | 18 | 9% |
| Knowledge | 14 | 7% |
| Handover to people | 12 | 6% |
| Measurement | 8 | 4% |
| Total | 200 | 100% |
In this example, speech recognition is the first layer to work on: it is the first break in 62 of 200 calls. The 22 business-side calls are not an AI defect. No change to the agent creates a free slot or corrects an outdated price list in your own system.
Speech recognition: most fixes sit in the decision layer
The first break shows where a call went off course, not always where the fix goes. On a phone line, many recognition errors cannot be fixed at the recognition layer: the audio is compressed, and a forklift can be louder than the voice. The conversation has to stay correct when the text is wrong, and that work happens in the decision layer:
- Do not answer a sentence that does not fit. When the heard text does not match the question the agent just asked, the agent does not act on it or open a new topic. It says it did not catch that and asks again.
- Keep a list of known mishearings. Words that recognition gets wrong in real calls go on a list, and the agent confirms them with a question instead of guessing.
- Use system data. When the order or record is already known, the agent reads it back and asks "Is it this one?"
- Add a keyword list last. A list of names and product terms is a cheap recognition setting, but it comes after the three steps above.
An example written for this article, not a real call: one misheard turn, handled two ways.
| Turn | What is said |
|---|---|
| Agent | "Your order arrives on Thursday between nine and twelve. Does that work for you?" |
| Caller, driving | "Yes, Thursday's fine. Can you leave it next door?" |
| What recognition heard | "Yes, Thursday's fine. Can you cancel the order?" |
| Agent, handled badly | "Of course, I've canceled your order. Is there anything else I can help you with?" |
| Agent, handled well | "Thank you, Thursday it is. Sorry, I didn't catch the last part. Could you say it again?" |
| Caller | "Can you leave it next door if I'm out?" |
| Agent | "No problem. Thursday between nine and twelve, and if you're out, we'll leave it with your neighbor." |
The first reply turns a recognition error into a canceled order. The second keeps the part that fits and asks again for the rest. For more on this layer, read how voice AI fixes what it mishears.
Tools and handover: trust the log, not the sentence
An agent that says "Your appointment is booked" has produced a sentence, not a booking. Check the tool log for every action: was the request sent, what came back, and was there a record ID? The agent should confirm an action only after a successful response. Our guide to CRM write-back covers what a call should leave behind.
Handover breaks just as quietly. The transfer happens, but the summary arrives late or not at all. Listen to the first seconds after each transfer. If your teammate asks why the caller is calling, the handover failed. See designing human handover for what the context should contain.
Endings: count the goodbye
Many flows end the call on a timer or as soon as the task is done. On live calls, this cuts people off in the middle of a question or a thank you. The agent should not hang up before it hears the caller's goodbye.
For every call the agent ended, read the caller's last sentence before the hang-up. Was there a goodbye in it? The share of calls without one is your cut-off rate, and it belongs to the decision layer.
Fix one layer at a time
The first-break table sets the order of work:
- Pick the AI layer with the largest count. The business-side gate goes to its owner.
- Make one change aimed at those calls. For recognition breaks, that change usually sits in the decision layer, as described above. With two changes at once, you cannot tell which one helped.
- Run the same failed calls again. They are now your regression set: the same caller lines and, where possible, the same audio conditions.
- Compare first-break counts before and after on the same kind of traffic: the same campaign, list type or inbound hours. An easier list makes any change look good.
Expect some calls to move rather than disappear. A call that now passes recognition may break later, at tools. That is progress: the next fix is now visible. Agent Factory, the way DRING builds, tests and improves agents, follows the same rule: one change at a time, tested before it is released.
Where DRING fits
If a live flow is underperforming, bring it to a callback. Have a sample of failed calls ready, with recordings where your policy allows, and the outcome each call should have reached. Our team can go through them with you as this article describes and look at where each one broke first.
Changes that touch your own systems are scoped separately. Which actions an agent can take in a CRM depends on the CRM's edition, permissions and API access. On DRING, additional analysis fields can be defined and returned through the API, CRM, dashboard and reports. The first layer that broke can be one of them.
Every agent DRING builds runs 1,000 to 10,000 simulated conversations, built for your company, before the first real call. Live calls still show what simulation missed, so improvement after launch starts from real callers.
Find the layer that is failing
Leave your number and DRING calls you in two minutes. Tell it which flow underperforms, and our team follows up with where to look first.