Skip to content
Share one workflow. DRING calls in two minutes and qualifies the need. Get a call in two minutes
Get a call in 2 minutes See Agent Factory
Launch guide · Pilot evaluation

How to judge a voice AI pilot on day 30

Decide what success means before the first live call. This guide gives you a pilot card, a weekly review template and a day 30 rule to Expand, Extend or Stop.

The short answer

A voice AI pilot has succeeded when every metric meets its target on day 30 and no stop line is crossed. The targets and stop lines sit on a pilot card signed before the first live call. If you set the criteria after the calls, day 30 becomes an argument about which calls count. If you set them before, day 30 is a short meeting.

Measure a baseline for the same call type in a comparable period. Put five metrics on one pilot card: resolution rate, repeat contact rate, wrong action rate, transfer quality and human effort. Give each one an exact definition, a target and a stop line, and read the same card in every weekly review.

On day 30, apply the rule you agreed on in advance. Expand when every target is met. Extend once, by two weeks, when one or two metrics are short of target and one known cause explains them. Stop when a stop line is crossed, more than two metrics are short or no single known cause explains the shortfall. Some stop lines need only one call, such as a wrong action with money or legal consequences.

For the build steps before the pilot, see our 90-day voice AI implementation plan.

Why impressions are not a verdict

After 30 days, the manager remembers the one call where the agent misunderstood an angry customer. The project team remembers the smooth calls it played in meetings. Neither says much about the hundreds of calls nobody heard.

Numbers chosen afterward have the same problem: it is easy to pick the metric that looks best, or to stretch the word "resolved" until it fits. So the operations owner, the quality reviewer, the systems owner and the sponsor sign the pilot card before day 1. During the pilot, the targets and stop lines do not change.

Step 1: measure the baseline

The baseline is how the same call type performs today, measured with the pilot definitions. Without it, a resolution rate of 70% is neither good nor bad. A call type is one workflow with one expected outcome. In the example used here, customers call to move a booked technician visit, and the call should end with a new slot the customer accepted.

Take the baseline from a comparable period:

  • Same weekdays and hours: four full weeks, covering the hours the agent will answer.
  • Same season or campaign: not a sale week against a quiet week, and not a holiday week against a normal one.
  • Same call mix: the same call type, languages and customer groups. If the pilot takes only evening calls, the baseline is what happened to evening calls before, even if that was voicemail.

Do not reuse old wrap-up codes: a "resolved" click can mean something different from a new slot in the booking system. Relabel a random sample of baseline calls with the pilot rules, and use it to measure your team's wrong actions too. Our article on containment vs resolution rate explains why a call without a transfer is not yet a resolved call.

If a metric was never measured, write "not measured" on the card, set its target from what the business case needs, and do not claim an improvement on it later.

Step 2: write the pilot card

The pilot card holds five metrics, each answering a different question, so a good number on one cannot hide a bad number on another. For each metric, write the numerator, the denominator and what is excluded. Then set two levels. The target is where the business case works. The stop line is where you stop, whatever the other metrics show. A metric between the two is short of target.

All values in the card are an example made up for this article, not data from a real company.

MetricExact definitionBaselineTargetStop line
Resolution rateCalls that end with the visit on a slot the customer accepted, shown in the booking system, with no transfer or callback, divided by eligible calls. Excluded: test calls and other call types.74%70% or moreBelow 60%
Repeat contact rateCustomers who contact you again about the same visit within 7 days, by any channel, divided by customers with an eligible call. Counted once the 7 days have passed.12%12% or lessAbove 18%
Wrong action rateReviewed calls with a wrong record, a wrong booking or wrong information, divided by reviewed calls. Review every call in week 1, then 100 random calls a week.1.5% (3 of 200 reviewed calls)1% or lessAbove 3%
Transfer qualityTransfers where the teammate did not have to ask again for the name, booking or reason, divided by all transfers.Not measured90% or moreBelow 75%
Human effortStaff minutes per 100 eligible calls on transfers, callbacks, reviews and corrections, judged on weeks 3 and 4. Baseline: minutes spent handling the calls.640 minutes250 minutes or lessAbove 400 minutes

The resolution rate target sits below the baseline on purpose. Write the reason on the card: the business case works at 70% if human effort falls from 640 to 250 minutes or less per 100 calls. Human effort counts only weeks 3 and 4, because reviewers read every call in week 1 and the first fix is made in week 2.

Stop lines that need only one call

Some failures cost too much to wait for a rate:

  • A wrong action with money or legal consequences, such as a charge, a refund or a canceled paid visit.
  • Personal data given to a caller who did not pass the identity check.
  • A caller who reports a danger, such as a gas smell, and is not passed to a person at once.

Any one of them ends the pilot early, and calls go back to the prior route the same day. A restart after a fix is a new pilot from day 1. A stop line on a metric works more slowly: the pilot ends early if a metric's measured value is across its stop line in two weekly reviews in a row. Count only the reviews in which that metric is judged: human effort from week 3, and repeat contacts once their 7 days have passed.

The random sample of 100 calls a week will miss most of the failures listed above, so search for them directly. Every week, read every call that ends in a charge, a refund or a cancellation, every call in which the agent gave out personal data, and every call that a customer or teammate reports as a problem. These reads come on top of the random sample.

Step 3: run the weekly review

The weekly review reads the same card on the same day each week, checks the counting and chooses at most one change to the agent.

WeekWhat to readQuestion to answerWho owns the decision or fix
Week 1Every call without the outcome, every transfer and every booking change, checked against the booking system.Do the labels match what a reviewer hears?The quality reviewer reads. The operations owner decides changes.
Week 2The card beside the baseline, with counts and margins. Unresolved calls grouped by cause.Which cause explains the most unresolved calls: the agent, the systems or the business?The operations owner picks one fix. The agent builder or systems owner makes it.
Week 3The metric the fix should move, and the other four. Repeat contacts for weeks 1 and 2, whose 7 days have passed.Did the fix move its metric without hurting another?The operations owner, with the quality reviewer.
Week 4The full card on all calls since day 1, the stop lines and the margin of each rate.Which day 30 decision does the card support: Expand, Extend or Stop?The operations owner prepares it. The sponsor signs it.

Three rules keep the weeks comparable:

  • One change per week, logged with its date. With two changes, you cannot tell which one moved the number.
  • Fixed definitions. If a definition proves wrong, change it once, write down why and recompute all earlier weeks.
  • Counts beside rates. Write "64% (256 of 400 calls)", not "64%".

Our guide to call center quality sampling covers which calls reviewers should read.

What the numbers can and cannot tell you

Every pilot rate is an estimate. For a rate p measured on n calls, an approximate 95% margin of error is 1.96 × √(p(1 - p) / n). At a rate of 70%, that gives:

  • 100 calls: about ±9.0 points
  • 400 calls: about ±4.5 points
  • 1,000 calls: about ±2.8 points

So a weekly figure on 400 calls is usually within about 4.5 points of the true rate, and a change of a few points from one week to the next can be chance alone. The formula assumes independent calls and a rate not close to 0% or 100%. A comparison with the baseline has a wider margin: for two rates of about 70% on 1,600 calls each, the difference carries about ±3.2 points.

In week 1 of the example, the resolution rate is 64% on 400 calls, about ±4.7 points. The true rate could be anywhere from about 59.3% to 68.7%: the range crosses the stop line and stays below the target. By day 30, the example has 1,600 calls at 68%, about ±2.3 points. The range, about 65.7% to 70.3%, is clearly above the stop line, with the target just inside it.

Example, not real data: the resolution rate against its stop line, target and baseline. The week 1 range (400 calls) crosses the stop line. The day 30 range (1,600 calls) is narrower, with the target just inside it.

Three more limits apply:

  • Rare events. No wrong action in 300 reviewed calls still fits a true rate of up to about 1%. This is the "rule of three": with no events in n calls, the 95% upper limit is about 3 / n. Read and classify every wrong action.
  • Small slices. One language or weekday with 60 calls has a margin of about ±11.6 points at 70%. Listen to those calls instead of comparing their rate.
  • Recent repeat contacts. On day 30, the repeat contact rate covers only calls up to day 23. Later calls have not had their 7 days yet.

The day 30 decision: Expand, Extend or Stop

On day 30, compute the pilot card on all calls since day 1. Then apply the rule in this order:

  • Stop if a stop line is crossed, more than two metrics are short, or no single known cause explains the shortfall. Calls go back to the prior route, and the reason is written down for the next pilot.
  • Expand if every target is met and no stop line is crossed. Move more calls of the same call type to the agent in steps, with the same card. A new call type gets its own pilot card.
  • Extend if one or two metrics are short for one known cause. Extend once, by two weeks, with one fix for that cause and the same card, and judge the extension on its own calls. If every target is then met, Expand. If not, the decision is Stop.
The day 30 decision: three paths, each with a condition agreed on before day 1.

How to read mixed results

A common day 30 result is mixed: four metrics meet their targets and one does not.

  • Do not average. A high resolution rate does not make up for a wrong booking.
  • Protect the customer first. Wrong action rate and the one-call stop lines are never traded against resolution rate or human effort. If the wrong action rate is short of target, Extend only if every wrong action has the same known cause and the fix removes it.
  • Read the miss against its margin. 68% does not meet a 70% target. But a miss smaller than the margin, with a known cause, is a typical case for Extend.
  • Find where the cause sits. If calls fail because no free slot exists, a better agent will not fix that. If human effort misses its target, check whether it is still falling from week 3 to week 4.

Example: one day 30 card

Example values, not real data.

MetricTargetDay 30 resultReading
Resolution rate70% or more68% (1,088 of 1,600 calls, ±2.3 points)Short of target, by less than the margin
Repeat contact rate12% or less11% (calls up to day 23)Met
Wrong action rate1% or less0.6% (4 of 700 reviewed calls)Met, each case read and classified
Transfer quality90% or more92% (230 of 250 transfers)Met
Human effort250 minutes or less230 minutes (weeks 3 and 4)Met
Stop lines that need only one call0 cases0 casesNo stop line crossed

One metric is short and no stop line is crossed. In the largest group of unresolved calls, the customer asks for an evening slot, which the booking system shows only to staff. The cause is known and one fix addresses it, so the decision is Extend: two weeks, the same card, and evening slots opened to the agent. At about 400 calls a week, the extension gives about 800 new calls to judge, with a margin of about ±3.2 points at 70%.

Where DRING fits

With DRING, an agent is live in a week. Use that week to relabel the baseline sample from past calls and sign the pilot card before the first live call. Every DRING agent runs 1,000 to 10,000 simulated conversations, built for your company, before the first real call. That is a test before launch, not the pilot. Simulation can find gaps early, but it gives no baseline and no resolution rate on real customers. Our article on simulated conversations explains what they cover.

During the pilot, after-call analysis can return extra fields that you define, such as the outcome of each call, through the API, your CRM, the dashboard and reports. Reviewers still read the calls, and the fields give every weekly review the same labels. Our call analytics page shows the dashboard and reports.

Agree on the criteria first

Leave your number and DRING calls you in two minutes. Tell it which call type you want to pilot, and our team follows up to agree on the pilot card with you before the first live call.