30 questions to ask before buying voice AI
Every demo sounds good. Put each vendor through the same calls from your own line, score the answers with the same weights, and the real differences become visible.
The short answer: give every vendor the same test
Voice AI demos are not comparable. Each vendor picks its own script, its own clean test data and a caller who behaves well. The most polished demo often wins, even when it does not fit your calls.
The fix has three parts. Send every vendor the same questions in writing, run every vendor through the same test scenario from your own line, and score the answers with weights agreed before the first demo. This template gives you all three. It assumes you have decided to buy; if not, start with our build vs buy framework.
How to score each answer
Send the 30 questions before the second meeting and ask for evidence with each answer: a document, a sample record, a contract clause or a live run. Have two people score separately, one from operations and one from procurement or IT, and discuss any question where they differ by two points or more.
| Score | What the vendor gave you |
|---|---|
| 0 | No answer, or "we can do that" with nothing behind it |
| 1 | A clear verbal answer, but no evidence yet |
| 2 | A written answer, a document or a partial demonstration |
| 3 | Shown in your shared test scenario, or written into the proposal or contract |
1. Scope and price coverage
Two quotes are comparable only when they cover the same work. These questions turn a price into a scope.
| # | Question | What earns a 3 |
|---|---|---|
| 1 | Which calls, intents, languages and channels does the quote cover, and which are excluded? | A written in-scope and out-of-scope list in the proposal, matching your real call mix. |
| 2 | What is the billing unit, and how is it measured and rounded? | One named unit, its rounding rule, and a sample invoice line for one typical call. |
| 3 | Which costs sit outside the quote? | Telephony, numbers, model usage, integrations and changes after launch, each named with a price or a rate. |
| 4 | What happens when volume runs above or below the plan? | Overage and unused-balance rules in writing, with worked bills at 50% and 150% of your expected volume. |
| 5 | Who builds and maintains the agent, and which changes after launch are included? | A named owner for prompts, tests and integrations, and a written rule for included and quoted changes. |
2. Data ownership, privacy and exit
Every call creates recordings, transcripts and fields about your customers. Settle who controls them before the pilot, not at renewal. Ask for controls that are tested, not only declared. DRING's security standards page shows one way to present them.
| # | Question | What earns a 3 |
|---|---|---|
| 6 | Who owns recordings, transcripts, summaries and extracted fields? | The contract says you do and limits the vendor's use to delivering your service. |
| 7 | Is our data used to train models that serve other customers? | A written no, or an opt-out in the contract, that also covers the vendor's model providers. |
| 8 | Where is data processed and stored, and which subprocessors touch it? | A named region and a current subprocessor list that states each provider's role. |
| 9 | Can we set retention per data type and see who accessed a call record? | Separate retention settings for recordings, transcripts and fields, and an access log you can request. |
| 10 | If we leave, what do we take with us, in what format and how fast? | A written exit list: the data you receive, its format, the deadline, and what stays with the vendor. |
3. System connections and write-back
A logo on an integrations page says a connection is possible, not which fields the agent can read or change in your edition. Ask for a field map, then check it in the scenario.
| # | Question | What earns a 3 |
|---|---|---|
| 11 | Which of our systems will the agent read from and write to? | A field-level map for each system (objects, fields, read or write), checked against your edition and permissions. |
| 12 | What exactly lands in our CRM or helpdesk after a call, and can we add our own fields? | A sample record from your scenario with outcome, summary, next action and your custom fields. |
| 13 | What happens when a write fails or a system is down during the call? | The caller hears an honest next step, the write is retried without duplicates, and a person is alerted. |
| 14 | How can our own systems start a call and read its result? | Documented services to start a call, check its state and fetch its result, plus webhooks and a test environment. |
| 15 | Which actions can the agent take without a person confirming? | A written permission list for each action, with every write enabled only after your sign-off. |
4. Human handover and control
Handover is part of the design, not a failure. Check that your teammate can continue without the caller starting again, and that your team can stop the agent.
| # | Question | What earns a 3 |
|---|---|---|
| 16 | When does the agent hand over, and who sets the rules? | Written triggers you can change, including a caller asking for a person, failed verification and out-of-policy requests. |
| 17 | What does our teammate see at the moment of transfer? | Caller identity, reason for the call, steps completed and the open question, visible before the teammate speaks. |
| 18 | What happens when nobody on our team can take the call? | A callback task or ticket with an owner and a time, created automatically and shown in the scenario. |
| 19 | Can we pause the agent ourselves? | One tested step that stops the automation and sends calls to your team, available to your named operators. |
| 20 | How are handovers reported afterwards? | Handover rate and reason for each intent, and whether the receiving teammate resolved the case. |
5. Quality testing and release
A good first demo says little about the hundredth change. Ask how the vendor proves an agent is ready and that a change broke nothing. DRING describes its testing pipeline step by step; ask each vendor to walk you through theirs.
| # | Question | What earns a 3 |
|---|---|---|
| 21 | What is tested before the first real call, and from what material? | A test set built from your calls, documents and policies, with its size, pass rule and results shared with you. |
| 22 | Who scores test conversations, and what happens when scorers disagree? | Written criteria, more than one evaluator, and a person who reviews every disagreement. |
| 23 | How is a change tested and released, and can it be rolled back? | Every change reruns the full test set, ships only with your approval, and the previous version can be restored. |
| 24 | How is live quality monitored after launch? | A stated share of live calls scored for outcome and policy, and reviewed with you on a fixed schedule. |
| 25 | What happens when the vendor changes the model or speech provider the agent runs on? | You are told in advance, and the agent reruns your test set on the new setup before it takes a live call. |
6. Telephony, service levels and failure
A browser demo skips the carrier, the phone audio and the switch you already run. Treat the telephony layer as its own part of the comparison.
| # | Question | What earns a 3 |
|---|---|---|
| 26 | Can we keep our numbers and our current switch? | A named connection path (call forwarding, SIP or number porting), tested on your carrier, with your current flow as fallback. |
| 27 | How many calls can run at the same time, and what does the next caller hear? | A stated concurrency limit and a demonstrated overflow path: queue, callback or transfer to your team. |
| 28 | How fast does the agent reply on a phone line, and how was that measured? | A figure measured end to end on real phone calls in your language, with the method stated. |
| 29 | What happens when a component fails during a call? | A tested failover path that closes or transfers the call in a controlled way, never leaving the caller in silence. |
| 30 | Which service levels are in the contract, and what happens when they are missed? | Uptime, support response and incident notice times in writing, with the remedy and a named contact. |
Build one test scenario from your own calls
The questions show what a vendor says; the scenario shows what the agent does. Start from a sample of recent calls, for example the last 50, and pick the most common request plus the moments that cause trouble. Create test records in a sandbox of your CRM or helpdesk, with the same fields as production.
Share the outline with every vendor at the same time: intents, systems and test records. Keep the caller's exact lines to yourself, so no vendor can tune the agent to a script.
| Moment | How to stage it | Pass when |
|---|---|---|
| Routine call | Your most common request, on a test record that exists | The answer matches the record and the outcome appears in the test system |
| Interruption | The caller cuts in mid-sentence and changes the request | The agent stops, takes the new request and does not restart its script |
| Missing record | The caller gives an order or account number that does not exist | The agent says so, asks once more, then offers a next step; it never invents a status |
| Angry or confused caller | The caller repeats, raises their voice or mixes two problems | The agent acknowledges the problem, takes one issue at a time and offers a person when your rule says so |
| Handover | The caller asks for a person; run it once with your test queue staffed and once with nobody available | Your teammate sees the reason without asking again; with nobody available, a callback task with an owner appears |
| Write-back check | Open the test system after all the calls | Each call has the right outcome, summary and next action, with no duplicates; the failed lookup is marked as failed |
Run it on a real phone line
- Call a real number from a mobile phone, not from a browser tab, with someone from your team as the caller.
- Record the session with consent and score it from the recording, not from memory.
- Count how often the caller had to repeat something, and time the gap before each agent reply.
- After the first run, ask for one change, such as a new policy rule, and run the scenario again.
How the vendor tests and releases that one change answers question 23 better than any slide. The scenario also provides the evidence for a 3 on questions 12, 17, 18 and 28.
Total the scores with the same weights
Agree the weights before the first demo, so nobody adjusts them to fit a favorite. The split below is an example; move points toward the areas where a failure would hurt your operation most.
| Area | Example weight |
|---|---|
| 1. Scope and price coverage | 15 |
| 2. Data ownership, privacy and exit | 15 |
| 3. System connections and write-back | 20 |
| 4. Human handover and control | 15 |
| 5. Quality testing and release | 20 |
| 6. Telephony, service levels and failure | 15 |
| Total | 100 |
Area score = (points in the area / 15) × area weight. Fifteen is the area maximum: five questions at 3 points each. The six area scores add up to a total out of 100.
Before scoring, mark three to five must-have questions, such as 6, 13 and 17. A vendor that scores 0 or 1 on any must-have is out, whatever its total.
A worked example
The scores below are invented to show the arithmetic. Vendor A gave the smoothest demo; Vendor B sounded plainer but showed its write-back and test process in the scenario.
| Area (weight) | Vendor A points | Vendor A score | Vendor B points | Vendor B score |
|---|---|---|---|---|
| Scope and price (15) | 12 | 12.0 | 10 | 10.0 |
| Data and exit (15) | 9 | 9.0 | 12 | 12.0 |
| Connections (20) | 6 | 8.0 | 12 | 16.0 |
| Handover (15) | 8 | 8.0 | 11 | 11.0 |
| Quality (20) | 7 | 9.3 | 12 | 16.0 |
| Telephony (15) | 12 | 12.0 | 10 | 10.0 |
| Total | 54 | 58.3 | 67 | 75.0 |
Vendor A also scored 1 on question 13, a must-have, so it is out whatever its total. Its demo was strong on the conversation and weak on what comes after it: the record, the failure path and the next release.
Red flags in the answers
- "We integrate with everything", with no field-level map for your systems.
- No answer to "what does the caller hear when your system is down?"
- Handover means forwarding the call to a number, with no context for your teammate.
- Test results reported as one overall percentage, with no method or sample size.
- Changes or model updates go live without your approval or notice.
- Exit terms that say data "can be made available", with no format or deadline.
One red flag is a reason to ask again in writing. Several in one area show you where the real work would land on your team.
How DRING answers these questions
Use the same template on DRING. Some answers we can already give in writing:
- Question 21: every DRING agent runs 1,000 to 10,000 simulated conversations, built for your company, before the first real call. They are generated from your workflow, documents and call recordings, not from a generic benchmark set.
- Questions 11 and 12: HubSpot is the documented reference implementation for choosing who to call, running the call and writing the result back. Connect links Salesforce, Freshdesk and other mainstream CRM and helpdesk systems on the same terms; the exact objects, permissions, custom fields and actions are confirmed during onboarding. Additional analysis fields can be defined and come back the same way through the API, CRM, dashboard and reports.
- Question 14: services to start a call, check its state and query its result, plus lifecycle and result webhooks, are standard for the scoped workflow.
- Questions 1 and 27: a DRING package is recommended by expected usage, concurrent call capacity, number of agent types, channels, live transfer needs, support level and integration depth. On languages, 62 are available and 10 are live today; yours is validated before launch.
Where DRING may score lower at first: a niche or heavily customized CRM is assessed separately before any non-standard development is quoted. Until that review is done, question 11 may stay at 2. A language outside the live set should count as unproven for your line until your scenario has run in it.
Keep the scorecard after you sign
The answers that earned a 3 become the acceptance criteria for the pilot. The six-moment scenario becomes the first regression test, so every later change is checked against the calls that decided the purchase.
Put DRING through your own test
Get a call in two minutes to share your workflow and the hard calls you would test. We will help turn them into one shared evaluation scenario.