Skip to content
Share one workflow. DRING calls in two minutes and qualifies the need. Get a call in two minutes
Get a call in 2 minutes See Agent Factory
Buying guide · Service model decision

AI agent, outsourcing or in-house team?

Do not pick one model for the whole company. Score each call type on the same factors with your own weights, and let the mix follow from the scores.

The short answer

Make this choice per call type, not once for the whole company. A call type is one workflow with one expected outcome, such as order status, a booking change, a complaint or a quote follow-up.

Score every call type on the same seven factors: volume, complexity, hours, languages, control, demand variability and exceptions. Give each factor your own weight, score each option from 1 to 5 and compare the weighted totals. The highest total wins that call type.

The result is often a deliberate mix. In the worked example below, order status goes to an AI agent, return requests go to an outsourced call center and damage complaints stay with the in-house team. Each option also has clear cases where it is the wrong choice. For the AI agent, these include calls that depend on judgment, carry high stakes or are emotionally loaded.

Why one model for the whole company fails

An order status call and a damage complaint share a phone number and little else. The first is a quick lookup that can come many times a day, at any hour. The second is rare, emotional and ends with a decision about money.

A company-wide choice averages these calls into one answer, so some call types end up with an option that fits them poorly. The AI agent takes the angry customer, skilled staff spend the day reading out delivery dates, or a contract written for simple calls has to cover your hardest ones. Make the decision at the level where calls actually differ: the call type.

Step 1: list your call types

"Where is my order?" is a call type. "Customer service" is not, because it contains many workflows with different outcomes.

Group one recent month of calls from your phone system or CRM by reason. For each call type, write down:

  • Volume: calls per month, and how much that changes in peak weeks.
  • Hours: the share of calls that arrive outside your office hours.
  • Languages: the languages your callers use.
  • Exceptions: how often a call leaves the expected path.
  • Stakes: what a wrong answer costs, in money, customers or legal risk. The stakes set how much complexity and control weigh in Step 2.

Monthly totals hide peaks, so also check the busiest week and hour, as our guide to seasonal call volume planning explains. Start with the five to ten call types that carry the most volume or cause the most trouble.

Step 2: set the weights before you score

Each factor gets a weight from 0 to 3: 0 means it does not matter for this call type, 3 means it matters a lot. Set the weights per call type, because the same factor matters differently for different calls. Control may weigh 3 for complaints and only 1 for order status.

Write the weights down, with one line of reasoning each, before anyone scores the options. Once the scores are visible, it is easy to adjust the weights until a favorite option wins. In our example, the weights for each call type add up to 10 to keep the arithmetic simple, but any total works.

Step 3: score each option from 1 to 5

For each factor, score how well each option handles this call type:

  • 5: a strong fit as it is.
  • 3: workable, with extra setup, cost or supervision.
  • 1: a poor fit that needs a workaround.

Use 2 and 4 for the steps in between. A score belongs to the option and the call type together: the same AI agent can score 5 on complexity for order status and 2 for complaints. The decision table shows when to weight a factor high and what earns each option a high score.

FactorYour weight (0 to 3)AI agent scores high whenOutsourcing scores high whenIn-house team scores high when
Volume3 if this call type is a large share of your callsHigh, repeatable volume spreads setup and testing over many callsThe volume fills the seats you contract; check minimum volumesThe volume fits your current team without new hires
Complexity3 if the call needs deep knowledge, several systems or a costly decisionThe steps can be written down and the systems can be connectedThe provider's staff can be trained on your policy and given system accessThe call needs deep product knowledge or decisions only your staff can make
Hours3 if many calls come in evenings, nights or weekendsCalls come at any hour, and nobody has to staff a night shiftThe contract covers the hours you need, including nights and weekendsCalls come during the hours your team already works
Languages3 if callers use languages your team does not speakEach language is tested for this workflow on a phone line, not only listedThe contract names native-speaking staff for each language and shiftYour own people speak every language this call type needs
Control (quality, data, brand voice)3 if the call carries brand risk, sensitive data or a decision about moneyYou need the approved wording on every call and a record of each oneThe contract gives you recordings, quality reports and clear data rulesYou want to hire, train, listen to calls and decide yourself
Demand variability (peaks)3 if volume jumps with campaigns, seasons or incidentsPeaks are sudden, and your agreed concurrent call capacity covers themYou can add trained seats at short notice; check the notice periodDemand is steady enough to staff without idle hours or long queues
Exceptions3 if calls often leave the expected path, or mistakes there are costlyExceptions are rare and known: the agent handles them or hands overThe contract lets the provider's staff resolve them; check what they can decideYour people can decide, ask a colleague or change the rule

To get the weighted total, multiply each score by its weight, add the results and divide by the sum of the weights. The total stays on the 1 to 5 scale. When two totals are close, do not let a decimal decide: look again at the factor with the highest weight, or split the call type, as shown below.

Cost is not a row, because volume, hours and peaks drive the bill for all three options. Once the scores narrow a call type to two options, compare them on cost per resolved call, not on price per minute or per seat.

Worked example: one retailer, three call types

All numbers in this section are an example made up for this article, not data from a real company.

The example company sells large home appliances online in two countries and delivers them with its own crews. Its in-house team works weekdays during office hours and speaks the home language; two teammates also speak the second language. Three call types carry most of the calls:

  • Order status: about 6,000 calls a month. 40% arrive outside office hours, and volume doubles in sale weeks. Few calls leave the expected path.
  • Return request: about 1,500 calls a month in both languages, many of them in the evening. About one call in five needs a judgment: a used item, missing parts or a request outside the return window.
  • Damage complaint: about 300 calls a month, mostly during office hours. The caller is upset, and the call ends with a decision on a refund or a replacement.

For order status, the AI agent's total is 3 × 5 + 0 × 5 + 2 × 5 + 1 × 4 + 1 × 4 + 2 × 5 + 1 × 3 = 46, divided by 10: 4.6.

Order statusWeightAI agentOutsourcingIn-house team
Volume3542
Complexity0545
Hours2542
Languages1452
Control1435
Demand variability2532
Exceptions1345
Weighted total4.63.82.6
Return requestWeightAI agentOutsourcingIn-house team
Volume1443
Complexity2345
Hours2541
Languages2452
Control1435
Demand variability0532
Exceptions2245
Weighted total3.64.13.4
Damage complaintWeightAI agentOutsourcingIn-house team
Volume0335
Complexity3235
Hours0544
Languages1452
Control3235
Demand variability0532
Exceptions3135
Weighted total1.93.24.7

Why these results:

  • Order status goes to the AI agent (4.6). Volume, hours and peaks carry most of the weight, and those are its strongest factors.
  • Return requests go to Outsourcing (4.1). The outsourced call center combines a person's judgment with evening hours and native speakers in both languages. The in-house team judges best but is not there in the evening.
  • Damage complaints stay with the in-house team (4.7). Complexity, control and exceptions carry the weight, and there the AI agent scores lowest.

The AI agent scores 4 on languages, not 5, because in this example its second language has not yet been tested on these calls.

Example, not real data: one retailer's three call types, each routed to the option with the highest weighted total. Order status goes to an AI agent (4.6), return requests to Outsourcing (4.1) and damage complaints to the in-house team (4.7).

No single model wins all three. The result is a deliberate mix on one customer line, and each part has a documented reason.

Change the weights, change the winner

Now keep every score for return requests and change only the weights. The weights above put judgment first: the company wants a person to decide each return. Suppose the company put coverage first instead and wanted every return call answered quickly, including evenings and the weeks after a sale. Hours and demand variability go up. To keep the weights at a total of 10, complexity, languages and exceptions each give up one point.

Return requestWeight, judgment firstWeight, coverage first
Volume11
Complexity21
Hours23
Languages21
Control11
Demand variability02
Exceptions21
AI agent, weighted total3.64.2
Outsourcing, weighted total4.13.8
In-house team, weighted total3.42.7

Same call type, same scores, different winner. The AI agent rises from 3.6 to 4.2 and passes Outsourcing, which falls from 4.1 to 3.8. Nothing about the options changed. Only the priorities did.

Example, not real data: return requests with the same scores under two sets of weights. With judgment first, Outsourcing wins with 4.1. With coverage first, the AI agent wins with 4.2.

Two lessons follow. First, the weights are the real decision. Second, when the winner depends on how you weigh judgment against coverage, the call type often holds two workflows. Split it: routine returns inside the return window go to one option, and returns that need a judgment go to another. A handover keeps both on the same line. Our guide to human handover in voice AI covers what the receiving person needs.

When each option is the wrong choice

Every option has calls it should not take. Use these lists to check your scores: if the winner of a call type matches one of its own warning signs, look again.

AI agent

  • Calls that depend on judgment: disputes, goodwill decisions and negotiations that no written rule covers.
  • High-stakes calls: safety incidents, medical or legal questions, or large sums, where one wrong answer costs too much.
  • Emotionally loaded calls: a caller who is upset, grieving or frightened and needs a person to listen.
  • Calls that are mostly exceptions: when almost every call is different, there is no path to follow.
  • Calls without system access: if the agent cannot read the order or change the booking, it can only take a message.

Outsourcing

  • Knowledge that changes every week: prices, rules or products that change faster than an outside team can be retrained.
  • Decisions you will not delegate: refunds above a limit or exceptions to policy.
  • Data a third party may not access: when your data rules keep the records this call needs inside the company.
  • Volume that does not fit the contract: a call type too small or too irregular for the minimum volumes and notice periods you can agree on.
  • Calls that define your brand: when you want to hear and shape every one of them yourself.

In-house team

  • Calls outside your hours: when evening and weekend calls mean lost sales or failed deliveries, and you do not want to staff nights.
  • Sharp peaks: you staff for the peak and pay for idle hours, or staff for the average and lose calls in the peak.
  • Languages your team does not speak: hiring for a language that brings a few calls a day is hard to justify.
  • High-volume lookups: skilled people spend their time on routine answers instead of solving problems.
  • Capacity needed fast: when you need more people sooner than you can hire and train them.

How to start: one call type, then the next

  1. Pick one call type. Take the one with the most volume, or the one that causes the most trouble today.
  2. Agree on the weights first. Operations and purchasing set them together, with the owner of the outcome.
  3. Score separately, then compare. Two people score on their own and talk through any factor where they differ by two points or more.
  4. Test before you move the whole call type. Run the winner on part of the traffic, keep the current model as a fallback and compare outcomes on the same kind of calls. If the AI agent wins, our build vs buy framework covers the next decision.
  5. Write the factors into the agreement. Whichever option wins, the hours, languages, peak capacity, exception rules and data access you scored belong in the contract or the scope.
  6. Move to the next call type. Score a call type again when something changes: a new market, a new language, a season or a new policy.

Where DRING fits

DRING provides the AI agent option in this table. If one of your call types scores high for it, test that one workflow with a working agent before you decide on the rest. Bring a month of calls for it, with recordings where your policy allows, and the outcome each call should reach. Every DRING agent runs 1,000 to 10,000 simulated conversations, built for your company, before the first real call.

Every DRING package includes human handover with the conversation attached, and live call transfer is available from the Pro package up. That lets the mix run on a single line: the agent takes the call types it scored well on and transfers the others to the people who handle them. If languages weigh heavily for a call type, DRING has 62 languages available, 10 live today; yours is validated before launch. Our guide to multilingual voice AI explains how a language is validated.

Start with one call type

Leave your number and DRING calls you in two minutes. Tell it which call type you want to assess, and our team follows up with what to bring.