AI Voice Agents in the Call Center: Which Calls the Agent Should Take, and Which It Must Not
AI voice agents in the call center are not a containment problem. They are a sorting problem. Every call ends in something being decided — a payment arrangement, a booking, a routing, a refusal — and the only question that matters before deployment is which of those decisions the agent is allowed to make on its own, which it should prepare for a person, and which it must never touch. Get the sort right and the technology is unremarkable. Get it wrong and you have a system that sounds confident on exactly the calls where being wrong is expensive.
The demo answers the wrong question
Voice agents are sold on containment rate and cost per call. Both are real numbers and both are downstream of a decision nobody makes explicitly: the perimeter.
The failure that operations leaders actually fear is not the agent that stalls and transfers. That is a mildly annoying call. It is the agent that handles a conversation smoothly, end to end, and commits the company to something it should not have — a payment arrangement outside policy, a confirmation that contradicts the record, a tone taken with someone in genuine distress. The call sounds good. The transcript reads well. The problem surfaces weeks later, in a complaint or an audit, at which point the containment number that justified the rollout is the thing pointing at the damage.
You cannot draw that perimeter after an incident. It is a design step, and it happens before the first call.
Three lanes, named
We sort every operation the same way, and calling is no different. The three lanes are:
- Delegate — the agent takes the call end to end and your team audits the log.
- Surface — the agent prepares or starts the call and a person finishes it.
- Hold — the call stays human, full stop.
The lanes are not maturity levels and Hold is not a waiting room. Some call types belong in Hold permanently, for reasons that have nothing to do with how good the model gets. Naming them is what turns a vague "route complex calls to humans" into something an operations team can actually run.
Sorting a real call center
Here is the sort for a reminders, service and collections operation. Yours will differ — that is the point of doing the exercise rather than adopting someone else's table.
| Call type | Lane | Why it sits there |
|---|---|---|
| Routine payment reminder, account a few days past due | Delegate | Short, scripted, bounded outcome. Every call logged and audited |
| Inbound triage — who is calling, about what, where it goes | Delegate | The agent understands the reason and routes with context instead of a menu tree |
| A call in a language and accent the model handles well | Delegate | Confirmed by evals before the language ships, not assumed from a demo |
| Order or delivery status from the system of record | Delegate | A retrieval problem, not a judgment problem |
| Appointment confirmation, cancellation, in-policy reschedule | Delegate | Policy-constrained and verifiable against the record |
| Routine account verification and profile maintenance | Delegate | Bounded once identity checks and policy limits are explicit |
| The moment a customer pushes back, hesitates, or gets emotional | Surface | Warm handoff mid-call, with full context. A person finishes the conversation |
| A payment arrangement inside a pre-approved range | Surface | The agent can assemble the options; committing to one is a policy act |
| A complaint that will end in a credit or a goodwill gesture | Surface | Precedent matters, and precedent is a human call |
| First contact on a multi-system failure with unclear cause | Surface | The agent collects symptoms and hands over a brief, not a transcript |
| A collections conversation past the sensitivity threshold | Hold | Disputed amounts, hardship, legal exposure. A machine should not make this call |
| A retention call to a high-value account | Hold | The relationship is the point. The agent can brief the caller; it does not dial |
| Anything involving a regulator, a lawyer, or a journalist | Hold | The downside is not operational, and the operation should not be deciding it |
Two things about that table are worth more than the rows themselves. First, the reason column is load-bearing — a lane assignment without a stated reason cannot be argued with later, which means it cannot be revised responsibly either. Second, the split does not follow difficulty. A collections conversation and a payment reminder are the same technical problem. They are in different lanes because one ends in a decision about a person under financial pressure and the other does not.
The two ways vendors get this wrong
Almost every voice AI pitch fails in one of two directions, and the direction tells you what you are buying.
Everything in Delegate. The vendor's answer to "which calls" is "all of them, and the containment number proves it." This builds a system you cannot trust, because trust is not a property of the model — it is a property of knowing where the boundary is. Klarna is the public example most people know: in 2024 the company said its AI assistant was handling two-thirds of customer service chats, and by 2025 its CEO was saying publicly that they had gone too far and needed more human coverage on quality-sensitive cases. That is not a story about a model underperforming. It is a story about a perimeter drawn by ambition instead of by decision risk.
Everything in Hold. The opposite failure is quieter and more common in large enterprises. The agent summarises, drafts, suggests, dashboards — and a person still makes and executes every decision. Nothing is automated, nothing is riskier than before, and nothing gets faster. This is an expensive dashboard. It survives procurement precisely because it changes nothing, and it is usually defended with the word "augmentation."
The useful vendor is the one who tells you which of your call types belongs in Hold and does not flinch when you ask why. If everything you describe gets waved into the first lane, you are talking to a demo.
What has to exist before a call type moves to Delegate
Moving a call type into the first lane is not a configuration change. It is a claim that the operation can catch the agent being wrong. Six things have to exist before that claim is true:
- A per-call audit record. Not a recording — a record. What the agent understood, what action it took, which rule permitted that action, and what the outcome was. If a call cannot be reconstructed six months later without listening to it, the lane is wrong.
- A bounded action set. The agent can do a listed set of things and nothing else. Guardrails are not prompt instructions; they are the absence of a code path to the thing you do not want done.
- Evals for the language and accent, passed before launch. Including code-switching and numbers read aloud, which is where most multilingual deployments actually break. No language ships on a demo.
- A stop condition the agent cannot talk its way past. Sentiment deterioration, a repeat contact within a short window, an identity mismatch, a request that falls outside the action set. Each of these ends the agent's turn rather than prompting another one.
- A mid-call handoff that transfers the conversation, not the transcript. The person picking up inherits the context and the reason for escalation. If they have to reread while the customer waits, you have built an IVR with better diction.
- A measured override rate from the Surface lane. You should know how often a person changed the agent's proposed action before you stop asking a person. Promoting a call type without that number is guessing.
The regulatory boundary sits underneath all six. The FCC's 2024 ruling that AI-generated voices in robocalls fall under the TCPA as artificial voices is a reminder that outbound voice is not only a product decision. Consent and disclosure design belong in the same conversation as lane assignment.
Lanes are recalibrated, not set once
The first sort is a hypothesis. The system that works is the one that revisits it on evidence.
Graduation. A call type moves from Surface to Delegate when the override rate has been low enough for long enough on a real sample — not a pilot's best week. The threshold is yours to set, but it must be set in advance and written down, because a threshold agreed after the fact is a rationalisation.
Demotion. The reverse must be as easy, and it is the part most programmes never build. A call type that starts producing complaints, or whose override rate climbs after a product change, goes back to Surface. If demotion requires a steering committee, it will not happen and the operation will absorb the errors instead.
The ones that never move. Some call types stay in Hold permanently, and the honesty here matters. The hardship conversation does not graduate when the model improves, because the reason it is held is not that the agent would handle it badly. It is that a person in financial distress is owed a person. Writing that reason down, once, prevents the annual conversation where someone reasonably asks why it is still manual.
Four numbers that tell you the sort is wrong
Containment rate will not tell you this, because containment goes up whether the sort is right or wrong. Four measures will:
Override rate by call type, in the Surface lane. How often a person changed the agent's proposed action rather than accepting it. Consistently near zero means the call type is ready to graduate. Consistently high means it was never a Surface candidate and belongs in Hold.
Repeat contact within a short window on Delegate calls. A customer calling back about the same thing is the cheapest available signal that a call was contained rather than resolved. If this number is rising while containment holds steady, the agent is closing calls that are not finished.
Complaint rate on Delegate calls, measured separately. Blended into the operation's overall complaint rate it is invisible. Split out, it is the earliest warning that a call type has drifted out of its lane — usually because a product or policy change moved the boundary and nobody moved the lane with it.
Attempts against Hold. How often the agent was routed a call that should never have reached it. This one is almost never instrumented, and it is the most diagnostic of the four, because it measures the routing rather than the agent. A rising count means the sort exists on paper and not in the call flow.
None of these requires a new platform. They require the per-call audit record to exist and to carry the call type, which is why that record is the first requirement rather than a reporting nicety.
Where this lands in an engagement
The artefact from all of this is a decision ledger: your call types, each with a lane, the reason it sits there, and what would have to be true for it to move. Not an assessment and not a roadmap — a list of decisions with an owner for each.
Building one is the first stage of how we work, a short, paid diagnostic. It is paid because it is the work, not a qualification call, and because a free version produces a document designed to sell the next stage. After the diagnostic comes one loop running in production, then scale, then your team owning it. If a calling operation is where you are starting, our calling page carries a worked ledger for a reminders-and-collections operation alongside what production voice actually requires: latency low enough that people do not talk over the agent, evals per language, mid-call handoff, and audio engineered for traffic and shop floors rather than quiet rooms.
We run a calling operation with a Series B fintech, where the work sits across Delegate and Hold. That combination — meaningful autonomy on routine contact, a firm perimeter around the sensitive conversations — is what a calibrated calling operation looks like in practice. We keep client detail off the public site; what we publish is who we work with, which operation, and where the decisions sit.
For the same calibration applied to text channels, see AI customer support for SaaS and the way the returns actually accrue in support AI ROI. For the underlying interaction layer, what conversational AI is; for how agents differ from the previous generation of automation, AI agents vs chatbots.
Start here
If you are evaluating voice agents this quarter, do this before you take another demo:
- List your call types. Twelve to twenty is usually the whole operation. Volume matters less than the decision at the end of each one.
- Assign a lane and write the reason. The reason is the artefact. A lane without one cannot be defended or revised.
- Pick one Delegate candidate and check the six requirements against it. Most operations discover the audit record is the missing piece, not the model.
- Instrument the Surface lane first. Its override rate is the evidence that lets anything move later.
- Agree the demotion rule before launch. The programme that cannot move a call type backwards will not move one forwards safely either.
Frequently Asked Questions
Which call center calls should an AI voice agent handle end to end?
Calls where the outcome is bounded and verifiable against a record: routine payment reminders, inbound triage and routing, order and delivery status, in-policy appointment changes, and routine account verification — and only in languages and accents that have passed evals. The shared property is that the decision at the end of the call is one a rule can express, so the log can be audited afterwards rather than trusted at the time.
Which calls should never be automated?
Calls where the decision is about a person rather than a transaction. Collections past a sensitivity threshold, hardship and dispute conversations, retention calls to high-value accounts, and anything involving a regulator, a lawyer or the press. These sit in Hold permanently, and the reason is not model capability — it is that the downside is relational and legal rather than operational, so a person should own it.
How do you decide when a call type is ready for full autonomy?
By measuring how often a person changed the agent's proposed action while the call type sat in the Surface lane, against a threshold agreed in advance. Alongside that, six things have to be in place: a per-call audit record, a bounded action set, evals passed for the language, stop conditions the agent cannot talk past, a mid-call handoff that transfers context, and a demotion rule that can move the call type back.
What is a decision ledger?
A list of an operation's decisions — here, its call types — each with a lane, the reason it sits in that lane, and what would have to be true for it to move. It is the deliverable from the paid diagnostic that starts an engagement, and it is what makes autonomy arguable rather than assumed.
Need help with AI implementation?
We build production AI systems that actually ship. Not demos, not POCs—real systems that run your business.
Get in Touch