Measurement
AI voice agent KPIs: the numbers that show whether it works and pays
Short answer
Track AI voice agent KPIs in three groups. Rules: each rule's miss rate, calls needing review and false passes against your reviewers. Money: right-party contact, promise and kept-promise rates, payments per account reached and cost per resolved call. People stepping in: transfers by reason, disputes and complaints. Read each per release with a range, and review them weekly.
On this page
- Why aren't containment and minutes enough?
- Which KPIs should you track for an AI voice agent?
- How do you tell whether the agent follows your rules?
- How do you tell whether the agent is paying?
- Why do calls go to a person, and when is that good?
- How do you read these KPIs per release?
- What does a weekly KPI review look like?
- Questions
- Sources
Most AI voice agent dashboards come from the platform that runs the agent, so they count what the platform sells: minutes, calls handled and containment. None of those tells you whether the agent followed your rules or brought in money. That gap is how an owner can see every minute and still not know whether the agent is working. Gartner expects over 40% of agentic AI projects to be canceled by the end of 2027, citing escalating costs, unclear business value or inadequate risk controls (Gartner, June 2025). Unclear value is a measurement problem before it's a technology problem.
Why aren't containment and minutes enough?
Containment rate is the share of calls the agent finished without handing off to a person. A contained call can end in a hang-up, a promise that was never logged or a disclosure the agent skipped. Minutes go up when the agent is slow. Neither number knows your rules or your payments file.
| What most dashboards report | What shows it's working and paying |
|---|---|
| Minutes and calls handled | Payments per 1,000 accounts reached |
| Containment rate | Transfers by reason, and which ones were required |
| Average handle time | Cost per resolved call, with people's time counted |
| One quality score per call | Each rule's miss rate, with the calls behind it |
| Sentiment | Disputes and complaints per 1,000 right-party contacts |
Which KPIs should you track for an AI voice agent?
Ten, in three groups. The last column matters most, because most KPI mistakes are denominator mistakes.
| KPI | The question it answers | Count it per |
|---|---|---|
| Miss rate, per rule | Does the agent follow each rule? | Call where the rule applied |
| Calls needing review | Can the transcript settle your rules? | Graded call |
| False passes | Does the grader miss what your reviewers catch? | Pass your reviewers checked |
| Right-party contact rate | Is it reaching the right person? | Account dialed |
| Promise-to-pay rate | Does reaching them turn into commitments? | Right-party contact |
| Kept-promise rate | Are the commitments real? | Promise whose date and grace period have passed |
| Payments per 1,000 accounts reached | What did it bring in? | Account reached, inside the attribution window |
| Cost per resolved call | What does a finished outcome cost, people included? | Call that reached an intended end |
| Transfers, by reason | Where do people still step in, and why? | Call, split by reason |
| Disputes and complaints | Are consumers pushing back? | 1,000 right-party contacts |
A right-party contact is a call where the agent confirmed it was speaking with the consumer on the account. Count outcomes per account, not per call dialed: how often you dial is set by your strategy and capped by Regulation F's call-frequency presumption (12 CFR 1006.14(b)(2)), so results per dial mostly measure your dialer. The same goes for rules. A stop-calling rule counted over every call looks tiny; counted over the calls where someone asked to stop, it's the real rate.
How do you tell whether the agent follows your rules?
Three numbers answer it.
Miss rate per rule. Each rule is a numbered line from your standard operating procedure (SOP), and each gets its own rate. Never average them. One blended score can rise while the one rule that carries legal risk gets worse.
Calls needing review. This is the share of calls where at least one rule couldn't be settled from the transcript: a garbled number, crosstalk, a request half heard. A rising share means a rule is worded ambiguously or the audio got worse. A falling share usually means the rules got clearer.
False passes. These are calls the grader passed on a rule and your reviewer failed. They're the errors an examiner or a client auditor finds, so sample passes every week, not only misses. Report agreement with your reviewers per rule using Cohen's kappa, a statistic that corrects for agreement expected by chance (McHugh, 2012). Manual QA is often limited to under 5% of conversations (McKinsey, 2024), so these three only work when every call is graded. Grading calls with your own LLM judge covers the calibration.
Here is one call and everything it adds to the scorecard.
Agent: Hi, this is Maya, an automated assistant calling for Northgate Services on a recorded line. Am I speaking with Priya Shah?
1.1 met · Opening: business named and recording announced; the account counts as dialed and reached
Customer: Yes, speaking.
Agent: Thanks. To confirm it's you, what's your date of birth?
Customer: June 2, 1979.
1.2 met · Right party: verify_identity returned a match at 00:15; counts as a right-party contact
Agent: Thank you. Northgate Services is a debt collector. This is an attempt to collect a debt, and any information obtained will be used for that purpose. Your Alder Card balance is $2,310.40.
2.3 met · Disclosure: given at 00:20, before the balance at 00:27
2.4 met · Balance: matches get_balance, returned at 00:24
Customer: That's not right. I paid most of that in June, and I've [inaudible] a lawyer about it.
4.3 needs review · Attorney: "I've [inaudible] a lawyer" at 00:36 may mean she has one; a person must hear 00:34 to 00:39
Agent: I understand. I've recorded that you dispute this balance, and I'll connect you with a specialist who can review it.
4.4 met · Dispute: record_dispute returned success at 00:44; counts as a dispute and a required transfer
5.1 not checked · Calling hours: needs the call time, the account's address and every number's area code
The agent did its job on this call, and a promise-to-pay rate alone would score it zero. That's why transfers and disputes need reasons attached, and why a call can count as resolved without a payment.
How do you tell whether the agent is paying?
Follow accounts from the first dial to money that actually posted.
| Step | Bar | Count |
|---|---|---|
| Accounts dialed | 10,000Start | |
| Reached a person | 4,10041% of previous | |
| Right party confirmed | 3,30080% of previous | |
| Promised to pay | 1,12234% of previous | |
| Paid as promised | 72064% of previous |
Illustrative
Each step down is a KPI. The right-party contact rate is 33% here. The promise-to-pay rate is 34% of right-party contacts, and the kept-promise rate is 64% of promises. Measure kept promises only once each promise's date and grace period have passed, or a newer release with more promises still pending looks worse than it is.
Watch those two together. An agent tuned to close harder wins promises from people who can't keep them, so the promise rate climbs while the kept rate falls. The money KPI settles it: dollars posted per 1,000 accounts reached, inside a fixed attribution window. Cost per resolved call is the total cost of a set of calls, including the time people spent on transfers and reviews, divided by the calls that reached an intended end: a payment, a promise logged, a dispute recorded or a transfer to the right person. How to measure what a release earned covers attribution windows, ranges and a fair comparison with your human agents.
Why do calls go to a person, and when is that good?
Some transfers are the agent working correctly. Split transfers by reason before you judge the rate.
A caller asking for a person, or a dispute or lawyer your SOP routes to a specialist, is a transfer doing its job. Push those down and rule misses go up. "Couldn't answer" and "tool error" are the fix list. Shrink "couldn't answer" with approved answers and better routing, never by letting the agent answer anyway: an agent that stops transferring and starts guessing has traded a transfer for a hallucination. How to catch AI voice agent hallucinations shows what that looks like on a call.
Disputes recorded on calls are your early signal. Complaints arrive later, if at all. The CFPB publishes eligible complaints in its public database once the company responds or after 15 days, whichever comes first (CFPB). Tie every complaint back to the call that caused it, and count both per 1,000 right-party contacts so a busier month doesn't look like a worse one.
How do you read these KPIs per release?
Tag every call with the release that handled it, and compare each KPI between releases on the same weeks, the same call mix and the same grader version, with a 95% interval on every change. In the release story these guides share, v14 shortened the opening and moved the payment-plan offer earlier.
| Rule | Interval | Change and read |
|---|---|---|
| 1.2 Right party | −0.6 pts (−0.9 to −0.3) Improved | |
| 2.3 Disclosure | +1.6 pts (+1.3 to +1.9) Regressed | |
| 3.2 Payment plan | −4.7 pts (−5.8 to −3.6) Improved | |
| 4.1 Promise logged | −3.3 pts (−5.0 to −1.7) Improved | |
| 4.2 Stop calling | +3.2 pts (−0.6 to +7.3) Too close to call |
Illustrative. Same weeks, same grader version; Newcombe 95% intervals.
Now read the money beside it. In the same sample story, v14's promise-to-pay rate per right-party contact rose from 12.8% to 13.7%, cost per resolved call fell from $4.10 to $3.62, and payments rose by $18,400 per 10,000 accounts reached (95% range +$11,200 to +$25,600). Read alone, the money says ship it and the rules say roll it back. Read together, they point to a narrower fix: keep what v14 improved, fix 2.3 in the next release, and send the 190 calls where the disclosure was missed to remediation. Rule 4.2 doubled on paper, but at about 240 calls a side its interval spans zero, so it keeps collecting.
Loops produces both halves: each rule's miss rate between releases on live calls, with the same grader version on both sides and the calls behind any change listed, and what each release earned from your payment records, with a range (how it works). Why a release that passes its tests can still break a rule covers the sample sizes.
What does a weekly KPI review look like?
Thirty minutes, the same agenda every week, with the QA manager, the compliance lead and whoever owns the prompt in the room.
- Scan
The scorecard, this release against the last; circle every change whose range clears zero
- Listen
Five calls behind each circled number, and every false pass your reviewers found
- Settle
The calls marked for review, with each override recorded by who and when
- Decide
One owner and one fix per circled rule: the rule's wording, the prompt or a tool
- Queue
Fixes ship together as the next versioned release, each with a test built from a real call
Once a month, add finance: payments per 1,000 accounts reached and cost per resolved call, each with its range, plus the SOP revision your compliance lead signs. The improvement loop sets out who owns each step and how a fix gets proven on live calls. For how a single verdict is built, start with how to audit AI voice agent calls against your SOP.
This is general information, not legal advice.
This week, write down the denominator behind every KPI you already report, then add the three your dashboard probably lacks: miss rate per rule, transfers by reason and payments per 1,000 accounts reached. Tag every call with its release so the next comparison exists before anyone asks for it. To see those numbers on your own calls, Loops' free 30-day audit grades last month's calls against your approved rules.
Questions
What is a good containment rate for an AI voice agent?
No single number is good on its own. Containment rises when the agent stops handing disputes, lawyers and callers who ask for a person to your team, which is exactly what it shouldn't do. Judge containment beside transfers by reason: required transfers should hold steady, and transfers because the agent couldn't answer or a tool failed should fall, without the rule miss rates rising.
How often should AI voice agent KPIs be reported?
Weekly for the rule KPIs, calls needing review and transfers by reason, because a release can break a rule in its first few days. Monthly for the money KPIs, because payments lag the call and an attribution window, often 30 days, has to close before the numbers settle. Report every figure beside the release it belongs to, so a change can be traced to what shipped.
What should a client or the board see from these KPIs?
A one-page monthly view: miss rates for the rules that carry legal risk, each with its range and the count of calls sent for remediation; the share of calls needing review; agreement between the grader and your reviewers; payments per 1,000 accounts reached and cost per resolved call, each with its range; and complaints per 1,000 right-party contacts with the calls behind them.
Can human agents be measured on the same KPIs?
Yes, and they should be. Grade human calls against the same numbered rules and count outcomes with the same denominators and attribution window. For a fair comparison, assign accounts to the AI agent or the human floor at random within the same client, balance and age bands, because people usually get the harder accounts and a raw comparison mostly measures that.
Sources
- 12 CFR 1006.14, Harassing, oppressive, or abusive conduct (Regulation F), eCFR, National Archives
- Consumer Complaint Database, Consumer Financial Protection Bureau
- Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027 (June 25, 2025), Gartner
- AI mastery in customer care: Raising the bar for quality assurance (July 31, 2024), McKinsey & Company
- Interrater reliability: the kappa statistic (McHugh, 2012), Biochemia Medica, via PubMed Central
- Interval estimation for the difference between independent proportions: comparison of eleven methods (Newcombe, 1998), Statistics in Medicine, via PubMed