1. Home
  2. Blog

Measurement

How to measure what an AI voice agent release earned

Short answer

Measure AI voice agent ROI from payment records, not minutes or containment. Credit each payment to the release that handled the account's last call before it posted, within a fixed window such as 30 days. Compare releases on the same weeks, clients, and balance and age bands, put a 95 percent range on every difference, and count human transfer and review time in cost per resolved call.

On this page
  1. How do you credit a payment to a release?
  2. How do you compare two releases like with like?
  3. When is a difference too small to call?
  4. What does a release comparison look like in numbers?
  5. Why track promise-to-pay and kept-promise rates separately?
  6. What does a resolved call cost?
  7. How do you compare the AI agent with human agents fairly?
  8. Which pitfalls make the numbers wrong?
  9. Questions
  10. Sources

Minutes handled and containment describe how much work an AI voice agent did. Containment rate is the share of calls the agent finished without handing off to a person, and a contained call can end in a hang-up with no money moving. Finance needs a different number: how much more, or less, was paid on the accounts this release spoke with than on comparable accounts the last release spoke with. That number only comes from joining calls to payment records.

Use the account reached as the unit, not the call dialed. A right-party contact is a call where the agent confirmed it was speaking with the consumer on the account. How often you dial is set by your strategy and limited by the call-frequency presumption in Regulation F, the CFPB's debt collection rule (12 CFR 1006.14(b)(2)), so results per call placed mostly measure your dialer. Four measures follow from right-party contacts:

  • Promise-to-pay rate: promises logged per right-party contact.
  • Kept-promise rate: the share of promises, with dates already past, paid in at least the promised amount within a few days of the promised date.
  • Dollars posted per 1,000 accounts reached, inside the attribution window. Count only amounts the agreement creating the debt or the law authorizes: for a debt collector, a fee outside those violates 12 CFR 1006.22(b), and it's a refund waiting to happen, not revenue.
  • Cost per resolved call, covered below.

How do you credit a payment to a release?

Pick one attribution rule, write it down and use it for every comparison. An attribution window is the number of days after a call during which a payment can still be credited to that call. One workable rule is last touch: a payment counts toward the release that handled the account's last call before the payment posted, as long as that call falls inside the window.

Run two checks beside it. A shorter window, such as 14 days, tests whether the result leans on slow payments the call may not have caused. Any-touch attribution, which credits every release that spoke with the account inside the window, shows whether a rollout that put two releases on the same accounts is muddying the picture. If the ranking of two releases flips between methods, the difference isn't solid enough to report.

Payments with no call inside the window stay unattributed. Some are self-cures, accounts that would have paid without any call. Others followed a letter, a text or a portal visit. Last-touch attribution still credits some of those to a call, so every release's total looks better than it is. The difference between two releases on comparable accounts is far more trustworthy than either total, because both sides get the same letters and the same self-cures.

How do you compare two releases like with like?

The cleanest design runs both releases at the same time on a random split of accounts, for example by a hash of the account ID, so both see the same weeks, clients and paydays. When releases ran one after the other, match instead: compare within the same client, balance band and age band, then weight each band's difference to a common mix.

Age alone moves results a long way. The FTC's 2013 study of the debt buying industry estimated, from its regression model, that buyers paid about 7.9 cents per dollar for credit card debt under three years old, 3.1 cents for debt three to six years old and 2.2 cents for debt six to 15 years old; debt past 15 years sold for virtually nothing. A release that happened to work fresher placements looks better for reasons that have nothing to do with the release.

Pooling can even reverse the answer. In Simpson's paradox, a relationship that holds inside every group flips when the groups are combined: release B can beat release A in every age band and still lose overall if it was handed more old accounts. Compliance comparisons between releases need the same matching; see why a release that passes its tests can still break a rule.

When is a difference too small to call?

When its range includes zero, or is so wide that "it cost us money" is still a plausible answer. Report every difference with a 95 percent range. For rates such as promise-to-pay or payment rate, the standard two-proportion method works at collections volumes. For dollars, resample accounts (a bootstrap), because payment amounts are skewed: a handful of settlements in full can move an average more than a thousand small installments.

Fix the length of the comparison before you look. Checking every morning and stopping on the first good day finds differences that aren't there. A range that crosses zero is still an answer. It says the release hasn't earned or cost enough to see yet, so run it to the length you set, or decide on other grounds, such as its compliance record.

What does a release comparison look like in numbers?

Illustrative example, not real data. Two releases ran on a random 50/50 split of one agency's accounts, in a comparison planned for eight weeks, with week four as an interim look that could stop early only for harm. At week four each had reached 12,000 accounts. Payments are credited to the release that handled the account's last call before posting, within 30 days.

Measure (illustrative, week four)v7v8Difference (95% range)Call it?
Promise-to-pay rate15.0%17.5%+2.5 pts (+1.6 to +3.4)Yes
Kept-promise rate50.0%48.0%-2.0 pts (-5.2 to +1.2)No
Accounts paying within 30 days8.0%8.9%+0.9 pts (+0.2 to +1.6)Yes
Average payment$310$305-$5Not tested
Dollars posted per 1,000 accounts reached$24,800$27,145+$2,345 (-$700 to +$5,300)Not yet

Read it row by row. v8 wins more promises, and more of its accounts paid. Its kept rate slipped, but within noise. Its average payment was $5 lower, but that compares different sets of payers, so it isn't a like-for-like test. Dollars per 1,000 accounts point up, yet the range runs from a small loss to a large gain, because payment sizes vary so much. The accurate summary at the interim look: v8 gets more people to pay, and it can't yet be said to collect more money. If the gap holds to week eight, twice the accounts shrink the range by about 1.4 times, to roughly +$200 to +$4,450, and you can call it.

Why track promise-to-pay and kept-promise rates separately?

Because a release can raise one by lowering the other. An agent tuned to close harder wins promises from people who can't keep them, so the promise-to-pay rate climbs while payments don't. The kept-promise rate is the check. Fix the grace period (three days, say), and measure it only on promises whose date and grace period have passed. Otherwise the newer release, with more promises still pending, is judged on an incomplete count.

Pressure has a compliance cost as well. An agent that wins promises by implying consequences the collector can't lawfully take or doesn't intend to take is making the representations 12 CFR 1006.18 prohibits (for a lender collecting its own debts, the deceptive or abusive practices 12 U.S.C. 5536(a)(1)(B) prohibits), and revenue earned that way returns as remediation. Read a release's rule miss rates beside its revenue: how to audit AI voice agent calls against your standard operating procedure (SOP) covers the grading, and FDCPA and Regulation F rules for AI collection calls covers the rules. And a promise only counts if it reached your system, so confirm the tool call that logged it returned success (rule 4.1 in the example collections SOP these guides share).

What does a resolved call cost?

Cost per resolved call is the total cost of a set of calls divided by the calls that reached an intended end: a payment taken, a promise logged, a dispute or cease request recorded, or a transfer to the right person. It beats cost per minute, because a cheap minute that resolves nothing is expensive.

Count everything the release causes. Illustrative: 10,000 AI calls cost $2,600 in platform, telephony and model charges. Of those, 400 transferred to people who spent six minutes each at a loaded $0.75 a minute ($1,800), and reviewers spent four minutes on each of 150 flagged calls ($450). With 1,200 resolved calls, that's $4,850 divided by 1,200, or $4.04 per resolved call. Leave out the people and it reads $2.17, which flatters the AI agent by nearly half.

How do you compare the AI agent with human agents fairly?

Your human floor doesn't get the same accounts. People handle callbacks, disputes, hardship cases and the largest balances, so a raw comparison mostly measures the work each side was handed. For a fair read:

  1. Assign accounts at random to AI-first or human-first for a set period, within the same client, balance and age bands.
  2. Credit outcomes to the assigned channel, including AI calls that transferred to a person. This is the intention-to-treat principle from clinical trials: judge the strategy you chose, not only the calls that went smoothly.
  3. Use the same attribution window, measures and cost accounting on both sides.
  4. Grade both sides on the same SOP rules, so a revenue gain bought with compliance misses shows up next to the gain.

Which pitfalls make the numbers wrong?

  • Seasonality. Tax refunds arrive in a rush: by March 13, 2026, the IRS had issued 50.4 million refunds worth $182.6 billion (IRS), and by May 8 it had issued 99.1 million worth $324.8 billion (IRS). A release shipped in February and one shipped in June faced different money.
  • Portfolio mix. A new client, a different balance mix or a fresh batch of placements changes results for reasons unrelated to the release. Compare within bands.
  • Placement changes. Clients recall accounts and add new ones mid-test. Freeze the account set for the comparison, or drop recalled accounts from both sides.
  • Self-cures. Some accounts would pay with or without a call. Where your strategy allows it, a small holdout with no AI call measures that baseline, so you can tell what the calls added.
  • Changes that ship with the release. A new dialing schedule, settlement offer or letter in the same week gets credited to the release unless you separate them.

This is general information, not legal advice.

Before your next release ships, write down the attribution rule and window, tag every call with the agent version, and set up a daily payments file (account ID, amount, date posted), so the comparison exists when someone asks what the release earned. Loops joins calls to that file and reports each release's earnings with a range, next to the rule changes that explain them (how the payments join works). You can start with a free 30-day audit.

Questions

What attribution window should a collections team use?

Thirty days after the last call is a common starting point, long enough to catch a payment made on the next payday or on a promised date later in the month. Whatever you choose, fix it before the comparison starts and rerun with a shorter window, such as 14 days. If two releases swap places when the window changes, the difference between them isn't solid enough to report.

How many accounts do you need to compare two AI agent releases?

It depends on the base rate and the size of difference you care about. In the illustrative example above, 12,000 accounts per release put a range of about plus or minus 0.7 points on the difference between two payment rates near 8 percent, enough to see the 0.9 point gain, though a gain that size would clear the bar only about 7 times in 10. Dollar figures need far more accounts, because payment sizes vary so widely.

Should payment plans count at full value or as installments arrive?

Count dollars as they post inside the attribution window, and report scheduled plan value separately. A plan set up on a call is a promise until the money arrives, and counting future installments as earned rewards a release that sets up plans consumers later abandon. The kept rate on first installments is the early signal of plan quality.

Can you measure AI voice agent ROI without a random split between releases?

Yes, with more care and less certainty. Compare within the same client, balance band and age band, weight each band's difference to a common mix, and use a group that didn't change over the same weeks, such as your human floor, as a check on seasonality. Report the result as weaker evidence than a randomized split would give.

Sources

  1. 12 CFR 1006.14, Harassing, oppressive, or abusive conduct (Regulation F), eCFR, National Archives
  2. 12 CFR 1006.18, False, deceptive, or misleading representations or means (Regulation F), eCFR, National Archives
  3. 12 CFR 1006.22, Unfair or unconscionable means (Regulation F), eCFR, National Archives
  4. 12 U.S.C. 5536, Prohibited acts, Legal Information Institute, Cornell Law School
  5. The Structure and Practices of the Debt Buying Industry (January 2013), Federal Trade Commission
  6. Filing season statistics for week ending March 13, 2026, Internal Revenue Service
  7. Filing season statistics for week ending May 8, 2026, Internal Revenue Service
  8. NIST/SEMATECH e-Handbook of Statistical Methods, 7.3.3: Comparing two proportions, National Institute of Standards and Technology
  9. Simpson's Paradox, Stanford Encyclopedia of Philosophy
  10. Intention-to-treat concept: A review (Gupta, 2011), Perspectives in Clinical Research, via PubMed Central