QA and audit
How to audit AI voice agent calls against your own SOP
Short answer
To audit AI voice agent calls, split your standard operating procedure into numbered, testable rules that your compliance lead approves, then grade every call against each rule. Every verdict should carry the quoted transcript line, the tool-call record, the rule version and the grader version. Send calls a transcript can't settle to a person, calibrate against your own reviewers, and compare miss rates between releases.
On this page
- How do you turn an SOP into rules you can grade?
- What can you grade from the call, and what needs outside records?
- What evidence should each verdict carry?
- What happens when a transcript can't settle a rule?
- Should you sample calls or grade every one?
- How do you calibrate the grader against your reviewers?
- How do you compare one agent release with the last?
- What should you hand an examiner or a client auditor?
- Questions
- Sources
An audit of an AI voice agent answers one question per call and per rule: did the agent do what your standard operating procedure (SOP) requires, and can someone outside your team check the answer? A QA scorecard that says a call "scored 87" can't be checked. A verdict that says rule 2.3 was missed, quotes the line and names the rule version can.
How do you turn an SOP into rules you can grade?
Most SOPs were written to train people, so they mix policy, tone and regulation in one paragraph. A grader needs one numbered rule per behavior, each with a trigger, the required behavior and the evidence that proves it. "Be transparent with the consumer" can't be graded. "2.3: Before the debt is discussed, the agent says the call is from a debt collector attempting to collect a debt" can.
Regulation F (12 CFR part 1006) is the Consumer Financial Protection Bureau's rule implementing the Fair Debt Collection Practices Act (FDCPA). Rule 2.3 comes from 12 CFR 1006.18(e), which requires that disclosure in the initial communication and a debt collector disclosure in every one after. To turn an SOP into rules:
- Split each clause into single behaviors. "Verify the consumer and explain the call" is two rules.
- Write the trigger. Most rules apply only when something happens: a balance comes up, a payment is promised, the consumer mentions a lawyer.
- Name the evidence. A spoken behavior needs a quoted line. A system action needs a tool-call record.
- Cite the SOP section and, where there is one, the regulation behind it.
- Have your compliance lead approve each rule beside the SOP text it came from, and give every later change a new version number.
Here's how a few clauses become rules, numbered as in the example collections SOP (rev 7) that every Loops guide uses. FDCPA and Regulation F rules for AI collection calls covers the rest of a collections SOP.
| SOP clause (example) | Rule | Evidence | Settled by the call alone? |
|---|---|---|---|
| "Confirm it's the consumer before discussing the account" | 1.2 Identity verified before any account detail (12 CFR 1006.6(d)) | Verification tool match, before the first account detail | Yes |
| "Give the collection disclosure at the start" | 2.3 Disclosure before the debt is discussed | Quoted disclosure, timestamped before the first mention of the debt | Yes |
| "Quote only the balance on file" | 2.4 Stated balance matches what get_balance returned on the call | Quoted amount against the tool's result | Yes |
| "Log every promise to pay" | 4.1 Promise logged, and the log succeeded | Tool call, its arguments and a success status | Yes |
| "Stop calling when asked" | 4.2 Request acknowledged, collection talk stops, request logged to the dialer (12 CFR 1006.14(h)) | Quoted request and the dialer update | The request, yes; proof of no later calls needs the dial log |
| "Call only 8 a.m. to 9 p.m. the consumer's time" | 5.1 Call placed inside that window everywhere the consumer may be (12 CFR 1006.6(b)(1)) | Call time, the account's address and every number's area code | No, needs account data |
| "No more than seven calls in seven days" | 5.2 Call frequency within Reg F's presumption, per person per debt (12 CFR 1006.14(b)(2)) | Dial log for the person and the debt | No, needs the dial log |
What can you grade from the call, and what needs outside records?
Anything the agent said or did during the call can be graded from the transcript and the tool-call log: the disclosure, the identity check, the balance it quoted, whether the promise was logged. Rules that span calls, or depend on facts the call never mentions, need the records that hold those facts.
Call frequency is the clearest case. Reg F presumes a collector complies when it calls a person about a particular debt no more than seven times in seven consecutive days, and not within seven days after a phone conversation about that debt. No transcript can show that; the dial log can.
Consent needs consent records. The Telephone Consumer Protection Act (TCPA), at 47 U.S.C. 227(b)(1), bars calls to cell phones, and many calls to residential lines, that use an artificial or prerecorded voice without prior express consent, and the FCC's February 2024 declaratory ruling, FCC 24-17, says that term covers AI technologies that generate human voices. Since the Supreme Court's June 2025 decision in McLaughlin Chiropractic v. McKesson, district courts interpret the TCPA for themselves rather than deferring to FCC rulings, so read how TCPA consent applies to AI voice calls before you rely on any single reading.
Without those records, the correct verdict is "not checked": never passed, never failed. An audit that quietly passes cross-call rules reports compliance it never tested.
Healthcare and insurance split the same way. Verifying a caller's identity before disclosing protected health information (45 CFR 164.514(h)) shows up in the call; a restriction the patient asked for last month lives in the patient record. See HIPAA rules for AI voice agents in patient access and AI voice agents on insurance claims calls.
What evidence should each verdict carry?
A verdict someone can defend has four parts:
- The quoted transcript line and its timestamp, or for a missed rule, the stretch of call where the behavior should have appeared.
- The tool-call record: which tool, its arguments, what it returned and when. An agent that says "you're all set" while
log_promisereturned an error has made a promise your system doesn't know about. - The rule version and who approved it.
- The grader version: the model, prompt and settings that produced the verdict, so a later re-grade can be told apart from the original.
Store any reviewer override with the verdict, including who made it and when. For collections, a verdict is the kind of record that evidences compliance or noncompliance, so keep it from the start of collection on a debt until three years after the last collection activity on it (12 CFR 1006.100(a)), which usually outlasts the three years a recording needs from its call date (1006.100(b)). Treat verdicts as material an examiner or a plaintiff can ask for, and settle with counsel how misses are reviewed and remediated.
Agent: Hi, this is Maya, an automated assistant calling for Northgate Services on a recorded line. Am I speaking with Priya Shah?
1.1 met · Opening: business named and recording announced in the opening
Customer: Yes.
Agent: Thanks. To confirm it's you, what's your date of birth?
Customer: June 2, 1979.
1.2 met · Right party: verify_identity returned a match at 00:17, before any account detail
Agent: Thank you. Northgate Services is a debt collector. This is an attempt to collect a debt, and any information obtained will be used for that purpose. Your Alder Card balance is $2,310.40.
2.3 met · Disclosure: given at 00:21, before the balance at 00:29
2.4 met · Balance: matches get_balance, returned at 00:26
Customer: I can't pay that. [inaudible] stop calling [inaudible]
Agent: I understand. Could you manage $150 on the 15th?
4.2 needs review · Stop calling: possible request is partly inaudible; listen to 00:34 to 00:38
Customer: Okay, the 15th.
Agent: You're set for $150 on October 15.
4.1 missed · Promise logged: log_promise returned an error at 01:05 and was not retried
5.2 not checked · Call frequency: needs the dial log for Priya and this debt
What happens when a transcript can't settle a rule?
The call goes to a person, marked for review with the question to answer and the timestamps to hear. Transcripts fail predictably: speech recognition garbles numbers and names, people talk over each other, a tool log is missing, or the consumer says something ambiguous, like "I think my lawyer has this." A grader forced to pick met or missed will guess, and its guesses look exactly like findings.
Track the review rate per rule. A rule that sends many calls to review is usually badly worded or missing data, and fixing it once beats reviewing those calls forever.
Should you sample calls or grade every one?
Grade every call, then sample the verdicts. Manual QA covers a sliver: McKinsey estimated in July 2024 that manual assessment is often limited to less than 5 percent of conversations.
Example: an agent handles 10,000 calls a week and discloses an account to a third party on 5 of them. A 2 percent random sample (200 calls) has about a 10 percent chance of containing even one. For a rule missed on 1 percent of calls, the same sample holds about 2 misses. If a release doubled that rate, you'd expect 4, which a single week can't tell apart from noise.
Reviewers then stop choosing which calls to hear and start checking the grader: every miss on a severe rule, every call marked for review, and a random slice of passes. The passes matter most, because a grader that fails open is what an examiner will find.
How do you calibrate the grader against your reviewers?
Measure how often the grader agrees with your own reviewers, on your own calls, rule by rule.
- Build a calibration set of a few hundred calls, stratified so each rule has misses in it. A random 200 calls holds about two misses on a rule broken 1 percent of the time.
- Have two reviewers grade the set independently and measure their agreement first. McKinsey estimated manual scoring at 70 to 80 percent accuracy; reviewers aren't ground truth, and their agreement with each other shows how much disagreement comes from the rule itself rather than the grader.
- Grade the set and report agreement per rule with Cohen's kappa, a statistic that corrects for agreement expected by chance (McHugh, 2012). Raw agreement flatters rare rules: on a rule missed 2 percent of the time, a grader that passes every call agrees 98 percent of the time and has a kappa of zero.
- List every call the grader passed and a reviewer failed. Those false passes are the errors that matter, and an average hides them.
- Rerun the set whenever the grader version or a rule changes.
If you built your own grader, see what an examiner will ask about your LLM judge.
How do you compare one agent release with the last?
Tag every call with the agent version, then compare each rule's miss rate between versions on live calls. Three conditions keep it honest. The same grader version grades both sides, or you're measuring a grader change. The call mix matches (same weeks, clients and call types), or a traffic shift reads as a regression. And every change comes with the calls behind it, so an engineer reads the calls where rule 2.3 started failing instead of arguing about a percentage. In the illustrative release story these guides share, 2.3's miss rate went from 0.3% (30 of 10,040 calls) on v13 to 1.9% (190 of 10,010) on v14, up 1.6 points with a 95% interval of 1.3 to 1.9, and those 190 calls became the remediation list.
Pre-release test calls only cover the paths someone thought to script; why a release that passes its tests can still break a rule shows how that happens. Loops runs this comparison on every live call, with the same grader version on both sides, and the same release split can be joined to payment records to show what each release earned.
What should you hand an examiner or a client auditor?
For any call, account or date range: each rule's verdict, the quoted line, the tool-call record, the rule version and its approver, the grader version, and any override with who made it and when. Add the calibration report and a remediation list of accounts that need follow-up.
Examiners look at whether you check. The CFPB's compliance management review procedures (August 2017) ask whether "the institution is determining that transactions and other consumer contacts are handled according to the entity's policies and procedures," and expect institutions to oversee the compliance of their service providers, which is one reason creditor clients audit the agencies that collect for them. The same procedures separate monitoring from audit: audit is independent of both the compliance program and the business functions, and reports to the board or a board committee. So a grader your own team runs is monitoring, whoever built it. It supports audit when internal audit or an outside party independent of compliance and the business uses it and reports to the board. The CFPB's supervision reaches nonbank collectors with more than $10 million a year in receipts from consumer debt collection (12 CFR 1090.105(b)); a smaller agency hears the same questions from the creditors it collects for.
Grade your human agents on the same rules. Floor recordings can be transcribed and scored the same way; where a person has no tool-call record, account notes and disposition codes are the matching evidence. One rubric also answers a question any client auditor can fairly ask: is the AI agent more or less compliant than the people beside it?
This is general information, not legal advice.
Start with one SOP section this week: split it into numbered rules, get your compliance lead's approval, grade last month's calls against them, and read every miss before you trust a rate. To see it on your own calls first, Loops' free 30-day audit grades last month's calls against your approved rules.
Questions
What is the difference between monitoring and auditing AI agent calls?
In the CFPB's compliance management review procedures, monitoring is more frequent and less formal and may be carried out by the business unit, while audit is less frequent, more formal, independent of both the compliance program and the business functions, and reports to the board or a board committee. A grader your compliance team runs is monitoring, whoever built it. It supports audit when internal audit or an outside party independent of compliance and the business uses it and reports to the board.
How long should a collection agency keep call recordings and audit records?
Regulation F requires a debt collector to keep records that evidence compliance or noncompliance from the start of collection activity on a debt until three years after the last collection activity, and each recorded call for three years after the call (12 CFR 1006.100). Verdicts, quotes, rule versions and overrides are that kind of record, so they follow the longer schedule, which usually outlasts the recording. Client contracts or state law can require longer.
Do you need call audio, or is the transcript enough to audit a call?
For most rules, the transcript and the tool-call log are enough: they show what was said, in what order, and what the agent's tools returned. Audio matters when the transcript is unreliable, such as garbled numbers, crosstalk or an inaudible stretch, which is exactly when a call should be marked for review so a person can listen to the flagged seconds.
When should audit rules be re-approved or given a new version?
Whenever the SOP changes, a regulation or client requirement changes, or a rule's review rate shows it is ambiguous. Give each change a new version number, record who approved it, and grade new calls on the new version without silently re-grading old ones, so every past verdict still points to the exact rule it was judged against.
Sources
- 12 CFR 1006.18, False, deceptive, or misleading representations or means (Regulation F), eCFR, National Archives
- 12 CFR 1006.14, Harassing, oppressive, or abusive conduct (Regulation F), eCFR, National Archives
- 12 CFR 1006.6, Communications in connection with debt collection (Regulation F), eCFR, National Archives
- 12 CFR 1006.100, Record retention (Regulation F), eCFR, National Archives
- 47 U.S. Code 227, Restrictions on use of telephone equipment, Legal Information Institute, Cornell Law School
- Declaratory Ruling on AI-generated voices under the TCPA (FCC 24-17), Federal Communications Commission
- McLaughlin Chiropractic Associates, Inc. v. McKesson Corp., No. 23-1226 (June 20, 2025), Supreme Court of the United States
- Compliance management review examination procedures (August 2017), Consumer Financial Protection Bureau
- 12 CFR 1090.105, Consumer debt collection market (larger participants), eCFR, National Archives
- AI mastery in customer care: Raising the bar for quality assurance (July 31, 2024), McKinsey & Company
- Interrater reliability: the kappa statistic (McHugh, 2012), Biochemia Medica, via PubMed Central