Compliance
Is it following our rules?
Manual QA typically reviews under 5% of calls. McKinsey, 2024
QA and compliance audit for AI voice agents
Loops is QA and compliance audit software for AI voice agents. It grades every call against your own standard operating procedures (SOP), catches the release that broke a rule, and shows what each release earned.
Live on every AI call at three US collection agencies
Works withRetellVapiBlandElevenLabsLiveKitPipecator any platform that sends finished calls by API or webhook
The gap
Compliance
Manual QA typically reviews under 5% of calls. McKinsey, 2024
Operations
Tests pass. Live calls drift.
Finance
Minutes and containment, not dollars.
$1,500
per call in TCPA statutory damages for a willful AI-voice call made without consent, and $500 when it isn't willful. The FDCPA adds up to $1,000 per consumer, plus attorney's fees. Loops shows, call by call, whether your agent followed the rules that prevent both. Sources: 47 U.S.C. § 227(b)(3) and 15 U.S.C. § 1692k Since February 2024, the FCC treats AI-generated voices as artificial voices under the TCPA, so AI calls need the same consent as prerecorded ones. FCC ruling 24-17 Gartner expects over 40% of agentic AI projects to be canceled by the end of 2027, over cost, unclear value or weak risk controls. Gartner, June 2025
How it works
Four steps, then the loop closes: fix a rule once, and the next release is proven against the last.
1Start from your rules
Upload the SOP you already have. Loops splits it into numbered rules, and your compliance lead approves them before anything is graded.
2Check every call
Loops grades each call, AI or human, rule by rule. Every verdict quotes the line it relied on, and Loops checks that required tool calls fired and returned success.
AgentHi, this is Maya, an automated assistant, on a recorded line. Am I speaking with Daniel Reyes?
CustomerYes, that's me.
AgentThanks. Please confirm your date of birth.
CustomerMarch 4, 1986.
AgentYour balance with Alder Card is $1,284.16, 94 days past due.
AgentYou're set for $107 on October 3.
3Prove every release
Compare each rule's miss rate, this release against the last, and open the calls behind any number that moved.
4See what it earned
Join calls to your payment records and see what each release earned or cost, next to the rule changes that explain it.
Loops drafts the updated prompt, tests and checks. Your team reviews and ships them. Loops never writes to your agent.
Watch
Two short explainers, captioned.
Your AI agents handle thousands of calls a week. Can you prove they follow your rules? Loops can. Upload your operating procedures, and your compliance lead approves them as numbered rules. Loops grades every call against them and quotes the line behind each verdict. Ship a release, and Loops shows which rules improved, which one broke, and what it earned. Fix the rule once. Prove the next release against the last. Loops. Every call checked. Every release proven.
Release 14 passed every test, so it shipped. A week later, Loops showed what the tests missed. The disclosure rule was skipped on 190 calls, each one listed with the line as evidence. A month in, the payments showed it also brought in $18,400 more per 10,000 calls. So the team fixed the rule, logged those accounts for remediation, and proved release 15 against 14 on live calls. That is what an independent audit gives you. Loops.
Independent by design
Your platform's QA grades its own calls on its own criteria. Loops grades every call against your SOP, the same way on any platform and on your human floor.
You
PlanWrite and approve the SOP.
Your platform or team
BuildRetell, Vapi, Bland, ElevenLabs or your own code runs the agent.
Loops
Test, check, proveGrades every call and every release against your rules.
You, with Loops' drafts
ImproveYour team reviews the drafted fix and decides what ships.
Compare
Platform QA and test tools are good at their jobs. Loops does a narrower one, and most teams use it alongside both.
| Capability | LoopsIndependent audit | Platform QABuilt into Retell, Vapi, Bland, ElevenLabs | Agent test toolsCoval, Cekura, Hamming | QA team todayManual sampling |
|---|---|---|---|---|
| Graded against your own SOP, rule by rule | Approved, versioned rules | Custom criteria, their format | Custom metrics, their format | On the calls they sample |
| Every production call, not a sample | Every call | Their own calls | Varies by tool | Under 5%, by hand |
| Human and AI calls on one rubric | Same rules, same screen | AI only | AI only | Humans only |
| This release against the last, per rule, in production | With the calls attached | Varies by platform; no per-rule diff | Mostly before release | No |
| Dollars per release from your payment records | Yes | No | No | No |
| Independent of whoever runs the agent | Yes | Built in, by the same vendor | Yes | Yes |
| Real-time latency and interruption monitoring | No, Loops grades after the call | Yes | Yes | No |
Based on each vendor's public documentation, September 2026. If a cell is out of date, tell us through the form and we'll fix it. Product names are trademarks of their owners; Loops is not affiliated with or endorsed by any of them.
Regulated floors
Start from a library of regulatory rules for your industry, then add your own. It's quality assurance and call monitoring on every call, not a sample.
FDCPA · Reg F · TCPA · state rules
Sample verdictMissedBalance quoted before the debt collection disclosure
State unfair claims practices laws
Sample verdictMissed"Your claim will be approved" promised an outcome
HIPAA · TCPA · state medical debt laws
Sample verdictMetDate of birth and ZIP confirmed before the balance
Security and data
Where a key can be limited to reading, as on ElevenLabs, we ask for that. Where it can't, we use a dedicated key only to read calls, log every request, and you can revoke it any time. Or send finished calls by webhook and give us no key at all.
Loops never calls anything that changes a prompt, flow, setting or number, and the request log shows it. Prompt changes arrive as drafts for your team.
Pre-release runs dial a test agent you name, using a separate key, and your platform bills those minutes as usual. Leave it off and Loops still compares releases on live calls.
Call data is stored and processed in the United States and encrypted in transit and at rest. You set how long we keep it.
Nothing you send trains any model, ours or a provider's.
NDA and DPA, plus a BAA for healthcare, signed before we receive a key. Our SOC 2 Type II report and bridge letter come with the security packet, under NDA.
The free audit
No build work on Retell, Vapi, Bland or ElevenLabs. LiveKit, Pipecat and in-house agents add a small SDK.
Sign the NDA and DPA (and a BAA for healthcare), connect your platform, choose which calls to include and send your SOP.
Your compliance lead reviews each rule beside the SOP text it came from.
Last month's calls, graded rule by rule, with the evidence for every verdict. With transcript storage off, as in HIPAA modes, audits start on new calls the day you connect.
Yours to keep, including how often Loops agreed with your QA reviewers and, if you sent human calls, how your AI agent compared with your human floor. Nothing renews on its own. If you don't continue, we delete your calls, transcripts and recordings within 30 days and confirm it in writing.
FAQ
Loops is independent QA and compliance audit software for AI voice agents. It grades every call against your own SOP, rule by rule, with the transcript line as evidence, compares each release with the last, and reports what each release earned from your payment records. How to audit AI voice agent calls against your SOP
Give Loops an API key from your platform and upload your SOP. Loops splits the SOP into numbered rules for your compliance lead to approve, backfills last month's calls, and returns the first audit within 72 hours of approval.
Yes, through the Loops SDK for Python or Node. After each call ends, it sends the transcript, tool calls and agent version, off the audio path, so it never slows a live call. You choose what it sends, and you can mask account numbers and dates of birth in your own process before anything leaves.
Every call carries the agent version your platform records, or a version tag you set in call metadata or through the SDK. Loops compares each rule's miss rate between versions on live calls and lists the calls behind any change. Both sides of a comparison are graded by the same grader version, and the report says so. Optional pre-release test calls catch problems before a release ships. Why a release that passes its tests can still break a rule
Send a daily payments file from your collection or billing system (account ID, amount, date posted). A payment counts toward the release that handled the account's last call before it posted, within a window you set, 30 days by default. Releases are compared on calls from the same weeks, the same clients and similar balance and age bands, and the report shows the range around every figure and says when a difference is too small to call. When your human calls are graded too, the same method compares your AI agent with your human floor. How to measure what a release earned
With transcript storage off, as in HIPAA modes, there's nothing to backfill. Send each finished call to Loops by webhook or through the SDK, and audits start that day.
You can, and many teams start there. Loops adds what an examiner asks about: rules your compliance lead approved and versioned, a grader your engineers didn't write and can't tune toward their own release, the same rules on your human floor, and the payments join. Your judge can keep running beside it. Grading calls with your own LLM judge
No. Loops never writes to your platform. When you fix a rule, it drafts the prompt change, tests and checks for your team to review, and your team decides what ships.
For any call, account or date range: each rule's verdict, the quoted line, the tool-call record, the rule version and who approved it, and any reviewer override with who made it and when. Export it as PDF or CSV.
Yes. Every verdict, with its rule number, the quoted line and the call ID, exports as CSV and is available to your CRM or compliance log by API.
Every verdict quotes the line and the tool-call record it relied on, so a reviewer can check it in seconds. When a transcript can't settle a rule, Loops marks the call for review rather than guessing. Your QA team can overturn any verdict, and your day-30 report measures how often Loops agreed with your reviewers on your own calls. It lists separately every call Loops passed that a reviewer failed, so the misses that matter most are never averaged away.
Yes. Send recordings from your human floor by upload or SFTP, and they're graded against the same rules as your AI agent, on the same screen.
Yes, when you send your dial log and consent records. Loops then checks call frequency against Reg F and that consent was on file before each AI-voice call. Without them, Loops grades what happens inside each call and marks cross-call and consent rules as not checked. TCPA consent for AI voice calls
Yes. Each of your clients approves its own rules and gets its own report, and their calls stay separate. Tell us about your platform through the form and we'll walk through how verdicts come back to you.
Free 30-day audit
Tell us what you run. A person on our team replies within one business day with next steps.