Compliance
Is it following our rules?
Manual QA typically reviews under 5% of calls. McKinsey, 2024
QA and compliance audit for AI voice agents
Loops is QA and compliance audit software for AI voice agents. It grades every call against your own standard operating procedures (SOP), catches the release that broke a rule, and shows what each release earned.
Live on every AI call at three US collection agencies
Works withRetellVapiBlandElevenLabsLiveKitPipecator any platform that sends finished calls by API or webhook
The gap
Compliance
Manual QA typically reviews under 5% of calls. McKinsey, 2024
Operations
Tests pass. Live calls drift.
Finance
Platform reports count minutes and containment, not payments collected.
Up to $1,500
per call in TCPA damages when a court finds an AI-voice call without consent willful or knowing, and $500 a call when it doesn't. The FDCPA adds actual damages, up to $1,000 in statutory damages per lawsuit, and attorney's fees. Loops shows, call by call, whether your agent followed the rules that prevent both, and checks consent before each call when you send your consent records. Sources: 47 U.S.C. § 227(b)(3) and 15 U.S.C. § 1692k Since February 2024, the FCC treats AI-generated voices as artificial voices under the TCPA, so AI calls need the same consent as prerecorded ones. FCC ruling 24-17 Gartner expects over 40% of agentic AI projects to be canceled by the end of 2027, over cost, unclear value or weak risk controls. Gartner, June 2025
How it works
Four steps, then the loop closes: fix a rule once, and the next release is proven against the last.
1Start from your rules
Upload the SOP you already have. Loops splits it into numbered rules, and your compliance lead approves them before anything is graded.
2Check every call
Loops grades each call, AI or human, rule by rule. Every verdict quotes the line it relied on, and Loops checks that required tool calls fired and returned success.
AgentHi, this is Maya, an automated assistant calling for Quillfeather Services, on a recorded line. Am I speaking with Daniel Reyes?
CustomerYes, that's me.
AgentThanks. Please confirm your date of birth.
CustomerMarch 4, 1986.
AgentYour balance with Alder Card is $1,284.16, 94 days past due.
CustomerI can do $107 on the 3rd.
AgentYou're set for $107 on October 3.
3Prove every release
Compare each rule's miss rate, this release against the last, and open the calls behind any number that moved.
4See what it earned
Join calls to your payment records and see what each release earned or cost, next to the rule changes that explain it.
Loops drafts the updated prompt, tests and checks. Your team reviews and ships them. Loops never writes to your agent.
Watch
Two short explainers, captioned. Both use sample data, not customer results.
Your AI agents handle thousands of calls a week. Can you prove they follow your rules? Loops can. Upload your operating procedures, and your compliance lead approves them as numbered rules. Loops grades every call against them and quotes the line behind each verdict. Ship a release, and Loops shows which rules improved, which one broke, and what it earned. Fix the rule once. Prove the next release against the last. Loops. Every call checked. Every release proven.
A sample story with illustrative figures. Release 14 passed every test, so it shipped. A week later, Loops showed what the tests missed. The disclosure rule was skipped on 190 calls, each one listed with the line as evidence. A month in, the payments showed it also brought in $18,400 more per 10,000 accounts reached. So the team fixed the rule, logged those accounts for remediation, and proved release 15 against 14 on live calls. That is what an independent audit gives you. Loops.
Independent by design
Your platform's QA scores calls on that platform, on criteria you write in its format. Loops grades every call against your approved SOP rules, the same way on any platform and on your human floor.
You
PlanWrite and approve the SOP.
Your platform or team
BuildRetell, Vapi, Bland, ElevenLabs or your own code runs the agent.
Loops
Test, check, proveGrades every call and every release against your rules.
You, with Loops' drafts
ImproveYour team reviews the drafted fix and decides what ships.
Compare
Platform QA and test tools are good at their jobs. Loops does a narrower one and runs alongside both.
| Capability | LoopsIndependent audit | Platform QABuilt into Retell, Vapi, Bland, ElevenLabs | Agent test toolsCoval, Cekura, Hamming | Your own LLM judgeBuilt in-house | QA team todayManual sampling |
|---|---|---|---|---|---|
| Graded against your own SOP, rule by rule | Approved, versioned rules | Custom criteria, their format | Custom metrics, their format | Whatever your team wrote | On the calls they sample |
| Every production call, not a sample | Every call | Every call or a sample you size, by platform | Varies by tool | Every call | Under 5%, by hand |
| Human and AI calls on one rubric | Same rules, same screen | AI only | AI only | If you build it | Humans only |
| This release against the last, per rule, in production | With the calls attached | Scores per criterion; no side-by-side release report | Mostly before release | If you build it | No |
| Dollars per release from your payment records | Yes | No | No | If you build the join | No |
| Independent of whoever runs the agent | Yes | Built in, by the same vendor | Yes | No, the team that ships grades it | Yes |
| Real-time latency and interruption monitoring | No, Loops grades after the call | Yes | Yes | No | No |
Based on each vendor's public documentation, September 2026. If a cell is out of date, tell us through the form and we'll fix it. Product names are trademarks of their owners; Loops is not affiliated with or endorsed by any of them.
Rule library
Start from a library of rules for your industry, each tagged with where it comes from, then add your own. It's quality assurance and call monitoring on every call, not a sample.
FDCPA · Reg F · TCPA · state rules
Sample verdictMissedBalance quoted before the debt collection disclosure
State unfair claims laws · privacy and recording laws
Sample verdictMissed"Your claim will be approved" promised an outcome
HIPAA · TCPA · No Surprises Act · state medical debt laws
Sample verdictMetDate of birth and ZIP confirmed before the balance
Security and data
Where a key can be limited to reading, as on Retell (History: Read) and ElevenLabs, that's the only key we accept. Where it can't, as on Vapi, we recommend the webhook: send finished calls and give us no key at all. If you'd rather give a key, it is a dedicated one we use only to read calls, every request is logged, and you can revoke it any time.
Loops never calls anything that changes a prompt, flow, setting or number, and the request log shows it. Prompt changes arrive as drafts for your team.
Pre-release runs dial a test agent you name, using a separate key. Where a platform's keys can't be limited to one agent, keep the test agent in its own workspace. Your platform bills those minutes as usual. Leave it off and Loops still compares releases on live calls.
Call data is stored and processed in the United States and encrypted in transit and at rest. You set how long we keep it.
Nothing you send trains any model, ours or a provider's.
NDA and DPA, plus a BAA for healthcare, signed before we receive a key. Our SOC 2 Type II report and bridge letter come with the security packet, under NDA.
The free audit
No build work on Retell, Vapi, Bland or ElevenLabs. LiveKit, Pipecat and in-house agents add a small SDK.
Sign the NDA and DPA (and a BAA for healthcare providers and health plans), connect your platform, choose which calls to include and send your SOP. For collections, add your dial log and consent records so frequency and consent get checked, and a daily payments file if you want dollars in the report.
Your compliance lead reviews each rule beside the SOP text it came from.
Last month's calls, graded rule by rule, with the evidence for every verdict. If your platform keeps no call history, as with zero data retention settings, audits start on new calls the day you connect.
Yours to keep, including how often Loops agreed with your QA reviewers, what each release in the period earned or cost if you sent payments, and, if you sent human calls, how your AI agent compared with your human floor. Nothing renews on its own. If you don't continue, we delete your calls, transcripts and recordings within 30 days and confirm it in writing.
FAQ
Loops is independent QA and compliance audit software for AI voice agents. It grades every call against your own SOP, rule by rule, with the transcript line as evidence, compares each release with the last, and reports what each release earned from your payment records. How to audit AI voice agent calls against your SOP
On Retell, create a key restricted to History: Read. On Vapi, send finished calls by webhook, or give a dedicated key. Then upload your SOP. Loops splits the SOP into numbered rules for your compliance lead to approve, backfills last month's calls, and returns the first audit within 72 hours of approval.
Yes, through the Loops SDK for Python or Node, which we share during onboarding (it isn't the loops package on npm, which belongs to loops.so). After each call ends, it sends the transcript, tool calls and agent version, off the audio path, so it never slows a live call. You choose what it sends, and you can mask account numbers and dates of birth in your own process before anything leaves.
Every call carries the agent version your platform records, or a version tag you set in call metadata or through the SDK. Loops compares each rule's miss rate between versions on live calls and lists the calls behind any change. Both sides of a comparison are graded by the same grader version, and the report says so. Optional pre-release test calls catch problems before a release ships. Why a release that passes its tests can still break a rule
Send a daily payments file from your collection or billing system (account ID, amount, date posted). A payment counts toward the release that handled the account's last call before it posted, within a window you set, 30 days by default. Releases are compared on calls from the same weeks, the same clients and similar balance and age bands, and the report shows the range around every figure and says when a difference is too small to call. When your human calls are graded too, the same method compares your AI agent with your human floor. How to measure what a release earned
If your platform keeps no transcripts, as with zero data retention settings, there's nothing to backfill. Send each finished call to Loops by webhook or through the SDK, and audits start that day. HIPAA modes that keep calls for a retention period you set can usually be backfilled.
You can, and many teams start there. Loops adds what an examiner asks about: rules your compliance lead approved and versioned, a grader your engineers didn't write and can't tune toward their own release, the same rules on your human floor, and the payments join. Your judge can keep running beside it. Grading calls with your own LLM judge
No. Loops never writes to your platform. When you fix a rule, it drafts the prompt change, tests and checks for your team to review, and your team decides what ships.
For any call, account or date range: each rule's verdict, the quoted line, the tool-call record, the rule version and who approved it, the grader version, and any reviewer override with who made it and when. Export it as PDF or CSV.
Yes. Every verdict, with its rule number, the quoted line and the call ID, exports as CSV and is available to your CRM or compliance log by API.
Every verdict quotes the line and the tool-call record it relied on, so a reviewer can check it in seconds. When a transcript can't settle a rule, Loops marks the call for review rather than guessing. Your QA team can overturn any verdict, and your day-30 report measures how often Loops agreed with your reviewers on your own calls. It lists separately every call Loops passed that a reviewer failed, so the misses that matter most are never averaged away.
Yes. Send recordings from your human floor by upload or SFTP, and they're graded against the same rules as your AI agent, on the same screen.
Yes, when you send your dial log and consent records. Loops then checks call frequency against Reg F and that consent was on file before each AI-voice call. Without them, Loops grades what happens inside each call and marks cross-call and consent rules as not checked. TCPA consent for AI voice calls
Yes. Approve your rules now and run pre-release test calls against a test agent you name, so your first release is graded before launch and every release after it is compared with the last.
Yes. Each of your clients approves its own rules and gets its own report, and their calls stay separate. Tell us about your platform through the form and we'll walk through how verdicts come back to you.
Free 30-day audit
Tell us what you run. A person on our team replies within one business day with next steps.