QA and compliance audit for AI voice agents

Every AI voice agent call checked.
Every release proven.

Loops is QA and compliance audit software for AI voice agents. It grades every call against your own standard operating procedures (SOP), catches the release that broke a rule, and shows what each release earned.

Live on every AI call at three US collection agencies

Works withRetellVapiBlandElevenLabsLiveKitPipecator any platform that sends finished calls by API or webhook

One call graded against the SOP, a release that broke one rule, and what the release earned. Sample data.
Every verdictquotes the transcript line it relied on
100% of callsgraded, where manual QA typically reviews under 5%
72 hoursfrom your rule approval to your first audit

The gap

Your agent is live. Can you prove it's working?

Compliance

Is it following our rules?

Manual QA typically reviews under 5% of calls. McKinsey, 2024

Operations

Is the new version better or worse?

Tests pass. Live calls drift.

Finance

What is it earning us?

Platform reports count minutes and containment, not payments collected.

Up to $1,500

per call in TCPA damages when a court finds an AI-voice call without consent willful or knowing, and $500 a call when it doesn't. The FDCPA adds actual damages, up to $1,000 in statutory damages per lawsuit, and attorney's fees. Loops shows, call by call, whether your agent followed the rules that prevent both, and checks consent before each call when you send your consent records. Sources: 47 U.S.C. § 227(b)(3) and 15 U.S.C. § 1692k Since February 2024, the FCC treats AI-generated voices as artificial voices under the TCPA, so AI calls need the same consent as prerecorded ones. FCC ruling 24-17 Gartner expects over 40% of agentic AI projects to be canceled by the end of 2027, over cost, unclear value or weak risk controls. Gartner, June 2025

How it works

One loop, from your rules to your results

Four steps, then the loop closes: fix a rule once, and the next release is proven against the last.

1Start from your rules

Your SOP is the rubric

Upload the SOP you already have. Loops splits it into numbered rules, and your compliance lead approves them before anything is graded.

Collections SOP · rev 7Approved
  1. 1.2Confirm the right person before any account detail
  2. 2.3Give the debt collection disclosure before discussing the debtSOP §3.1, page 4
  3. 3.2Offer a payment plan before asking for the full balance
  4. 4.1Log the promise to pay with log_promise
4 of 41 rules shownEach rule cites its SOP line

2Check every call

Every call graded, with evidence

Loops grades each call, AI or human, rule by rule. Every verdict quotes the line it relied on, and Loops checks that required tool calls fired and returned success.

Call 8f3a-2291 · agent v14Sample

AgentHi, this is Maya, an automated assistant calling for Quillfeather Services, on a recorded line. Am I speaking with Daniel Reyes?

CustomerYes, that's me.

AgentThanks. Please confirm your date of birth.

CustomerMarch 4, 1986.

1.2 met · verify_identity returned a match at 00:14

AgentYour balance with Alder Card is $1,284.16, 94 days past due.

2.3 missed · no disclosure before the debt

CustomerI can do $107 on the 3rd.

AgentYou're set for $107 on October 3.

4.1 met · log_promise returned success at 00:39
8 rules checked7 met · 1 missed

3Prove every release

See exactly what a release broke

Compare each rule's miss rate, this release against the last, and open the calls behind any number that moved.

Missed-rule rate · v13 → v14 · scale 0 to 10%Sample
1.2Identity1.4% → 0.8%
2.3Disclosure0.3% → 1.9%
3.2Offer a plan9.8% → 5.1%
4.1Log the promise6.5% → 3.2%
Tests passed. v14 still broke 2.3: +1.6 pts (95% +1.3 to +1.9).190 calls listed for remediation

4See what it earned

A dollar figure on every release

Join calls to your payment records and see what each release earned or cost, next to the rule changes that explain it.

v14 vs v13 · per 10,000 accounts reachedSample
+$18,400more in payments posted95% range +$11,200 to +$25,600 · about 200,000 accounts a side, same weeks, same clients
Promise-to-pay rate (per right-party contact)
12.8% → 13.7%
Cost per resolved call
$4.10 → $3.62
To fix before v15
2.3 on 190 calls

Fix the rule once.

Loops drafts the updated prompt, tests and checks. Your team reviews and ships them. Loops never writes to your agent.

Start a free 30-day audit

v15: 2.3 fixed, v14's gains kept

Watch

Loops in 30 seconds

Two short explainers, captioned. Both use sample data, not customer results.

How Loops worksFrom your operating procedures to a proven release, in one loop.
Read the transcript

Your AI agents handle thousands of calls a week. Can you prove they follow your rules? Loops can. Upload your operating procedures, and your compliance lead approves them as numbered rules. Loops grades every call against them and quotes the line behind each verdict. Ship a release, and Loops shows which rules improved, which one broke, and what it earned. Fix the rule once. Prove the next release against the last. Loops. Every call checked. Every release proven.

The release that passed every testAn illustrative story: what the tests missed, and what the release earned.
Read the transcript

A sample story with illustrative figures. Release 14 passed every test, so it shipped. A week later, Loops showed what the tests missed. The disclosure rule was skipped on 190 calls, each one listed with the line as evidence. A month in, the payments showed it also brought in $18,400 more per 10,000 accounts reached. So the team fixed the rule, logged those accounts for remediation, and proved release 15 against 14 on live calls. That is what an independent audit gives you. Loops.

Independent by design

The builder builds. Loops keeps score.

Your platform's QA scores calls on that platform, on criteria you write in its format. Loops grades every call against your approved SOP rules, the same way on any platform and on your human floor.

You

Plan

Write and approve the SOP.

Your platform or team

Build

Retell, Vapi, Bland, ElevenLabs or your own code runs the agent.

Loops

Test, check, prove

Grades every call and every release against your rules.

You, with Loops' drafts

Improve

Your team reviews the drafted fix and decides what ships.

Compare

Loops vs platform QA and voice agent testing tools

Platform QA and test tools are good at their jobs. Loops does a narrower one and runs alongside both.

Capability LoopsIndependent audit Platform QABuilt into Retell, Vapi, Bland, ElevenLabs Agent test toolsCoval, Cekura, Hamming Your own LLM judgeBuilt in-house QA team todayManual sampling
Graded against your own SOP, rule by rule Approved, versioned rules Custom criteria, their format Custom metrics, their format Whatever your team wrote On the calls they sample
Every production call, not a sample Every call Every call or a sample you size, by platform Varies by tool Every call Under 5%, by hand
Human and AI calls on one rubric Same rules, same screen AI only AI only If you build it Humans only
This release against the last, per rule, in production With the calls attached Scores per criterion; no side-by-side release report Mostly before release If you build it No
Dollars per release from your payment records Yes No No If you build the join No
Independent of whoever runs the agent Yes Built in, by the same vendor Yes No, the team that ships grades it Yes
Real-time latency and interruption monitoring No, Loops grades after the call Yes Yes No No

Based on each vendor's public documentation, September 2026. If a cell is out of date, tell us through the form and we'll fix it. Product names are trademarks of their owners; Loops is not affiliated with or endorsed by any of them.

Rule library

Compliance rules for collections, insurance and healthcare calls

Start from a library of rules for your industry, each tagged with where it comes from, then add your own. It's quality assurance and call monitoring on every call, not a sample.

Live today

Collections and loan servicing

FDCPA · Reg F · TCPA · state rules

  • Disclosure before discussing the debt Reg F
  • Right party before any account detail FDCPA
  • Seven calls in seven days, from your dial log Reg F
  • State call caps, like two in seven days in Massachusetts State law
  • Calling hours in every time zone the consumer may be in Reg F
  • Stop-calling, attorney and dispute requests, each handled the way Reg F treats it Reg F
  • Voicemails kept to the limited-content message Reg F
  • Consent on file before an AI-voice call, from your consent records TCPA

Sample verdictMissedBalance quoted before the debt collection disclosure

Collections rules in detail

Rule library ready

Insurance

State unfair claims laws · privacy and recording laws

  • Verify the policyholder before discussing a claim Privacy law
  • Never promise a claim outcome Claims law
  • Acknowledge the claim and explain what the policy requires next Claims law
  • Tell a caller who disputes a decision how to get it reviewed Claims law

Sample verdictMissed"Your claim will be approved" promised an outcome

Claims call rules in detail

Rule library ready

Healthcare patient access

HIPAA · TCPA · No Surprises Act · state medical debt laws

  • Verify identity and authority before any PHI HIPAA
  • Honor confidential-communication requests HIPAA
  • Name the caller at the start and give a callback number TCPA
  • Tell self-pay patients a good faith estimate is available when scheduling No Surprises Act
  • No diagnosis, procedure or balance on voicemail HHS guidance
  • Quote only the balance your billing system returned Your SOP
  • Mention financial assistance before asking for payment Your SOP

Sample verdictMetDate of birth and ZIP confirmed before the balance

Patient-access rules in detail

Security and data

What Loops can touch, and what it can't

Reads calls. Places test calls only if you ask.

Where a key can be limited to reading, as on Retell (History: Read) and ElevenLabs, that's the only key we accept. Where it can't, as on Vapi, we recommend the webhook: send finished calls and give us no key at all. If you'd rather give a key, it is a dedicated one we use only to read calls, every request is logged, and you can revoke it any time.

Never writes to your agent

Loops never calls anything that changes a prompt, flow, setting or number, and the request log shows it. Prompt changes arrive as drafts for your team.

Test calls are opt-in

Pre-release runs dial a test agent you name, using a separate key. Where a platform's keys can't be limited to one agent, keep the test agent in its own workspace. Your platform bills those minutes as usual. Leave it off and Loops still compares releases on live calls.

US data, encrypted

Call data is stored and processed in the United States and encrypted in transit and at rest. You set how long we keep it.

No training on your calls

Nothing you send trains any model, ours or a provider's.

Paperwork before access

NDA and DPA, plus a BAA for healthcare, signed before we receive a key. Our SOC 2 Type II report and bridge letter come with the security packet, under NDA.

Request the security packetVisit the trust center

The free audit

Your first 30 days

No build work on Retell, Vapi, Bland or ElevenLabs. LiveKit, Pipecat and in-house agents add a small SDK.

  1. Day 0
    Connect

    Sign the NDA and DPA (and a BAA for healthcare providers and health plans), connect your platform, choose which calls to include and send your SOP. For collections, add your dial log and consent records so frequency and consent get checked, and a daily payments file if you want dollars in the report.

  2. Days 1 to 2
    Approve the rules

    Your compliance lead reviews each rule beside the SOP text it came from.

  3. +72 hours
    First audit

    Last month's calls, graded rule by rule, with the evidence for every verdict. If your platform keeps no call history, as with zero data retention settings, audits start on new calls the day you connect.

  4. Day 30
    Your report

    Yours to keep, including how often Loops agreed with your QA reviewers, what each release in the period earned or cost if you sent payments, and, if you sent human calls, how your AI agent compared with your human floor. Nothing renews on its own. If you don't continue, we delete your calls, transcripts and recordings within 30 days and confirm it in writing.

FAQ

Questions about auditing AI voice agent calls

What is Loops?

Loops is independent QA and compliance audit software for AI voice agents. It grades every call against your own SOP, rule by rule, with the transcript line as evidence, compares each release with the last, and reports what each release earned from your payment records. How to audit AI voice agent calls against your SOP

How do I audit a Retell or Vapi agent against my SOP?

On Retell, create a key restricted to History: Read. On Vapi, send finished calls by webhook, or give a dedicated key. Then upload your SOP. Loops splits the SOP into numbered rules for your compliance lead to approve, backfills last month's calls, and returns the first audit within 72 hours of approval.

Does Loops work with LiveKit, Pipecat or in-house agents?

Yes, through the Loops SDK for Python or Node, which we share during onboarding (it isn't the loops package on npm, which belongs to loops.so). After each call ends, it sends the transcript, tool calls and agent version, off the audio path, so it never slows a live call. You choose what it sends, and you can mask account numbers and dates of birth in your own process before anything leaves.

How does Loops know a new release regressed?

Every call carries the agent version your platform records, or a version tag you set in call metadata or through the SDK. Loops compares each rule's miss rate between versions on live calls and lists the calls behind any change. Both sides of a comparison are graded by the same grader version, and the report says so. Optional pre-release test calls catch problems before a release ships. Why a release that passes its tests can still break a rule

How is a dollar figure tied to a release?

Send a daily payments file from your collection or billing system (account ID, amount, date posted). A payment counts toward the release that handled the account's last call before it posted, within a window you set, 30 days by default. Releases are compared on calls from the same weeks, the same clients and similar balance and age bands, and the report shows the range around every figure and says when a difference is too small to call. When your human calls are graded too, the same method compares your AI agent with your human floor. How to measure what a release earned

What if my platform doesn't keep call history?

If your platform keeps no transcripts, as with zero data retention settings, there's nothing to backfill. Send each finished call to Loops by webhook or through the SDK, and audits start that day. HIPAA modes that keep calls for a retention period you set can usually be backfilled.

Why not grade calls with our own LLM judge?

You can, and many teams start there. Loops adds what an examiner asks about: rules your compliance lead approved and versioned, a grader your engineers didn't write and can't tune toward their own release, the same rules on your human floor, and the payments join. Your judge can keep running beside it. Grading calls with your own LLM judge

Does Loops change my agent or write my prompts?

No. Loops never writes to your platform. When you fix a rule, it drafts the prompt change, tests and checks for your team to review, and your team decides what ships.

What can I hand an examiner or a client auditor?

For any call, account or date range: each rule's verdict, the quoted line, the tool-call record, the rule version and who approved it, the grader version, and any reviewer override with who made it and when. Export it as PDF or CSV.

Can verdicts flow into our own systems?

Yes. Every verdict, with its rule number, the quoted line and the call ID, exports as CSV and is available to your CRM or compliance log by API.

How accurate is the audit?

Every verdict quotes the line and the tool-call record it relied on, so a reviewer can check it in seconds. When a transcript can't settle a rule, Loops marks the call for review rather than guessing. Your QA team can overturn any verdict, and your day-30 report measures how often Loops agreed with your reviewers on your own calls. It lists separately every call Loops passed that a reviewer failed, so the misses that matter most are never averaged away.

Can Loops audit human agents too?

Yes. Send recordings from your human floor by upload or SFTP, and they're graded against the same rules as your AI agent, on the same screen.

Does it catch rules that span several calls, like Reg F's seven-in-seven or TCPA consent?

Yes, when you send your dial log and consent records. Loops then checks call frequency against Reg F and that consent was on file before each AI-voice call. Without them, Loops grades what happens inside each call and marks cross-call and consent rules as not checked. TCPA consent for AI voice calls

Our agent isn't live yet. Can Loops help?

Yes. Approve your rules now and run pre-release test calls against a test agent you name, so your first release is graded before launch and every release after it is compared with the last.

Can a CRM or voice agent platform offer Loops to its clients?

Yes. Each of your clients approves its own rules and gets its own report, and their calls stay separate. Tell us about your platform through the form and we'll walk through how verdicts come back to you.

Free 30-day audit

Find out what your agent said last month.

Tell us what you run. A person on our team replies within one business day with next steps.

  • Every call from last month, graded against your SOP
  • Paperwork signed before we receive any access
  • The report is yours to keep, whether or not you continue
What runs your agents? Pick any
What would you like first?

We use your email only to reply about this audit. See our privacy policy.