
Picture a lending company whose contact center quality assurance team scores five calls per agent per month. Good scores across the board. Compliance looks clean.
Then it turns on 100% monitoring, and the calls nobody scored tell a different story: agents skipping mandatory risk disclosures on loan products. Not sometimes. Routinely. The QA team had been scoring the same five calls where agents knew they were being watched. The other 4,500 calls per month? Nobody heard them.
That gap between what you sample and what actually happens is where compliance violations live, where churn signals hide, and where the revenue intelligence sits that your CRM will never capture. This is the core problem with traditional contact center quality assurance: you’re making decisions from 2% of the data and hoping the other 98% looks the same.
It doesn’t.
What is call center quality assurance?
Call center quality assurance is the operational framework that checks whether customer conversations meet the company's standards — what agents say, how they solve the problem, what they must disclose — and turns the findings into coaching and process fixes. It borrows from manufacturing quality management: it is process-oriented, it aims to prevent defects rather than find them after the fact, and it is everyone's responsibility on the floor, not only the QA team's.
In practice, most QA programs work from a sample: a reviewer listens to a handful of calls per agent, scores each against a scorecard, and shares the result with the agent and the team lead. The rest of this article is about what that sample misses.
Contact Center Quality Assurance by the Numbers: Why 2% Sampling Fails
Here’s why the sampling model breaks down. A typical contact center handles somewhere between 2,000 and 50,000 calls per month, depending on size. QA teams manually review 2-5 calls per agent per month. In a 200-seat center running 20,000 monthly calls, that’s 400-1,000 reviews. At best, 5% coverage. At worst, 2%.
Once you know which conversations matter, reviewing them is the fast part — here is how to find and filter conversations and pick the right scorecard in Ender Turing.
That’s not quality assurance. That’s a lottery.
The statistical problem is severe. With a 2% sample, you need a compliance violation to occur in roughly 1 in 50 calls for your QA team to have a reasonable chance of catching it in any given month. If the violation rate is 5% (which is common for soft compliance issues like missing disclosures), your manual QA program has about a 10% chance of flagging it for a specific agent.
Manual review costs $5-10 per evaluation, according to industry benchmarks. At those rates, a 200-seat center spending $6 per review on 800 calls monthly is paying $57,600 a year to monitor 4% of conversations. The return on that investment is guesswork dressed up as quality management.
And the people doing the reviews? They’re inconsistent. McKinsey estimates that manual QA scoring is 70–80% accurate, against more than 90% for a largely automated process. The machine doesn’t have a bad Monday.
How many calls should QA actually review?
The honest answer: as many as you can learn from — and manual review can’t get there. At $5–10 per evaluation, a 200-seat center can afford to review 400–1,000 of its 20,000 monthly calls: 2–5% coverage. With a 5% violation rate, a 2% sample gives you roughly a 1-in-10 chance of catching it for a specific agent in a given month. Automated scoring covers 100% of calls for under $0.50 per conversation — which is why the question is shifting from “how many can we afford to review” to “why would we leave 98% unheard.”
What Lives in the 98%
Here is what 100% call monitoring surfaces in the calls nobody hears.
Compliance drift. Agents who pass manual QA consistently are skipping required disclosures on calls that aren’t being monitored. In financial services, this is a ticking regulatory bomb. The FCA’s Consumer Duty (fully enforced since July 2024) now requires firms to monitor outcomes across all customer interactions, not just samples. The FCA imposed 176 million GBP in fines in 2024, a 230% increase from the previous year. Nine of the fifteen highest fines were tied to poor internal management and control failures.
Revenue signals nobody hears. Cross-sell opportunities mentioned by customers. Churn warnings expressed in frustration patterns. Product feedback that never makes it to the product team. McKinsey found that inbound service call centers generated up to 25% of new revenue at some credit card companies and up to 60% at some telecom companies. But that revenue intelligence sits in conversations nobody analyzes.
Agent gaming. Call avoidance patterns. Handle time manipulation. Cherry-picking easy calls. These behaviors are invisible in a 2% sample because agents know when they’re being scored. In behavioral economics, this is called the Hawthorne effect. And it’s rampant in contact centers.
Coaching blind spots. Your best agent closes 3x more upsells than average. What do they say differently in the first 90 seconds? With 2% sampling, you’ll never know. With 100% analysis, you can extract the exact phrases, tonality patterns, and conversation structures that separate top performers from the rest.
The Compliance Pressure Is Accelerating
Three regulatory shifts are making 2% monitoring untenable. Not in five years. Now.
PCI DSS 4.0.1 (mandatory since March 2025). It forbids keeping the card security code in any call recording after the payment is authorized. Pause-and-resume leaves that to an agent remembering on every call. Every playback, download, and administrative action needs an audit trail entry with timestamps and user IDs. You can’t audit what you don’t monitor.
EU AI Act high-risk obligations (from 2 December 2027). AI systems used in contact centers for automated performance monitoring or employment decisions are classified as high-risk. A company that uses AI to score or monitor its agents must keep the system’s logs and tell the agents before it is used, or face fines of up to €15 million or 3% of worldwide turnover. That reaches companies outside the EU too when the results are used in the EU. If you’re already using AI for agent scoring, routing, or workforce management, you’re in scope. If you’re not monitoring the AI’s outputs across 100% of interactions, you have no compliance story.
SEC and CFTC off-channel enforcement. Since December 2021, the SEC and the CFTC have imposed more than $3.2 billion in penalties on financial firms for recordkeeping failures; in fiscal year 2024 alone, the SEC charged more than 70 firms and collected more than $600 million. Only 39% of firms enable and monitor all of their communication channels, up from 10.3% in 2023. The message from regulators is clear: if you can’t prove you monitored it, you’re liable for it.
From 2% to 100%: What Actually Changes
Cresta published a case study in December 2025 on how CVS Health went from scoring 5% of calls to 100% using AI-powered conversation intelligence. The results came fast. Visibility into customer satisfaction trends right away instead of waiting weeks for survey data. Almost immediate reductions in after-call work time. Agent-level performance insights from the customer’s perspective that were previously impossible.
Here’s what changes when a contact center makes the switch.
Week 1: The shock. QA scores that looked healthy under sampling often drop when every call is scored. This isn’t because agents suddenly got worse. It’s because the sample was never representative. Leaders realize they’ve been flying blind, and the recalibration is uncomfortable but necessary.
Month 1: Pattern recognition kicks in. With full data, you start seeing things that sampling hides. Which product generates the most confused calls. Which shift has the highest compliance drift. Which agents are strong on empathy but weak on resolution. These aren’t anomalies you catch with five calls a month. They’re systemic patterns that require thousands of data points to see.
Month 3: Coaching becomes targeted. Instead of generic coaching sessions based on a handful of cherry-picked calls, managers can identify specific skill gaps per agent and assign targeted training. Self-coaching dashboards let agents review their own calls against benchmarks. SQM Group reports that clients of its own Auto QA product have seen up to 600% ROI with payback inside three months, and most 300-400% ROI in year one.
Month 6: The data compounds. You have enough longitudinal data to spot trends. Agent attrition risk from conversation patterns. Seasonal compliance drift. Product issues surfacing in call topics weeks before they appear in CSAT surveys. This is where contact center quality assurance stops being about scoring calls and starts being about running the business.
The Mid-Market Trap
Large enterprises are adopting fast. Speech analytics deployment sits at roughly 44% across the industry, but the split is telling. Future Market Insights found that 55.6% of the $25.3 billion conversation intelligence market goes to large enterprises. Mid-market companies (200-1,000 seats) are dramatically underserved.
The reason is straightforward. First-generation speech analytics platforms were built for 5,000-seat deployments. Enterprise pricing. Enterprise implementation timelines. Six-month rollouts with a team of consultants. A 300-seat center can’t justify that investment or absorb that disruption.
But the compliance obligations don’t scale with company size. The FCA doesn’t give mid-market firms a lighter Consumer Duty. PCI DSS 4.0.1 applies whether you have 50 agents or 5,000. The regulatory pressure is identical. The tooling accessibility was not. That gap is closing fast as cloud-native platforms drop implementation timelines from months to weeks, but there’s a window right now where mid-market centers are carrying enterprise-grade compliance risk with startup-grade monitoring.
What Contact Center Quality Assurance Actually Requires in 2026
The old model was simple. Hire QA analysts. Score some calls. Coach agents quarterly. Hope for the best.
The new model looks different.
Monitor every interaction. Not 5%. Not 20%. All of them. Voice, chat, email, and bot conversations. The technology exists and the cost per interaction at scale is under $0.50, compared to $5-10 for manual review. If you’re not analyzing 100% of interactions, you don’t have quality assurance. You have quality sampling.
Score automatically, coach in real time. Automated scoring with over 90% accuracy means QA teams can focus on coaching instead of listening. Real-time alerts catch compliance issues as they happen, not three weeks later when the QA cycle catches up. The difference between catching a disclosure violation on Monday versus discovering it in Thursday’s coaching session is the difference between one call and 200.
Connect QA to business outcomes. Quality scores in isolation tell you nothing. When quality management connects to CRM data, CSAT results, and revenue outcomes, you can answer the questions that matter. Which agent behaviors drive retention? Which coaching interventions actually move CSAT? Where is the revenue hiding in your conversations?
Monitor the monitors. If you’re using AI chatbots or voice bots, who’s QA-ing them? We wrote about this recently in our post on AI agent quality assurance. Bots handle thousands of conversations daily with zero human oversight in most deployments. That’s the same 2% problem, except the bot doesn’t learn from being caught.
Build the audit trail. PCI DSS 4.0.1 and the EU AI Act both demand comprehensive logging. Every automated decision, every score, every flag needs documentation. This isn’t optional compliance overhead. It’s the foundation of defensibility when a regulator asks how you monitor quality. “We listen to five calls a month” is not an answer in 2026.
How to run call center quality assurance: KPIs, scorecard, framework
1. Set the KPIs that define quality. Pick the few numbers that describe a good conversation in your operation and set a target for each — for example first call resolution (FCR) of at least 70%, customer satisfaction (CSAT) of at least 75%, and goals for average handle time (AHT), after-call work (ACW), average speed of answer (ASA) and wait time. Targets depend on the queue: a collections line and a technical support line should not share one AHT goal.
2. Build the scorecard. Turn your standards into points an evaluator — a person or an AI — can answer consistently: greeting and verification, discovery, resolution, required disclosures, tone and empathy, closing. Weight the points that carry risk, such as compliance and disclosures, above the ones that carry style.
3. Decide how conversations get evaluated. Who scores, how often, and how the results reach agents. Manual review caps coverage at the sample described above; scoring 100% of conversations automatically turns the scorecard into a measurement of the whole operation, with human reviewers calibrating the AI and handling the exceptions.
4. Write the QA framework down. One page that names the quality standards, how outcomes are measured, the drivers you control (scripts, training, tools), who owns each part, and how success is measured.
5. Close the loop with coaching. Share scores with the agent and the team lead, coach on the calls that show the gap, encourage agent self-assessment, and check whether next week's scores move. Quality management in Ender Turing scores every conversation on your scorecard and turns the findings into coaching.
Five Things You Can Do This Week
You don’t need a six-month transformation plan. Start here.
1. Audit your actual coverage. Pull the numbers. How many interactions happened last month? How many did QA review? Divide. If the answer is under 10%, you have a gap that’s larger than your QA team can see.
2. Map your compliance exposure. List every regulatory requirement that touches customer conversations. PCI DSS. CFPB. FCA Consumer Duty. State-level privacy laws. For each one, answer: “If a regulator asked for evidence of monitoring, what would we produce?” If the answer is “a spreadsheet of 200 scored calls out of 15,000,” that’s your risk.
3. Calculate the cost of manual QA. Take your QA team headcount, fully loaded compensation, and divide by calls reviewed. Compare that per-evaluation cost against automated alternatives at $0.30-0.50 per interaction. The business case usually writes itself.
4. Run a 30-day pilot on 100% of calls. Most automated QA platforms can deploy alongside existing tools without disrupting operations. The pilot data alone is valuable. You’ll see patterns in the first week that your QA team has never surfaced.
5. Connect QA data to one business metric. Pick one: CSAT, first-call resolution, agent attrition, or revenue per call. Track the correlation between QA scores and that metric for 90 days. This is how you build the executive case for full investment. Not with vendor promises. With your own data.
The 98% of conversations nobody hears aren’t silent. They’re full of signals. Compliance violations accumulating. Revenue opportunities passing by. Agents developing habits that will cost you in six months. The only question is whether you’re listening.