The two percent problem in call center quality assurance

What a ten-call sample can actually tell you about an agent, worked out properly

It is the last Thursday of the month. A team lead opens the queue, filters to one agent, and picks calls until there are ten of them. Six minutes each, a rubric with fourteen lines, a score at the end. Then the same for the next agent, and the next. By Friday there are forty scores in a spreadsheet, and on Monday somebody will be coached against one of them.

Here is the finding. If you review ten of an agent's five hundred monthly calls and eight of them pass, the honest statement of what you know is that the agent's true pass rate is somewhere between 50 and 96 percent. That is not a quality score. It is a range so wide that it contains both your best performer and the person you are about to put on a plan. Every downstream use of that number, the leaderboard, the coaching session, the bonus, inherits the range even though nobody ever writes it down.

Manual quality assurance samples because manual quality assurance has no choice. A reviewer gets through eight to ten calls in a day. Published benchmarks put typical coverage at 1 to 3 percent, and broader estimates at 2 to 5 percent, which works out at roughly four to eight reviewed calls per agent per month. This article is about what that sample can and cannot support, and the arithmetic is less forgiving than most QA programmes assume.

What does a ten-call sample actually tell you?

Take an agent who handles 500 calls a month. At 2 percent coverage you review 10. Suppose 8 pass.

The observed score is 80 percent. The 95 percent confidence interval around it runs from 50 percent to 96 percent. Drop to 7 out of 10 and the interval runs from 39 to 91. Even a perfect 10 out of 10 only tells you the true rate is somewhere above 78 percent.

And 10 is the generous case. At the coverage most contact centres actually run, four to eight calls, it gets worse. Six passes out of eight gives you a range of 41 to 94 percent. Three out of four gives you 28 to 97 percent, which is very nearly the entire scale.

Calls reviewedObserved scoreWhat you can actually say (95%)
475 percentSomewhere between 28 and 97 percent
580 percentSomewhere between 37 and 98 percent
875 percentSomewhere between 41 and 94 percent
1080 percentSomewhere between 50 and 96 percent
10100 percentSomewhere above 78 percent

To get an agent's score to plus or minus five points, which is roughly the precision people assume they already have when they compare two agents on a dashboard, you need about 196 calls. That is 39 percent of the month. Not two percent, and not something a human QA team is ever going to do.

Can a sample tell two agents apart?

This is the question that matters most, because almost nothing in a QA programme uses a score on its own. Scores get ranked, compared, and turned into who gets coached.

Take two agents. One is genuinely good, at 85 percent. One is genuinely struggling, at 70 percent. A fifteen-point gap in true quality, which is large: this is the difference you built the programme to find.

Review ten calls each. The sample puts the struggling agent ahead of the good one 13.5 percent of the time, and ties them another 15.6 percent. So on 29 percent of months, a ten-call sample either misranks them or cannot separate them at all. At the more common five-call coverage it is 42 percent, which is close to a coin toss.

Put the other way round: the chance that ten calls correctly detects a fifteen-point quality gap is about 13 percent. You would need around 200 calls per agent per month to reach 95 percent.

A leaderboard built on a two percent sample is not a ranking of your agents. It is a ranking of your agents plus a large amount of noise, and the noise does not average out, because the coaching decision gets made monthly on that month's sample.

What happens to an agent whose work never changes?

This is the part QA programmes rarely model, and it is the one with a human cost.

Take one agent whose true quality is fixed at 80 percent and never moves. Same person, same skill, every month, reviewed on ten calls.

In 90 percent of months their observed score lands somewhere between 60 and 100 percent. Nothing about their behaviour changed. That spread is the measuring instrument, not the agent.

In any given month, there is a 12 percent chance they score 60 percent or below. Over twelve months, the probability that this happens at least once is 79 percent.

So roughly four out of five agents who are performing perfectly consistently will have at least one bad month a year that was manufactured entirely by which calls got pulled. Someone will sit down with them about it. They will be told to focus on something. If their score recovers the next month, which it very probably will, regression to the mean gets recorded as coaching working.

What does a two percent sample never catch?

Compliance is where the arithmetic gets genuinely uncomfortable, because rare events are exactly what sampling is worst at.

Suppose an agent mishandles something on one percent of their calls. Five bad calls hidden in five hundred. Not a catastrophic rate, and precisely the kind of thing a QA programme exists to surface.

Review ten calls at random. The probability you catch even one of the five is 9.6 percent. You miss all of them 90 percent of the time. Run that for six months and there is still a 54 percent chance you have never seen a single one. Run it for a year and it is 30 percent.

To be 95 percent confident of catching a one percent problem you would need to review 225 calls, which is 45 percent of the month.

The uncomfortable implication: for any behaviour rarer than about one call in twenty, a two percent sample is not monitoring. It is waiting for a complaint.

The sample is not even random

Everything above assumes the ten calls were drawn at random. They almost never are, and which calls get chosen turns out to encode more policy than the sample size does.

Reviewers pull the calls that scheduling allows: a particular shift, a particular queue, the start of a week. They pull escalations because escalations are flagged. Industry write-ups on this are consistent that manual sampling skews toward the easy calls and the escalated ones, and that reviewers unconsciously avoid the difficult middle, where the systemic patterns actually live. One vendor analysis reports that teams moving to full coverage typically find their scores drop by eight to twelve points, which is a claim from a company selling automated QA and worth treating as directional rather than exact, but the direction is the point.

Selection bias does not widen the confidence interval. It moves the centre of it, and it moves it in a direction you cannot measure from inside the sample. That makes it worse than the noise, because noise at least announces itself.

Why not just review more calls?

Because the arithmetic on headcount is as unforgiving as the arithmetic on samples.

Reviewing a six-minute call properly, with scoring, takes around twelve minutes. For an operation of 100 agents at 500 calls a month, full coverage is 10,000 QA hours a month, which is about 63 full-time reviewers for 100 agents. At 400 agents it is 250 reviewers.

Nobody is going to staff that, and nobody should. This is why the two percent number has been stable for two decades: it is not a standard anyone chose, it is the largest number a human team can reach. The rubric, the calibration sessions, the monthly cadence, all of it was designed around a constraint, and then the constraint quietly became the method.

That is the actual claim of this article. Call sampling was never a measurement decision. It was a staffing decision that everyone stopped questioning, and the statistics were never really on its side.

What changes when the sample is all of it

When every conversation is scored against the same criteria, three specific things change, and it is worth being precise rather than triumphant about them.

The confidence interval collapses. Not because the scoring is smarter, but because n went from 10 to 500. A score becomes a measurement instead of an estimate.

Selection bias disappears, because there is no selection. The middle of the distribution, where the patterns are, gets looked at for the first time.

Rare events become findable. A behaviour occurring on one percent of calls appears five times a month per agent instead of once every ten months.

What does not change, and this is the part vendors skip: an automated score is only as good as the criteria you wrote, and it can be confidently wrong in ways a human reviewer would catch instantly. Full coverage removes the sampling problem. It does not remove the rubric problem, and it introduces a new one, which is that nobody reads a score they did not have to earn. Any system that scores everything needs a human review path and an override, with the reviewer and the timestamp recorded, or you have replaced a noisy measurement with an unaccountable one.

We build Harmony for the first part of that. You write the criteria once in plain language, they run on every conversation rather than a sample, every answer clicks back to the second in the transcript that produced it, and a human can override any score with the override recorded. Rules can be marked skippable so an agent is never penalised for a call where the situation never arose. We are not going to tell you it improves quality by some percentage, because we have not verified a number we would stand behind, and this article is about not making claims the sample cannot support.

How we calculated this

Every statistic in this article about what a sample can tell you is our own calculation, not an industry benchmark. Here is the method, so you can check it.

We model an agent handling 500 calls per month, with QA reviewing a fixed number at random. Pass or fail on a call is treated as a Bernoulli trial and an agent's quality as a fixed underlying pass rate. Confidence intervals are Jeffreys intervals on a binomial proportion at 95 percent. The ranking figures come from the full joint distribution of two independent binomials, summing the probability mass where the weaker agent's observed count meets or exceeds the stronger one's. Rare-event detection uses the hypergeometric distribution, sampling without replacement from 500 calls containing 5 bad ones. The month-to-month spread is the 90 percent binomial interval at a fixed true rate of 80 percent, and the twelve-month figure assumes months are independent.

Two limitations worth naming. Real quality is not a fixed constant, so treating it as one is a simplification that flatters the sample rather than the argument. And real QA scores are usually weighted percentages rather than pass or fail, which changes the exact intervals but not their order of magnitude.

The coverage benchmarks cited (1 to 3 percent, 2 to 5 percent, four to eight calls per agent per month, eight to ten calls per reviewer per day) come from published industry material read on 8 September 2026, most of it produced by companies that sell automated QA. We report them as a range for that reason. The 12 minutes per review figure is our own estimate for a six-minute call plus scoring, and is stated as an assumption rather than a benchmark.

Frequently asked questions

What percentage of calls should a call center quality assurance programme review?

The honest answer is that any percentage a human team can sustain is too small to support the conclusions QA programmes draw from it. Published benchmarks put manual coverage at 1 to 5 percent. To estimate a single agent's score to plus or minus five points you need roughly 196 calls a month, and to be 95 percent confident of catching a problem occurring on 1 percent of calls you need about 225. Both are close to half the month.

Is a 2 percent sample statistically valid?

It is valid for describing a whole operation, where the sample size is agents multiplied by calls and can run into thousands. It is not valid at the level almost everyone uses it, which is the individual agent, where two percent means four to ten calls and the confidence interval spans forty to seventy points.

Why do QA scores jump around month to month?

Usually because of the sample, not the agent. An agent whose true quality is a constant 80 percent will land anywhere between 60 and 100 percent in 90 percent of months on a ten-call review, and has a 79 percent chance of at least one month at 60 percent or below over a year. If scores move without a matching change in behaviour, the measurement is the most likely explanation.

Does automated QA remove sampling bias?

It removes selection bias, because there is nothing left to select. It does not remove rubric error: a badly written criterion applied to 100 percent of calls is now wrong 100 percent of the time instead of 2 percent of the time. Full coverage raises the stakes on the scorecard itself, which is why a human review and override path matters more, not less.

How many QA reviewers would it take to review every call?

For 100 agents handling 500 calls a month, at roughly 12 minutes per review, about 10,000 hours a month, or 63 full-time reviewers. For 400 agents it is 250. This is the reason sampling exists, and it is a staffing constraint rather than a methodological choice.

Veja o Harmony nas suas próprias reuniões

O Harmony transforma as conversas da empresa em trabalho pronto.

Feito para o trabalho depois da call, não só para a gravação. Traga uma reunião real e veja o follow-up, a atualização do CRM e as tarefas ficarem prontos antes de você fechar a aba.

Etapa 1 de 4

Quantas calls e reuniões por semana?