Call sampling is a decision, not a method

Every QA programme has a policy about whose work gets looked at. Almost none of them have written it down.

Thursday afternoon. A QA lead has ninety minutes and four agents to get through. She opens the recordings list, filters to last week because this week is not fully processed yet, sorts by duration because the very short ones are hang-ups, skips the two flagged escalations because those are already being handled elsewhere, and takes the next eight that will fit before the end of the day.

That sequence took her about forty seconds and she would not describe any of it as a decision. It is the decision. Everything the QA programme will conclude about those four agents this month was determined in those forty seconds, by a set of rules nobody wrote down, nobody reviews, and nobody could reproduce.

Here is the argument. Call sampling is treated as a neutral technique, a practical way to approximate a whole from a part. It is not. Every sampling scheme is a policy about who bears scrutiny and which parts of your operation are allowed to go unobserved, and in most contact centres that policy is an accident of scheduling rather than a choice anyone made. The size of the sample is a separate problem, and a well-documented one. This is about the selection.

Who actually chose two percent?

Nobody did.

A QA analyst gets through eight to ten calls in a working day when the review is done properly. Published benchmarks put manual coverage somewhere between 1 and 5 percent of interactions, which works out at roughly four to eight reviewed calls per agent per month. That number has been stable for two decades, and it is stable because it is a division problem: reviewer hours divided by review time.

What happened next is the interesting part. The constraint did not stay a constraint. It became a method. The monthly cadence, the rubric length, the calibration session, the coaching conversation built around a specific call: all of it was designed around how many calls a human could get through, and then the whole apparatus acquired the vocabulary of measurement. People now say "our QA sample" the way an epidemiologist would, and mean something with none of the same properties.

If you want the arithmetic on what a sample that size can support, we worked it out separately in the two percent problem. The short version is that at four to ten calls per agent the confidence interval is wide enough to contain both your best and worst performer. This article assumes that and asks a different question: given that you are looking at eight calls, which eight, and who decided?

What is your sampling scheme actually selecting for?

Four schemes cover almost everything running in production. Each one encodes a different, usually unstated, policy.

Convenience. The calls that scheduling allowed. A particular shift, a particular queue, the start of a week, the ones already processed by Thursday. This is by far the most common scheme and almost nobody names it, because naming it would require admitting it exists. What it encodes: agents who work the hours your QA team works are observed. Agents on the night shift, on the overflow queue, or in the market three time zones away are not, or are observed by someone else with a different rubric.

Escalation-triggered. Review the calls that already went wrong. Efficient, and completely circular: you are studying failures selected on the basis of having failed. This tells you a great deal about complaints and nothing at all about the ordinary calls where quality actually lives. Industry write-ups on manual QA are consistent that sampling skews toward the easy calls and the escalated ones, and that the systemic patterns sit in the middle of the distribution, in the calls nobody flagged.

Targeted. Review the people you are already worried about. This is a legitimate management activity and a catastrophic measurement one, because it produces exactly the finding it was designed to look for. More on the loop below.

Random. The only scheme with defensible statistical properties and the rarest one in practice. Genuine randomness means reviewing the 4am call, the one in the language your reviewer does not speak, and the twelve-second hang-up, and it means doing so when you would rather be looking at something else.

Most programmes run a blend, undocumented, and report one number at the end of it.

The contradiction inside the standard advice

Open any contact centre QA best-practice guide and you will find a version of this instruction: randomise your call selection, and also make sure to include high-risk or flagged calls for focused review.

Both halves are sensible. Together they are incoherent, because they produce two different objects. A random sample supports an estimate of an agent's typical performance. A targeted set supports an investigation of a specific risk. Mixing them produces something that is neither, and then the result gets reported as a single percentage score, put on a dashboard, and compared with last month.

This is the most common unforced error in QA design and it is entirely fixable. The two jobs are different jobs. Assurance asks "what is normally happening here." Investigation asks "did this specific thing go wrong." They need different call selection, different rubrics, and above all different reporting, because the moment they share a number the number means nothing.

The loop nobody designed

Targeted sampling has a property that makes it worse than merely imprecise: it confirms itself.

Suppose an agent has a poor month, for real reasons or for the sampling reasons described in the other article. They are now flagged. Being flagged means more calls get reviewed. More reviewed calls means more findings, because findings scale with looking. More findings confirms the flag.

Recent academic work on fairness in contact centre QA systems names the mechanism directly: an agent's past performance can anchor future judgments, producing a negative halo, and role metadata such as trainee versus senior specialist can create undue scrutiny or undue leniency that distorts comparisons across levels. The paper studies model-based scoring, but nothing in that mechanism requires a model. Human reviewers have been doing it for thirty years.

The same paper notes a second contamination worth naming: customer profile and the difficulty of the situation bleeding into the agent's score, so that an agent who handles hard calls looks worse than one who handles easy ones. If your sampling scheme is not controlling for call difficulty, and almost none are, then agents on the hardest queue carry a permanent penalty that no coaching will fix, because it was never about them.

New joiners get the worst of this. They are watched more closely, which is reasonable, and they therefore accumulate more findings than an experienced agent doing the same work, which is not.

Why agents do not trust the score

Ask anyone who has been on the receiving end and the objection is never statistical. It is always the same sentence: you picked the wrong call.

Sometimes that is defensiveness. Often it is correct, and the QA programme has no way to tell the difference, because it cannot show what the other 492 calls looked like. A score derived from eight unexplained selections is not evidence, it is an assertion with a decimal point, and agents work this out quickly. The documented consequences are the ones you would expect: scoring is perceived as subjective and inconsistently applied, and some agents optimise for the parts of the rubric they know get checked rather than for the customer.

There is a version of this that vendors report from the other side, and it is the most interesting argument for full coverage that has nothing to do with accuracy. When every call is looked at, agents stop feeling singled out by an isolated negative review, because there is no selection left to dispute. The complaint changes from "you picked the wrong call" to "I disagree with that criterion," which is an argument worth having.

What sampling is actually optimising for

Here is the uncomfortable framing, and it is the one that explains why sampling survived so long without much scrutiny.

A two percent QA programme is not optimised for discovering what is happening in your operation. It is optimised for being able to demonstrate that you have a QA programme. It produces a monthly number, a coaching record, a documented process, and an artefact you can show a client or an auditor. All of that is real value, and none of it requires the number to be a good estimate of anything.

That is not cynicism about QA teams, who generally know all of this perfectly well and are working inside a budget. It is a claim about what the institution rewards. Nobody is ever asked to defend the sampling scheme. They are asked whether the reviews were completed on time.

How to sample deliberately

If you are going to sample, and most operations still have to, six things make it a decision rather than an accident. None of them costs money.

Write the scheme down. One paragraph. Which calls are eligible, how they are chosen, by whom, how often. The act of writing it usually exposes the convenience sampling immediately.

Separate assurance from investigation. Two processes, two selections, two reports. Never one score built from both.

Stratify before you randomise. Split by queue, shift, call type and duration band, then randomise inside each stratum. This is the single highest-value change available to a manual programme, because it stops the night shift and the hard queue from disappearing.

Make eligibility total. Every call in the period is eligible, including the 4am one and the twelve-second one, or your scheme is not what you think it is.

Record what was not eligible, and why. Calls still processing, calls in a language the reviewer cannot score, calls already under investigation. The exclusions are the policy. Undocumented exclusions are how convenience sampling reappears.

Report the selection alongside the score. "82 percent, from 8 calls, random within queue and shift, excluding 14 calls under separate review." That sentence is longer than a number and it is the difference between evidence and an assertion.

The decision does not disappear when the sample does

Score every conversation and the selection problem is genuinely gone. There is nothing left to choose, so there is nothing left to choose wrongly. The night shift is in. The hard queue is in. The middle of the distribution, where the patterns were always hiding, is in for the first time.

But the decision does not vanish. It moves. It stops being "which calls do we look at" and becomes "which questions do we ask of all of them," and that is a better place for it, because a question is a written artefact that somebody signed off, while a sampling scheme was a set of habits nobody could reproduce.

This is what Harmony is built around. You write the questions once, in plain language, as though briefing a very good analyst: did the agent verify identity before acting on the account, was the resolution timeframe stated, were the diagnostic steps covered. Those questions then run on every conversation from that point on, not on a sample and not just once, and they run the same way on the night shift and the overflow queue as on the calls a reviewer would have got round to.

Three things make the result usable rather than merely comprehensive. Every answer clicks back to the second in the transcript that produced it, so the argument moves from "you picked the wrong call" to a specific sentence either party can look at. Rules can be marked required or skippable, so an agent is never scored zero on a criterion for a situation that never arose, which removes the largest source of false gaps between queues. And a human can override any score, with the reviewer and the timestamp recorded, because a system that scores everything and cannot be argued with is worse than one that scores a little and can be.

We also do not infer emotional state, from voice or otherwise, which removes an entire category of the bias the fairness research describes, and which has been prohibited for workers in the EU since February 2025.

The honest version

Full coverage fixes selection bias. It does not fix the rubric, and it raises the stakes on it: a badly written criterion applied to two percent of calls is wrong occasionally, and applied to all of them is wrong constantly. Everything this article says about invisible policy applies just as much to a question set nobody reviews. If your criteria are written once at rollout and never looked at again, you have replaced an undocumented sampling scheme with an undocumented rubric, and the second one has more reach.

So the recommendation is the same in both worlds. Write down the judgment, put a date on it, and review it on a cadence. The technology changes where the judgment lives. It does not remove the obligation to make it deliberately.

How we checked this

Coverage benchmarks (1 to 5 percent of interactions, roughly four to eight calls per agent per month, eight to ten calls per reviewer per day) come from published industry material read on 8 September 2026, most of it produced by companies that sell automated QA. We report them as ranges for that reason. The observations about sampling skewing toward easy and escalated calls, and about agents feeling singled out by isolated reviews, come from the same category of source and carry the same caveat.

The fairness mechanisms described, anchoring on past performance, role-based scrutiny, and situation difficulty contaminating agent scores, come from published academic work on fairness evaluation in contact centre quality assurance systems. That work studies model-based scoring; we extend the mechanisms to human review, which is our inference rather than the paper's finding.

The statistical claims referenced in passing are worked out with their method in a separate article and are our own calculation, not an industry benchmark.

We make Harmony, which sells automated scoring, so we benefit from you concluding that sampling is a problem. Four of the six practices in the deliberate-sampling section require no vendor at all, and if you adopt only those we would still consider this page to have done its job.

Frequently asked questions

What is call sampling in a contact centre?

Selecting a subset of recorded conversations for quality review, because reviewing all of them by hand is not possible. In practice it covers 1 to 5 percent of interactions. The term borrows credibility from statistical sampling, but most operational schemes are not random and therefore do not have the properties that word implies.

How should calls be selected for QA review?

Stratify first by queue, shift, call type and duration, then randomise within each stratum, and keep every call in the period eligible. Run investigation of specific incidents as a separate process with its own report, never blended into the same score as routine assurance.

Is random call sampling actually fair to agents?

Genuine random sampling is fair in expectation but noisy in any given month, and very few programmes are genuinely random. The more common failure is convenience sampling, where the calls reviewed are the ones the QA team's schedule allowed, which systematically under-observes some shifts and queues and over-observes others.

Why do agents dispute QA scores?

Usually because the selection cannot be defended rather than because the rubric is wrong. A score built on eight unexplained choices out of five hundred calls is difficult to justify to the person it is about. Reporting the selection method alongside the score, or removing the selection entirely, changes the nature of the disagreement.

Does scoring 100 percent of calls remove bias from QA?

It removes selection bias, because there is nothing left to select. It does not remove rubric bias, and it amplifies it: a flawed criterion now applies to every conversation instead of a few. Full coverage moves the judgment from which calls you look at to which questions you ask, which is an improvement only if somebody reviews the questions.

How many calls should we review per agent?

That is a different question from this article's, and the arithmetic is unforgiving: reaching plus or minus five points on an individual agent's score takes around 196 calls a month. The practical answer for a manual programme is to sample deliberately and report the uncertainty honestly, rather than to find a number that makes the score feel solid.

Veja o Harmony nas suas próprias reuniões

O Harmony transforma as conversas da empresa em trabalho pronto.

Feito para o trabalho depois da call, não só para a gravação. Traga uma reunião real e veja o follow-up, a atualização do CRM e as tarefas ficarem prontos antes de você fechar a aba.

Etapa 1 de 4

Quantas calls e reuniões por semana?