Multilingual call QA: seven languages, scored the same way

Why the same rubric produces different scores in different languages, and what to do about it

Seven country leads, seven dashboards, one number each. Poland is at 82. Portugal is at 74. The quarterly review is on Thursday and somebody is going to ask why Portugal is eight points down. Nobody in the room can say whether Portugal is worse, or whether Portugal was scored by a stricter reviewer against a rubric that lost something in translation eighteen months ago.

Here is the short answer. In most multilingual QA programmes the scores are not comparable across languages, and the dashboard that puts them side by side is the thing creating the problem. Four separate mechanisms push the same rubric apart when it crosses a language boundary: the translation of the criteria, the local norms of the conversation, the absence of cross-language calibration between reviewers, and the fact that the models doing the scoring are measurably better in English than in your smaller markets. Fixing it is mostly method, not technology, and the method is specific enough to write down.

Why does one rubric produce different scores in different languages?

The rubric was translated, and translation moves the criterion

A criterion like "the agent acknowledged the customer's frustration before proposing a solution" is precise in English and slippery everywhere else. Translated into Polish it may become a test of a particular polite formula. Translated into Brazilian Portuguese it may become a test of warmth. The words survive the translation. The behaviour being measured does not.

This happens once, usually at rollout, usually by whoever was available, and then it is invisible for years because each country only ever sees its own version.

Local norms are not scoring criteria, but they get scored

Directness reads as efficiency in one market and rudeness in another. The length of an opening, the use of a first name, how much hedging precedes bad news: all of these vary by market and none of them is a quality difference. When a rubric contains anything tonal, it is measuring the local norm at least as much as the agent.

Reviewers calibrate within a language and never across

Good QA programmes run calibration sessions where reviewers score the same call and argue about the difference. Almost all of them run those sessions inside one language, because that is who is in the room. The Polish reviewers converge with each other. The Portuguese reviewers converge with each other. The two groups drift apart and nothing in the process is designed to notice.

The model is genuinely worse in your smaller languages

If any part of the scoring is automated, this one is measurable and it is larger than most buyers expect.

Research from Oracle AI on enterprise deployments reports accuracy drops of up to 29 percent in non-English languages compared to English on semantically identical content, and finds that retrieval-augmented setups do not close the gap, because the underlying reasoning is English-centric. Academic benchmarks show the same shape. On MMLU-ProX, leading models score above 70 percent in English and around 40 percent in Swahili. On a clinical question set, GPT-4 scored 63.4 percent in English, 51.8 percent in Filipino and 50.6 percent in Hindi. On a parallel Irish and English benchmark, the best model reached 76.2 percent in English and 55.8 percent in Irish.

Two caveats, because this is the number most likely to be misquoted. These are general task-accuracy benchmarks, not measurements of call scoring, and the gap between a major European language and English is much smaller than the gap to a low-resource language. But the direction is consistent across every benchmark anyone has run, and it means the score in your smallest market is your least reliable score, which is exactly backwards from how much scrutiny it usually receives.

Code-switching makes it harder again. Published work on multilingual recognition reports language identification error rising from about 2 percent when an utterance is in one language to about 8.5 percent when one to three languages appear inside a single utterance. Agents in Warsaw, Lisbon and Bangalore switch mid-sentence constantly.

What actually breaks

Not the scores. The comparison.

A per-language score used inside its own language is fine. Polish agents ranked against Polish agents, coached by a Polish team lead, against a rubric that means one thing in Polish. That works.

The failure begins the moment those numbers land in one table. The table implies a comparison it cannot support, and the comparison then drives things that matter: which market gets the headcount, which team lead is under pressure, which agent is on a plan. Nobody decides to compare across languages. The dashboard just puts the numbers next to each other and the human brain does the rest.

The second failure is quieter. Because scores are not trusted across markets, they stop being used across markets, and the whole programme collapses into seven local exercises. The pattern that shows up in three countries at once is the most valuable thing a multilingual operation could find, and it is precisely what a set of incomparable local scores can never surface.

How do you score seven languages the same way?

Six practices. None of them requires a specific vendor, and the first two do most of the work.

1. Write the criteria once, in one language, and never translate them. Keep a single canonical rubric. If the scoring is automated, the criteria stay in the language you wrote them in while the conversation stays in the language it happened in. Separating the question language from the transcript language is the single change that removes the translation drift, because the criterion is only ever expressed once.

2. Never translate the call. Score the source. A translated transcript has already lost hedging, register and the exact words that a criterion about acknowledgement or consent depends on. Translate for a human who needs to read it later if you must, but score the original.

3. Write behavioural criteria, not tonal ones. "Did the agent state the resolution timeframe?" travels across every language you operate in. "Was the agent empathetic?" does not, and in the EU it also runs into the AI Act's restrictions on inferring emotional state at work. The test for a good multilingual criterion: could two reviewers who share no common language both point at the same moment in the call as the evidence?

4. Make situational rules skippable. If a criterion about objection handling is scored zero on every call where no objection arose, markets with different call mixes get different scores for reasons that have nothing to do with quality. A rule that can be marked not applicable, and excluded rather than failed, removes a whole class of false cross-market gaps.

5. Run one calibration set across all languages. This is the practice almost nobody runs and it is the one that makes the numbers comparable. Take twenty calls per language. Have the local reviewers score them and have the automated system score them. Then measure two things: agreement between reviewers within each language, and agreement between human and system per language. You now have a per-language reliability figure, and you will find it is not the same everywhere.

6. Publish that figure next to the score. If Portuguese human-to-system agreement is 91 percent and Polish is 78 percent, an eight-point gap between the two markets is not a finding. Reporting reliability alongside the score is what stops a dashboard from implying a precision it does not have. Re-run the calibration quarterly, because models change under you.

Where Harmony fits

Two of the six practices above are product decisions rather than process ones, and they are the two that no amount of QA discipline can fix on its own.

Question language and transcript language are set independently. You write one canonical set of criteria, in English or Polish or whatever your head office runs on, and they evaluate against conversations in Portuguese, Spanish or Japanese without the criteria ever being translated per market. One rubric, seven languages, no drift, because there is only one version of the criterion in existence. That is practices one and two, solved at the product level rather than by asking seven country leads to be disciplined.

Transcription runs in over 100 languages with automatic detection, including code-switching inside a single sentence. This matters more than a language count suggests. Supporting a hundred languages one at a time is a different capability from handling two of them in one utterance, and an agent's mid-sentence switch to English is exactly where a transcript quietly degrades and takes every downstream score with it.

The rest is what makes the scores usable rather than merely produced.

A standing question set, not a report you request. You write the questions once, in plain language, as though briefing a very good analyst. They then run on every conversation from that point forward, in every market, without anyone rerunning anything, and one question set covers all seven countries rather than being rebuilt per country.

Every answer clicks back to the second that produced it. A country lead who disputes a score does not argue with a number, they open the moment in the transcript, in their own language. This is the single fastest way to end a cross-market dispute, and it is also what makes the calibration exercise in practice five cheap to run rather than a project.

Required and skippable rules. A criterion about objection handling can be excluded on calls where no objection arose, rather than scored zero. Without that, markets with different call mixes get different scores for reasons that have nothing to do with quality, which is one of the most common false cross-market gaps.

Human override, recorded. Weekly cycles give each person an average, median and spread with every scored conversation listed, and a reviewer can override any score with the reviewer and timestamp captured.

No sentiment analysis, anywhere. We do not infer emotional state, from voice or otherwise, on agents or on customers. That is a deliberate product decision and it happens to solve the hardest problem in multilingual QA: emotional register is the single least portable thing across languages, so a tonal criterion is guaranteed to measure the local norm rather than the agent. It also means the EU AI Act's Article 5 prohibition on inferring worker emotions, in force since February 2025, is not a question we have to answer with a caveat.

Capture is not the constraint. Companion joins from a pasted link with no install and no calendar connection, and it works on conversations it never recorded: upload a back catalogue, paste a transcript, connect IVR and work phones. For a multilingual operation this matters because your seven markets almost certainly do not run on one telephony stack.

What we still cannot tell you

We do not publish per-language accuracy figures for our own scoring, because we have not run the calibration study that would let us stand behind them. That is a real gap, and it is the same gap every vendor in this category has: nobody publishes per-language reliability, and everybody sells to multilingual operations.

Until someone does, practice five above is not optional and no product feature replaces it. Run the calibration yourself, on your own rubric, in your own markets, and treat any vendor's single global accuracy claim as an English number until they show you otherwise. That includes ours.

We are also behind on certifications relative to the incumbents: SOC 2 Type I is complete and ISO 27001 is in progress as of September 2026.

How we checked this

The cross-language accuracy figures come from published research read on 8 September 2026: an Oracle AI paper on multilingual consistency in enterprise applications for the 29 percent figure, MMLU-ProX for the English-to-Swahili comparison, a multilingual clinical question benchmark for the GPT-4 figures, and a parallel Irish-English reasoning benchmark for the Irish figure. The code-switching language identification figures come from published work on multilingual end-to-end recognition.

All of these measure general task accuracy, not call-scoring accuracy. We use them to establish that a cross-language gap exists and is consistent in direction, not to predict the size of the gap in any particular QA deployment. Anyone who quotes the 29 percent figure as a call-scoring number is quoting it wrongly, including us if we ever do it.

We make Harmony, which sells into multilingual operations, so the practices above are ones we benefit from you adopting. Four of the six can be run with a spreadsheet and no vendor at all.

Frequently asked questions

What is multilingual call QA?

Scoring customer conversations against the same quality criteria across more than one language, in a way that makes the resulting scores comparable between markets. The second half is the hard part. Most programmes achieve consistent scoring inside each language and comparable scoring across none of them.

Can you use the same QA scorecard in every language?

You can and you should, but only if the scorecard itself is never translated per market and the criteria are behavioural rather than tonal. A criterion about whether a resolution timeframe was stated travels. A criterion about warmth or empathy measures the local norm, and in the EU it also runs into the AI Act's limits on inferring emotion at work.

Should call transcripts be translated into English before scoring?

No. Translation strips hedging, register and the specific wording that criteria about consent, acknowledgement or commitment depend on. Score the conversation in the language it happened in and keep the criteria in one canonical language instead.

Are AI QA scores less accurate in non-English languages?

Published benchmarks consistently show large language models performing better in English than in other languages on the same task, with the gap widening as a language gets less well represented in training data. Reported drops range from a few points for major European languages to more than thirty points for low-resource ones. No vendor in this category currently publishes per-language accuracy for call scoring, which is why running your own calibration set is the only way to know your figure.

How do you compare QA scores between countries fairly?

Publish a per-language reliability figure alongside each score, produced by having both local reviewers and the automated system score the same twenty calls per language. A gap between two markets only means something if it is larger than the measurement difference between them. Without that figure, a cross-country comparison is a guess with a decimal point on it.

What about agents who switch languages mid-call?

Ask specifically about code-switching rather than about the language list. Supporting 100 languages one at a time is a different capability from handling two of them in one sentence, which is what actually happens in Warsaw, Lisbon and Bangalore. Published work shows language identification error rising roughly fourfold when multiple languages appear inside a single utterance.

See Harmony on your own meetings

Harmony turns enterprise conversations into finished work.

Built for the work after the call, not just the recording. Bring one real meeting and watch the follow-up, the CRM update, and the action items land before you've closed the tab.

Step 1 of 4

How many calls and meetings a week?