Engines are not equally good at every language, and almost nobody has checked

We concentrate on the pairs where automated output still fails in ways that matter — and we can show you the evidence for each one.

--:--:-- [PST]
S01

The finding most companies have not acted on

In a 2025 benchmark of how well translation systems handle names, brands and cultural references, the top-performing system overall ranked first in German, French, Italian and Chinese — and ninth in Korean, eighth in Arabic. The researchers concluded there is no universal solution.

Most organisations standardise on one engine for everything, and they choose it on the basis of how it performs in the European languages they see most. That decision is then silently applied to Arabic, Korean, Thai and the Indic languages, where it may be badly wrong, and where nobody in the building can read the output well enough to notice.

S02

Difficulty does not work the way people assume

There are four different reasons a language pair resists automation, and they call for different arguments and different evidence. We have grouped our language pages accordingly.

S03
  • ### 1. Where professional humans still beat every machine tested
  • The 2025 WMT evaluation put professional translators ahead of all sixty systems tested in these pairs. This is the strongest evidence available for keeping a qualified person in the workflow.
  • Arabic — not one of sixty systems matched a human into Egyptian Arabic. The hardest major commercial language for machine translation. Japanese — humans still ahead. The failures cluster in honorifics and the speaker–listener relationship, which English source text does not encode. Korean — humans tied for first, and the widest spread between systems of any language measured. Italian — named jointly the most challenging pair in the whole evaluation, alongside Egyptian Arabic. A well-resourced European language. Worth reading if you think difficulty tracks distance from English.
S04
  • ### 2. Where the training data runs out
  • Low-resource means there is little digitised parallel text for a system to learn from. It is a statement about data infrastructure, not about speakers — Hindi has 344 million native speakers and is low-resource. This is also where vendors' language-count marketing is least trustworthy.
  • Thai — no spaces between words, no inflection, a layered register system, thin data. Indic languages — Hindi, Bengali, Tamil, Telugu, Marathi and the rest. Script rendering, honorifics and code-mixing on top of scarce data. Southeast Asian languages — Vietnamese, Indonesian, Malay, Filipino, Khmer, Lao, Burmese. African languages — Swahili, Amharic, Hausa, Yoruba, isiZulu, Somali and others. Where the gap between claimed and actual engine quality is widest.
S05
  • ### 3. Where the benchmarks mislead
  • Plenty of data, fluent output, and structural problems that more data has not fixed. These are the pairs where output passes a fluency read and fails a structured one.
  • Chinese — machines beat humans on the benchmark and score lowest of ten languages on real client projects. Both are true, and patents are where it matters. Russian — well-resourced, fluent, and still near the bottom on real commercial content. Case agreement and verbal aspect are the reasons.
S06
  • ### 4. Where the problem is consistency, not comprehension
  • Engines handle these languages well. What they cannot do is hold one term steady across a document family for a decade — which is what technical and regulated documentation actually requires.
  • German — the work is terminology discipline across large technical sets. Nordic languages — Swedish, Danish, Norwegian and Finnish, where post-editing barely saves any time at all and the standard discount is hardest to justify.
S07

How we choose what to take on

We would rather tell you a pair is well served by your current setup than sell you a review you do not need. The test we apply is whether human judgement measurably changes the outcome — because of how the language works, because of what the content is, or because of what an error would cost.

Where the answer is no, we will say so. And where we cannot staff a language to a standard we would stand behind, we will say that too rather than take the work.

S08

Find out how your engine performs on your content

Send source content in the pairs you care about. We will run it through more than one system, score the output on the same error framework, and give you the comparison.

Ask for an engine comparison