The engine that wins in German comes ninth in Korean

Korean has the widest spread between the best and worst machine systems of any language measured. If you standardised on one engine across your whole language set, this is the pair where that decision is most likely to be costing you.

--:--:-- [PST]
1st → 9th
where the top-ranked overall system placed in German versus in Korean
Entity benchmark, 2025
Human first
professional translators still tied for first place into Korean, ahead of every machine system
WMT25
S01

The evidence

Two findings, and together they make the strongest engine-selection argument of any language we work in.

In the 2025 WMT evaluation, professional human translators tied for first place into Korean, with the best machine systems in second. As in Arabic and Japanese, the human still wins.

In a separate 2025 benchmark measuring how systems handle names, brands and cultural references, the system that ranked first overall — first in German, French, Italian and Chinese — ranked ninth in Korean. The researchers' conclusion was blunt: there is no universal solution.

That second finding is the one worth acting on. Almost every organisation picks an engine on the basis of how it performs in the languages its staff can read, then applies that choice to Korean, where nobody in the building can check it.

S02

Why Korean is hard for machines

Speech levels are grammar, not style. Korean encodes the relationship between speaker and listener in verb endings — several distinct levels of formality, chosen according to relative seniority, age, setting and familiarity. English source text contains none of this information. The system guesses, and a wrong guess is not a slightly awkward sentence; it is a social error that a Korean reader registers immediately.

Honorific marking runs through the whole sentence. Subject honorifics, object honorifics and verb endings have to agree with each other and with the situation. Systems get part of the chain right and part wrong, producing output that is internally inconsistent in a way that is invisible to anyone reading a back-translation.

Named entities and transliteration are unstable. This is what the entity benchmark measured, and where the spread between systems was widest. Brand names, product lines, personal names and loanwords each have conventions, and the conventions are not always what the engine produces.

Word spacing and particles carry meaning that shifts with small changes, and errors here are immediately obvious to native readers while being entirely invisible to a reviewer working from English.

S03

What we do

  • We agree the speech level before we start, based on audience and setting rather than letting the engine choose.
  • We score honorific and speech-level errors as their own category, separate from accuracy, with their own severity scale — because these are the errors that damage a brand without ever being technically "wrong".
  • We hold entity and transliteration decisions across the document set and maintain them as a termbase you keep.
  • We can benchmark engines on your Korean content specifically, which given the evidence above is the single highest-value hour you can spend on this pair.
S04

Where Korean work is worth most

Korean crossed with technical and regulated content: automotive, electronics and semiconductor documentation, medical devices, patents. Korea is a major filing jurisdiction and a major manufacturing market, and the documentation volumes are large and consequential.

Published 2026 rate data puts English↔Korean second only to Japanese among the major pairs, at roughly one and a half to two times English↔Spanish. The cited reason is translator supply.

S05

Find out how your engine performs in Korean

It takes one batch of representative content. We will score the output of more than one system on the same framework and show you the difference.

Ask for an engine comparison