The engine that wins in German comes ninth in Korean
Korean has the widest spread between the best and worst machine systems of any language measured. If you standardised on one engine across your whole language set, this is the pair where that decision is most likely to be costing you.
The evidence
Two findings, and together they make the strongest engine-selection argument of any language we work in.
In the 2025 WMT evaluation, professional human translators tied for first place into Korean, with the best machine systems in second. As in Arabic and Japanese, the human still wins.
In a separate 2025 benchmark measuring how systems handle names, brands and cultural references, the system that ranked first overall — first in German, French, Italian and Chinese — ranked ninth in Korean. The researchers' conclusion was blunt: there is no universal solution.
That second finding is the one worth acting on. Almost every organisation picks an engine on the basis of how it performs in the languages its staff can read, then applies that choice to Korean, where nobody in the building can check it.
Why Korean is hard for machines
Speech levels are grammar, not style. Korean encodes the relationship between speaker and listener in verb endings — several distinct levels of formality, chosen according to relative seniority, age, setting and familiarity. English source text contains none of this information. The system guesses, and a wrong guess is not a slightly awkward sentence; it is a social error that a Korean reader registers immediately.
Honorific marking runs through the whole sentence. Subject honorifics, object honorifics and verb endings have to agree with each other and with the situation. Systems get part of the chain right and part wrong, producing output that is internally inconsistent in a way that is invisible to anyone reading a back-translation.
Named entities and transliteration are unstable. This is what the entity benchmark measured, and where the spread between systems was widest. Brand names, product lines, personal names and loanwords each have conventions, and the conventions are not always what the engine produces.
Word spacing and particles carry meaning that shifts with small changes, and errors here are immediately obvious to native readers while being entirely invisible to a reviewer working from English.
What we do
- We agree the speech level before we start, based on audience and setting rather than letting the engine choose.
- We score honorific and speech-level errors as their own category, separate from accuracy, with their own severity scale — because these are the errors that damage a brand without ever being technically "wrong".
- We hold entity and transliteration decisions across the document set and maintain them as a termbase you keep.
- We can benchmark engines on your Korean content specifically, which given the evidence above is the single highest-value hour you can spend on this pair.
Where Korean work is worth most
Korean crossed with technical and regulated content: automotive, electronics and semiconductor documentation, medical devices, patents. Korea is a major filing jurisdiction and a major manufacturing market, and the documentation volumes are large and consequential.
Published 2026 rate data puts English↔Korean second only to Japanese among the major pairs, at roughly one and a half to two times English↔Spanish. The cited reason is translator supply.