You may know this as engine benchmarking, MT evaluation, engine selection, or AI comparative analysis.
What it is
We take a real sample of your own content — not a benchmark set — and run it through your current engine and three or four alternatives. The same reviewers score every output against the same scorecard, blind to which engine produced which file. You get a comparison and a recommendation, language by language.
Why it is worth doing
Most organisations chose an engine once, standardised on it, and have never re-checked. Published comparisons of engine performance on real client projects show different engines winning different languages, and the overall best-performing engine losing outright in several — with gaps between the best and second-best choice large enough to decide whether output is publishable or not.
We would ask you to treat those published comparisons the way we do: as one dataset rather than as settled fact. They are usually published by companies with something to sell, several of their language samples are small, and much of the scoring is automatic — which, in some languages, is exactly the thing that cannot be trusted. That is the argument for measuring on your content rather than reading somebody's table, including ours.
What you get
- A scored comparison across the engines tested, per language, using the method on how we score.
- A recommendation per language, with the size of the difference stated, so you can judge whether it justifies a change.
- The error patterns each engine produces on your material — which is often more useful than the ranking, because it tells you what to fix in the glossary or the prompt rather than which vendor to switch to.
- Where relevant, a view on whether your quality-estimation thresholds are set correctly for each language. They are frequently calibrated on European results and applied everywhere.
Our position
We do not sell a translation engine, we do not resell anybody's licences, and we take no referral fee from any vendor named in a report. We have no answer to defend. If the finding is that your current engine is the right one, that is the report you get.
What it costs
A fixed fee per study, agreed in advance from the number of languages and the size of the sample.
Delivered, or staffed
Send us the work and we return the deliverable described above. Or place our reviewers into your own process and tooling and run them under your rubric, your schema and your brand. Same people, same documentation, different contract. Tell us which you want and we will price it that way.