What you get
Qualified native speakers judging model output against criteria you define, with the judgement recorded: ratings, rankings, or pass-fail against a rubric, delivered in your schema, with reviewer identity and qualifications attached to every judgement.
Three kinds of work
Why language is the hard part
Automatic metrics are unreliable in exactly the languages where you most need to know.
A 2026 study evaluating four frontier models on Hausa and Fongbe, ten thousand sentences each, found human ratings of 4.0 to 4.5 out of 5 for Hausa and 1.0 to 2.2 for Fongbe. For Fongbe, one standard metric ranked the worst-performing system best. Another returned near-identical scores for every text, because the model underneath could not tell two texts apart. For Hausa, every automatic metric picked one system while human judges preferred a different one. The researchers' own conclusion was that human evaluation is mandatory for these languages.
A separate 2025 audit of the 200-language benchmark most multilingual claims rest on found its own reference translations falling well below the quality standard claimed for them, and showed a model scoring 13.95 on the benchmark and 2.29 on realistic text while a better model scored 4.87 and 13.40. The benchmark ranked the worse model higher.
If your evaluation coverage is thinnest where your metrics are least trustworthy, that is not a measurement problem. It is a staffing problem.
Adversarial testing
We also take red teaming engagements in languages where English-first safety evaluation does not transfer — where the failures are cultural rather than linguistic, and a translated prompt list does not find them. Arabic makes the point: there is no single Arabic to test in, and a suite written in one variety tells you little about the others.
This is a different discipline from linguistic review and we scope it as one, in conversation rather than off a service page. We are not a general AI safety consultancy. What we bring is the part that is hardest to buy — native speakers who can tell you why a response that looks fine is not — and we brief reviewers on what an engagement involves before they accept it.
How we staff it
From a network of linguists we know personally, whose subject expertise and qualifications are on record, in named languages. Not an open crowd. For work of this kind the difference shows up as consistency between reviewers, which we measure and report rather than assert.
Reliability
We calibrate reviewers against a shared rubric before production work begins and report inter-rater agreement with the delivery. Where agreement is low, that is a finding about the rubric as often as about the reviewers, and we will say so.
What it costs
Per task or per hour, depending on the shape of the work. Volume and turnaround are agreed per language, because our answer is different in Japanese than it is in Fongbe and we would rather tell you that up front.