Translation quality is the question everyone asks and the least dangerous of the failures. The one that reaches closing is subtler: a clause translated perfectly, understood correctly, and then scored against the wrong jurisdiction's sense of what is serious — producing a confident low-severity finding with nothing on the page to suggest anything went wrong.
Four failures, only one of which is translation
| Failure | Visible? | Cost |
|---|---|---|
| Mistranslation | Yes — reads oddly, catches attention | Low; usually caught |
| Untranslatable concept mapped silently | No | Moderate to high |
| Wrong materiality baseline | No | Highest |
| Language not processed at all | No — looks like a clean document | High |
Three of the four are invisible in the output. That is the governing fact about multilingual review, and it should shape what you ask a vendor.
Failure 1 — Mistranslation (the manageable one)
Modern translation of contractual language is good, and errors that do occur tend to announce themselves: a clause that reads strangely gets a second look. This is a real risk but a self-revealing one.
The control is simple and worth insisting on: every finding must carry the source sentence in its original language, alongside any translation. Two benefits — a reviewer qualified in that language reads what the document actually says rather than a round-trip through English, and translation errors become visible instead of baked into the record.
Failure 2 — Concepts that do not map
Some legal concepts have no clean equivalent across systems. Civil-law good-faith obligations, security interests that perfect differently, employment protections with no analogue, corporate forms that translate to a misleadingly familiar English term.
The dangerous behaviour is silent mapping: the system picks the nearest English concept and proceeds as though the match were exact. Downstream, the clause is scored as if it were the English concept — which it is not.
The correct behaviour is to flag it for local review rather than resolve it. Ask a vendor directly what their system does here. Many map silently, and it will never show up in a demo run on English documents.
Failure 3 — The wrong materiality baseline
This is where multilingual deals actually go wrong, and it is not a language problem at all — it is a scoring problem wearing a language problem's clothes.
Severity encodes a judgment about what counts as serious, and that judgment came from some jurisdiction's norms. A provision that is unremarkable in one market can be a genuine transaction blocker in another. The model reads it correctly, extracts it correctly, quotes it correctly, and ranks it low — because the rubric is aimed at a different legal world.
Per-deal jurisdiction settings are common and structurally wrong for cross-border work. A target's customer contracts may sit under three governing laws and its employment documents under a fourth. One setting applies the wrong baseline to most of the room.
Failure 4 — The language that was never processed
The quiet one. A tool's language support is usually described as a list, not as a performance curve — but performance varies enormously across that list, and a language handled poorly may produce few findings or none.
Zero findings is indistinguishable from a clean document in almost every report format. So a data room where the Japanese documents were effectively skipped looks like a data room where the Japanese documents were fine.
Require coverage reporting broken down by language. Documents submitted, processed, failed, per language. Then reconcile it against the room before reading a single finding. This is the two-minute control that prevents the most expensive multilingual failure.
What to require
- Original-language source text on every finding, alongside translation
- Per-document jurisdiction, not per-deal
- Per-jurisdiction severity configuration — the same clause type must be able to score differently under different governing law
- Explicit uncertainty on concepts that do not map, rather than silent substitution
- Coverage reporting by language
- Named language performance — not a support list, but which languages are strong and which are weak. A vendor who claims uniform performance across twenty languages is describing a list, not a capability
Where the value actually is on multilingual deals
The bottleneck on cross-border diligence is not translation. It is local counsel time — expensive, engaged late, in another time zone, and on the critical path.
Teams manage this by sending local counsel whatever the deal team could tell was relevant, which across a language barrier is a weak filter. A screening pass that scores every document per jurisdiction changes what reaches them: a ranked set with the original-language clause attached and a stated reason for the rank.
Three consequences, and the third is the durable one:
- Scarce local-counsel hours go to the highest-ranked material rather than to whatever was recognisable to the deal team
- The ranking itself becomes reviewable — local counsel can say “routine here” or, far more valuably, “this one you ranked low is a blocker”
- Those corrections, captured as overrides, become the jurisdiction's calibration for the next deal
That reframes local counsel from a reviewer of documents into a calibrator of the baseline — a better use of the scarcest resource on the transaction, and the only mechanism by which the wrong-baseline failure gets fixed at all.
The verification problem, multiplied
Grounded commercial legal AI has been measured hallucinating between 17% and 33% of the time in an adjacent task, with providers' hallucination-free claims judged overstated.[1] Verification is permanent — and on multilingual deals it is harder, because the person who can verify a Japanese finding is not the person running the screen.
This raises the value of original-language quotes considerably. A finding that carries the source sentence can be verified by whoever reads that language, asynchronously, in seconds. A finding that links to a document in a language your reviewer does not read cannot be verified by them at all.
Bottom line
Stop leading with translation quality. Ask instead how jurisdiction is set, whether severity varies by it, what happens to concepts that do not map, and whether coverage is reported per language.
Then use the ranking to point expensive local counsel at the right material first — and capture their corrections, because the jurisdictional baseline you need cannot be purchased, only accumulated.
Sources
- Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., & Ho, D. E. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. arXiv:2405.20362; Journal of Empirical Legal Studies (2025). Measures English-language legal research — see the scoping note above. arxiv.org/abs/2405.20362
We publish no per-language accuracy figures because we are aware of no independent measurement and will not invent one. Nothing here is legal advice — cross-border materiality is a question for counsel qualified in the relevant jurisdiction. See our methodology.