Multi-jurisdiction deals carry a layer that single-jurisdiction deals do not: a regulatory surface of filings, approvals and thresholds that must be identified correctly or the transaction cannot complete. This layer is rules, not judgment — and handing it to a language model is a category error that a surprising number of deployments make.
Separate the two layers before anything else
| Document layer | Regulatory layer | |
|---|---|---|
| Question | What do these contracts say, and what matters? | Which filings are required, where, by when? |
| Nature | Probabilistic — extraction and ranking | Deterministic — thresholds and tests |
| Right tool | Screening with severity calibrated per jurisdiction | Rules engine against maintained thresholds |
| Failure cost | A missed provision — expensive | A missed filing — can unwind a completed deal |
The bottom row is why the separation matters commercially. A missed contractual provision costs money. A missed merger-control or foreign-investment filing is among the few diligence failures that can undo a transaction after closing, and it is not a place for probabilistic output.
The correct architecture for the regulatory layer
The model's job is extraction with citation, not determination. It should pull the inputs a threshold test needs — turnover by jurisdiction, asset values, shareholding percentages, sector classification, employee counts — and quote where each figure came from.
A deterministic rules layer then applies maintained thresholds to those inputs. Where a figure sits close to a limit, or where an input could not be extracted with confidence, the correct output is escalate, not resolve.
Ask also who maintains the thresholds and how often. Filing thresholds are revised, sometimes annually. A rules layer with stale numbers is worse than none, because it produces a confident clear where there should be a filing.
Why per-deal jurisdiction settings break the document layer
The most common configuration error in cross-border work: setting one jurisdiction for the transaction.
A target's customer contracts may sit under three governing laws, its employment documents under a fourth, and its financing under a fifth. A single deal-level setting applies the wrong materiality baseline to most of the room — and produces low-severity findings that are indistinguishable from correct ones, because a wrongly-calibrated low score looks exactly like a right one.
Jurisdiction must be a per-document property, and severity must be configurable against it. If severity is a fixed global attribute of a clause type, the tool cannot represent the problem and no amount of tuning will fix it.
Complexity that is structural, not linguistic
Four things make multi-jurisdiction deals hard in ways that have nothing to do with translation:
- Concepts that do not map. Security interests that perfect differently, employment protections with no analogue, corporate forms whose English translation is misleadingly familiar. The dangerous behaviour is silent mapping to the nearest equivalent; the correct behaviour is flagging for local review.
- Enforceability versus text. Identical clauses can be enforceable in one jurisdiction and void in another. The document does not say which, and no reading of it will reveal the answer.
- Conditions precedent that interact. Approvals in different jurisdictions on different clocks, some conditional on others. This is a sequencing problem, not a document problem.
- Unwritten local practice. What is customary is frequently undocumented, and no system reads what was never written.
Where screening genuinely earns its place
The binding constraint on multi-jurisdiction diligence is local counsel time — expensive, engaged late, in another time zone, on the critical path. Teams manage it by sending local counsel whatever the deal team could tell was relevant, which across a language and legal-system barrier is a weak filter.
A screening pass that scores every document per jurisdiction changes what reaches them: a ranked set, with the source clause in its original language, and a stated reason for the rank. Three effects:
- Scarce local-counsel hours go to the highest-ranked material rather than to whatever was recognisable to the deal team
- The ranking becomes reviewable — local counsel can say “routine here” or, more valuably, “this one you ranked low is a blocker in this jurisdiction”
- Those corrections, captured as overrides, become that jurisdiction's calibration for the next deal
The third point is the durable one. It is also the only mechanism by which the wrong-baseline problem gets fixed at all, since the required knowledge cannot be purchased — only accumulated.
Coverage, per jurisdiction
Require coverage reporting broken down by jurisdiction and language, and reconcile it before reading findings. A language or format handled poorly produces few findings or none — and zero findings is indistinguishable from a clean document in almost every report format.
On a multi-jurisdiction deal this failure concentrates: it is usually one jurisdiction's document set that gets silently under-processed, which is exactly the pattern that looks like “nothing much of concern in Germany.”
What stays human, permanently
- Local materiality. Whether a provision is serious in a given market is a question for someone qualified there. A score is a prompt for that question, not an answer.
- Whether an approval will be granted. That is a forecast about an institution, not a fact in a filing.
- Sequencing conditions precedent across jurisdictions with interacting clocks.
- Anything not disclosed — and local disclosure norms differ from what a buyer's team expects to be given, which bites harder cross-border than domestically.
The verification constraint
Grounded commercial legal AI has been measured hallucinating between 17% and 33% of the time in an adjacent task, with providers' hallucination-free claims judged overstated.[1] On multi-jurisdiction work verification is harder, because the person who can confirm a finding is not the person who ran the scan and may be eight time zones away.
This makes original-language quoted source text more valuable here than anywhere else. A finding carrying the source sentence can be verified asynchronously in seconds by whoever reads that language. A finding that merely links to a document in a language your reviewer does not read cannot be verified by them at all.
Bottom line
Split the layers. Let a rules engine handle filings and thresholds, with the model extracting and citing inputs rather than deciding — a missed filing is the one diligence failure that can unwind a closed deal.
On the document layer, set jurisdiction per document, configure severity against it, demand original-language quotes and per-jurisdiction coverage reporting, and use the ranking to point expensive local counsel at the right material first. Then capture what they tell you, because that is the only way the baseline ever becomes correct.
Sources
- Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., & Ho, D. E. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. arXiv:2405.20362; Journal of Empirical Legal Studies (2025). See the scoping note above. arxiv.org/abs/2405.20362
We publish no merger-control or foreign-investment filing thresholds on this page. They vary by jurisdiction and are revised periodically; rely only on current primary sources or local counsel. Nothing here is legal advice — see our methodology.