Bias in document review does not usually look like discrimination. It looks like a system that is quietly better at some kinds of documents than others — and reports the same confident output for both. The danger is not that it is wrong. It is that the output gives you no way to tell which documents it was wrong about.
Five biases that actually occur here
| Bias | How it shows up | Who it disadvantages |
|---|---|---|
| Drafting-style | Better on documents resembling the training corpus | Smaller counterparties, unusual drafters, older paper |
| Jurisdictional | Severity calibrated to one legal tradition | Non-US/UK targets, civil-law documents |
| Sector | “Market standard” means standard in the dominant sector | Specialist industries with different norms |
| Format | Clean digital text handled well, scans poorly | Older agreements — often the most interesting ones |
| Language | Strongest in English, degrading elsewhere | Cross-border deals, non-English source documents |
Notice these are not fairness problems in the usual sense. They are uneven competence presented as even confidence, and that is what makes them dangerous in a diligence context: a document the system handled poorly produces output indistinguishable from one it handled well.
The one that costs money: severity calibrated to the wrong baseline
Severity is a judgment about what counts as serious, and that judgment came from somewhere — a corpus, a rubric, a set of assumptions about typical deals. When your deal does not match, the system reads the clause correctly and ranks it wrongly.
A clause that is routine in one market may be a genuine transaction blocker in another. An employment provision that is unremarkable in one jurisdiction may create substantial transfer liability elsewhere. The model is not confused; the rubric is aimed at a different world.
Where fairness bias does genuinely apply
Two areas in diligence touch protected characteristics directly, and they deserve separate treatment:
- Employment document review at volume — analysis touching compensation, classification, or termination patterns can surface or obscure disparities depending on how the categories are framed.
- Any output feeding a decision about individuals — retention, key-person assessment, integration planning.
These sit in a different regulatory register. The EU AI Act, applying from 2 August 2025, carries administrative fines under Article 99 reaching €35,000,000 or 7% of total worldwide annual turnover for prohibited practices, with employment-related uses among the areas drawing heightened scrutiny.[1] Whether a specific deployment falls in a high-risk category is a classification question for counsel, but the direction of travel is clear enough that employment-adjacent analysis warrants its own review rather than riding along with contract screening.
Why bias is worse when output is narrative
Uneven performance is manageable if you can see it. Prose output makes it invisible; structured output makes it measurable.
With scored output — a value per defined category, carrying the source sentence, measured against a threshold — you can slice performance by document characteristic:
- Override rate on scanned documents versus digital-native ones
- Override rate by governing law
- Override rate by counterparty size or document age
- Findings per document by language
Each of those is a bias test, and each is a simple aggregation over data you already hold. If overrides cluster on civil-law documents, you have found jurisdictional bias with evidence. With narrative summaries there is no category to slice by and no score to compare — the bias is still there, and there is no instrument that would reveal it.
Detecting it before it costs you
Stratify your evaluation set deliberately
Most pilots use a convenience sample, which usually means recent, clean, English, from one sector. Deliberately include older scanned agreements, documents under a second governing law, a specialist-sector contract, and paper drafted by a small counterparty. Compare recall across those strata rather than in aggregate — the aggregate is what hides the problem.
Track override rate by document characteristic
The single most practical ongoing control, and free once you capture overrides. A category or document type routinely corrected in one direction is a calibration defect with a name.
Test the same clause in two framings
Take a provision, present it in plain modern drafting and in dense traditional drafting, and check whether severity moves. If it does, the system is responding to style rather than substance — which is drafting-style bias, demonstrated in ten minutes.
Check coverage by stratum
Format bias frequently shows up as processing failure rather than scoring error. Reconcile processed counts separately for scans, for long documents, and for non-English files. A system that silently fails on 30% of scanned documents is not scoring them badly — it is not scoring them at all.
What to do about it
- Per-document jurisdiction, not per-deal. A target's contracts may sit under four governing laws; one deal-level setting applies the wrong baseline to most of the set.
- Configurable severity by sector. If severity is a fixed global property of a clause type, the tool cannot represent your world and calibration cannot fix it.
- Original-language source text in the output, so a reviewer qualified in that jurisdiction reads what the document says rather than a round-trip through English.
- Explicit uncertainty on concepts that do not map. The correct behaviour for an untranslatable legal concept is to flag it for local review, not to map it to the nearest English term and proceed.
- Feed overrides back. Local and specialist reviewers correcting severity is the debiasing mechanism. A system that improves through your use is worth substantially more over three years than one that improves only when the vendor ships.
The honest limit
You cannot eliminate this. Every system encodes a view of what is normal, and yours will never exactly match. Grounded commercial legal AI has been measured hallucinating between 17% and 33% of the time in an adjacent task, with providers' hallucination-free claims judged overstated[2] — uneven performance is a permanent feature to be measured, not a defect to be fixed.
So the goal is not an unbiased system. It is a system whose bias you can measure, because a known, quantified skew is manageable — you sample more heavily where the system is weak — and an unknown one is not.
Bottom line
The bias that matters here is uneven competence reported with even confidence: strongest on clean, recent, English, dominant-sector documents, weakest on exactly the older and unusual paper that tends to be interesting.
Stratify your evaluation set, track override rates by document characteristic, check coverage separately for scans and non-English files, and require per-document jurisdiction with configurable severity. Measurable bias is manageable. Invisible bias is what reaches closing.
Sources
- Regulation (EU) 2024/1689 (EU AI Act), Article 99; applies from 2 August 2025. artificialintelligenceact.eu/article/99
- Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., & Ho, D. E. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. arXiv:2405.20362; Journal of Empirical Legal Studies (2025). arxiv.org/abs/2405.20362
Nothing here is legal advice; AI Act classification requires your own counsel. We cite only sources we have retrieved and read — see our methodology.