What Should Our Team Know About AI Bias in Document Review?

Learning about fairness, consistency, and practical bias management in diligence AI

Updated August 2026 · 6 min read · Deal Room Intelligence Series

Bias in document review does not usually look like discrimination. It looks like a system that is quietly better at some kinds of documents than others — and reports the same confident output for both. The danger is not that it is wrong. It is that the output gives you no way to tell which documents it was wrong about.

Five biases that actually occur here

BiasHow it shows upWho it disadvantages
Drafting-styleBetter on documents resembling the training corpusSmaller counterparties, unusual drafters, older paper
JurisdictionalSeverity calibrated to one legal traditionNon-US/UK targets, civil-law documents
Sector“Market standard” means standard in the dominant sectorSpecialist industries with different norms
FormatClean digital text handled well, scans poorlyOlder agreements — often the most interesting ones
LanguageStrongest in English, degrading elsewhereCross-border deals, non-English source documents

Notice these are not fairness problems in the usual sense. They are uneven competence presented as even confidence, and that is what makes them dangerous in a diligence context: a document the system handled poorly produces output indistinguishable from one it handled well.

The one that costs money: severity calibrated to the wrong baseline

Severity is a judgment about what counts as serious, and that judgment came from somewhere — a corpus, a rubric, a set of assumptions about typical deals. When your deal does not match, the system reads the clause correctly and ranks it wrongly.

This is the failure mode to design against: not a misread provision, but a correctly-read provision scored against the wrong sector's or jurisdiction's norms — producing a low-severity finding that looks exactly like a correct one, with nothing on the face of the report to indicate otherwise.

A clause that is routine in one market may be a genuine transaction blocker in another. An employment provision that is unremarkable in one jurisdiction may create substantial transfer liability elsewhere. The model is not confused; the rubric is aimed at a different world.

Where fairness bias does genuinely apply

Two areas in diligence touch protected characteristics directly, and they deserve separate treatment:

These sit in a different regulatory register. The EU AI Act, applying from 2 August 2025, carries administrative fines under Article 99 reaching €35,000,000 or 7% of total worldwide annual turnover for prohibited practices, with employment-related uses among the areas drawing heightened scrutiny.[1] Whether a specific deployment falls in a high-risk category is a classification question for counsel, but the direction of travel is clear enough that employment-adjacent analysis warrants its own review rather than riding along with contract screening.

Why bias is worse when output is narrative

Uneven performance is manageable if you can see it. Prose output makes it invisible; structured output makes it measurable.

With scored output — a value per defined category, carrying the source sentence, measured against a threshold — you can slice performance by document characteristic:

Each of those is a bias test, and each is a simple aggregation over data you already hold. If overrides cluster on civil-law documents, you have found jurisdictional bias with evidence. With narrative summaries there is no category to slice by and no score to compare — the bias is still there, and there is no instrument that would reveal it.

Detecting it before it costs you

Stratify your evaluation set deliberately

Most pilots use a convenience sample, which usually means recent, clean, English, from one sector. Deliberately include older scanned agreements, documents under a second governing law, a specialist-sector contract, and paper drafted by a small counterparty. Compare recall across those strata rather than in aggregate — the aggregate is what hides the problem.

Track override rate by document characteristic

The single most practical ongoing control, and free once you capture overrides. A category or document type routinely corrected in one direction is a calibration defect with a name.

Test the same clause in two framings

Take a provision, present it in plain modern drafting and in dense traditional drafting, and check whether severity moves. If it does, the system is responding to style rather than substance — which is drafting-style bias, demonstrated in ten minutes.

Check coverage by stratum

Format bias frequently shows up as processing failure rather than scoring error. Reconcile processed counts separately for scans, for long documents, and for non-English files. A system that silently fails on 30% of scanned documents is not scoring them badly — it is not scoring them at all.

What to do about it

  1. Per-document jurisdiction, not per-deal. A target's contracts may sit under four governing laws; one deal-level setting applies the wrong baseline to most of the set.
  2. Configurable severity by sector. If severity is a fixed global property of a clause type, the tool cannot represent your world and calibration cannot fix it.
  3. Original-language source text in the output, so a reviewer qualified in that jurisdiction reads what the document says rather than a round-trip through English.
  4. Explicit uncertainty on concepts that do not map. The correct behaviour for an untranslatable legal concept is to flag it for local review, not to map it to the nearest English term and proceed.
  5. Feed overrides back. Local and specialist reviewers correcting severity is the debiasing mechanism. A system that improves through your use is worth substantially more over three years than one that improves only when the vendor ships.

The honest limit

You cannot eliminate this. Every system encodes a view of what is normal, and yours will never exactly match. Grounded commercial legal AI has been measured hallucinating between 17% and 33% of the time in an adjacent task, with providers' hallucination-free claims judged overstated[2] — uneven performance is a permanent feature to be measured, not a defect to be fixed.

Scoping that figure. It measures open-ended legal research against case law, not review of a document you supplied — a harder retrieval problem, and it is not a bias measurement. It supports only the general point that error rates are real and unsolved.

So the goal is not an unbiased system. It is a system whose bias you can measure, because a known, quantified skew is manageable — you sample more heavily where the system is weak — and an unknown one is not.

Bottom line

The bias that matters here is uneven competence reported with even confidence: strongest on clean, recent, English, dominant-sector documents, weakest on exactly the older and unusual paper that tends to be interesting.

Stratify your evaluation set, track override rates by document characteristic, check coverage separately for scans and non-English files, and require per-document jurisdiction with configurable severity. Measurable bias is manageable. Invisible bias is what reaches closing.

Sources

  1. Regulation (EU) 2024/1689 (EU AI Act), Article 99; applies from 2 August 2025. artificialintelligenceact.eu/article/99
  2. Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., & Ho, D. E. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. arXiv:2405.20362; Journal of Empirical Legal Studies (2025). arxiv.org/abs/2405.20362

Nothing here is legal advice; AI Act classification requires your own counsel. We cite only sources we have retrieved and read — see our methodology.

Test it on your least typical document →

Anweshna Portal
Anweshna Demo