The validation step that matters is not the one in your process document. It is the one your team still performs in week three of a live deal, at 10pm, with the investment committee meeting on Thursday. Any control that costs meaningful effort per finding will be abandoned exactly when volume is highest — which is exactly when it was load-bearing.
Design for the worst night, not the first week
Most validation protocols are written in calm conditions by someone thinking clearly, and they assume a reviewer will trace each finding to source, assess materiality, and record a decision. Under deadline, with 400 findings, that protocol becomes: skim the red ones, trust the rest.
So the design constraint is not thoroughness. It is cost per finding. A control that takes fifteen seconds survives; a control that takes four minutes does not, no matter how well-reasoned. Everything below is chosen for that property.
The four failure modes you are validating against
Validation is usually framed as “check whether the AI is right.” That is too coarse to design around. There are four distinct ways AI output misleads a committee, and they need different controls.
| Failure | What leadership sees | Control that catches it |
|---|---|---|
| Fabricated support | A cited clause that does not exist | Source-trace — cheap, catches it instantly |
| Misgrounded support | A real clause cited for something it does not say | Read the quote against the claim, not just the link |
| Confident omission | A clean-looking report on an unexamined area | Coverage reporting — cannot be caught by reviewing output |
| Severity inflation | Twelve “critical” findings, three of which are market-standard | Calibration against your own precedent set |
The third row is the one that ruins committee meetings, and note what it says in the right-hand column: it cannot be caught by reviewing the output. Reviewing output validates what the system said. Omission is what it did not say. If your entire validation process operates on the findings list, you have no control at all against the failure mode most likely to embarrass you.
A five-stage protocol that holds under pressure
Stage 1 — Verify coverage before verifying findings
Before anyone reads a single finding, confirm what was actually examined: how many documents were processed, how many failed extraction, which categories returned results, and whether anything was truncated. A document that silently failed to parse produces no findings, and no findings looks identical to a clean document in every report format.
This takes two minutes for an entire data room and it is the highest-value stage in the protocol. Do it first, because if coverage is broken, validating findings is wasted effort.
Stage 2 — Source-trace every finding above your escalation threshold
Not every finding. The ones that will appear in the committee pack. Open the cited document, read the cited sentence, and confirm it supports the assertion — not that it exists, that it supports the claim. Misgrounded citations pass an existence check and fail a reading check, and they are the most common error to survive casual review.
Cost: seconds per finding, if the tool quotes source text. Prohibitive, if it only links to a document. This is the difference that determines whether the protocol survives.
Stage 3 — Calibrate severity against your own precedent
Vendor-default severity does not know your risk appetite, your sector, or what your committee has waved through before. Maintain a short internal reference — this clause type, in this sector, at this deal size, is routine — and re-rank against it. Without this, the first AI-assisted deal you take to committee will over-flag, and the committee will discount every subsequent one.
Stage 4 — Sample the low-severity tail
Pull a fixed random sample of findings nobody would otherwise read — say twenty — and check them properly. This is the only mechanism that measures your real error rate rather than assuming it, because everything else in the protocol is concentrated on the findings you already suspect matter. Errors live where attention does not.
Stage 5 — Record the disagreements
When a reviewer overrides a finding, capture it: what was flagged, what the reviewer concluded, why. Two payoffs. It becomes your calibration set for stage 3, and it is the evidence that human judgment was actually applied — which is what makes the process defensible if the deal goes wrong later.
Why output shape decides whether any of this is affordable
Read the protocol again with an eye on cost. Stages 1, 2 and 4 are cheap or expensive depending entirely on one thing: whether the system produces structured, scored output with source text attached, or narrative prose.
With scored output — a value per defined category, each carrying the quote behind it, measured against a stated threshold:
- Stage 1 is nearly free. Every category reports a value, so an unexamined area shows up as a gap rather than as silence. Coverage becomes readable at a glance instead of being an audit exercise.
- Stage 2 is seconds per finding. The claim and its supporting sentence sit together. Verification is reading two things next to each other.
- Stage 3 becomes mechanical. You are adjusting a number against a threshold, which is a recordable act, rather than arguing with a paragraph.
- Stage 5 writes itself. An override is a changed value with a reason attached — a structured record, not a memo someone has to compose.
With narrative output, every one of those stages requires re-reading source documents, which is the work the tool was bought to reduce. This is the practical reason a scoring stage between intake and review earns its place: it is not that scores are inherently wiser than prose, it is that scores are the format in which verification is cheap enough to actually happen.
What to put in front of the committee
Three things, in this order:
- What was examined — document count, categories covered, anything that failed or was truncated. Lead with the boundary of the work, not the findings. A committee that learns about a coverage gap after making a decision will not trust the next pack.
- What was found, ranked, with source — material issues first, each traceable to the sentence it came from.
- What was overridden and why — the findings your team downgraded, and the reasoning. This is the evidence of judgment, and it is what distinguishes a diligence process from a tool output.
Never present AI output as a finished risk assessment. Present it as a screened and verified exception list, with its coverage stated. The first framing collapses the moment one finding turns out to be wrong. The second survives it, because it never claimed more than it could support.
The duty this sits under
ABA Formal Opinion 512 (29 July 2024) holds that lawyers using generative AI must “fully consider their applicable ethical obligations” — competence, confidentiality, client communication, supervisory responsibility, and fees consistent with time actually spent.[2]
The supervisory duty is the relevant one here. Validation is not a quality-control nicety layered on top of a tool; for regulated professionals it is the mechanism by which reliance on the tool becomes defensible in the first place.
Bottom line
Check coverage before findings. Trace the escalated ones to actual source text. Calibrate severity against your own history rather than the vendor's defaults. Sample the tail nobody reads. Record every override.
And choose tools whose output makes those five things cheap — because a validation protocol that is expensive per finding is not a protocol, it is a document nobody will follow on the night it matters.
Sources
- Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., & Ho, D. E. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. arXiv:2405.20362; published in the Journal of Empirical Legal Studies (2025). Measures legal research, not document review — see the scoping note above. arxiv.org/abs/2405.20362
- American Bar Association Standing Committee on Ethics and Professional Responsibility, Formal Opinion 512: Generative Artificial Intelligence Tools, 29 July 2024. americanbar.org
We cite only sources we have retrieved and read. Where a figure could not be verified, we changed the figure rather than the citation — see our methodology.