The QA framework that fails is the one written in calm conditions and abandoned in week three of a live deal. Every control in it has to survive 10pm on a Tuesday with a committee meeting on Thursday — which means the binding design constraint is not thoroughness. It is cost per finding. Anything expensive gets skipped exactly when volume is highest, which is exactly when it was load-bearing.
Build on a structure that already exists
You do not need to invent a taxonomy. The NIST AI Risk Management Framework 1.0 (26 January 2023) is voluntary, public, and organised around four functions — Govern, Map, Measure, Manage — which map cleanly onto what a diligence QA programme actually needs.[1]
| Function | In practice, for diligence | Typical maturity |
|---|---|---|
| Govern | Who owns AI risk, who approves tools, escalation path | Usually present, often nominal |
| Map | Where AI is actually used, on what data, in which decisions | Frequently incomplete — shadow usage is common |
| Measure | Error rates, override rates, coverage gaps, grounding checks | Almost always absent |
| Manage | Thresholds, verification protocol, retention, response to findings | Partial |
Measure is where nearly every firm is weakest, and it is the function that makes the other three credible. Without measurement you have a policy document, not a quality programme — and a policy that has never been tested against data is indistinguishable from one that does not work.
The five controls that survive a live deal
Chosen for cost, not for completeness. Each takes seconds to minutes.
1. Coverage reconciliation — before anyone reads a finding
Documents submitted, processed in full, partially processed, failed, and why. Two minutes for an entire data room, and the highest-value control available.
The reason it comes first: a document that failed extraction produces zero findings, and zero findings is indistinguishable from a clean document in almost every report format. If coverage is broken, everything downstream is auditing a sample you did not choose.
2. Grounding trace on escalations
For every finding heading to a committee, open the cited document and read the cited sentence against the claim. Not that the reference exists — that it supports the assertion. Misgrounded citations pass an existence check and fail only a reading check.
Cost: seconds, if the tool quotes source text. Prohibitive, if it only links to a document. This single property decides whether your QA framework is affordable.
3. Fixed random sample of the low-severity tail
Twenty findings nobody would otherwise read, every deal, checked properly. This is the only control that measures your error rate rather than assuming it, because every other control concentrates on findings you already suspect matter. Errors live where attention does not.
4. Override capture
What was flagged, what the reviewer concluded, why. Three payoffs from one habit: evidence that judgment was applied, calibration data for the rubric, and an answer to “why was this not escalated?” months later.
5. Reproducibility spot-check
Periodically re-run a previously processed document and compare. Divergence means either an undisclosed model change or non-determinism — both of which break your ability to reconstruct any earlier decision.
What to actually measure
Four numbers, tracked over time. Absolute values matter less than movement.
- Coverage gap — documents in the room versus documents processed. Should trend to zero. If it does not, you have a technical problem masquerading as a quality problem.
- Grounding failure rate — share of sampled findings where the quote does not support the claim. Expect this to be non-zero. Grounded commercial legal AI has been measured hallucinating between 17% and 33% of the time in an adjacent task, with providers' hallucination-free claims judged overstated.[2] A sample that finds nothing is more likely a badly designed sample than a perfect system.
- Override rate by category — a category routinely downgraded is mis-tuned for your sector and its scores are noise. Routinely upgraded means systematic under-detection.
- Escalation-to-materiality rate — of findings escalated, how many turned out to matter? Falling numbers mean the threshold is too loose and your team is being trained to skim.
Decide the response before you run the measurement
An audit with no pre-committed response is a document-generating exercise. Define in advance:
- What grounding failure rate triggers a rubric review
- What triggers a vendor conversation
- What triggers suspending use of a category
- Who is authorised to make that call mid-deal
Deciding afterwards means deciding while looking at a number you do not like, with a live transaction depending on the answer. That is not a decision, it is a negotiation with yourself.
The framework has to shape procurement
Most QA programmes are written after the tool is bought, which guarantees some controls are unimplementable. Three properties are not add-ons — if the tool lacks them, the corresponding control cannot exist at any price:
- Source quotes on findings. Without this, grounding traces cost a re-read and will be skipped.
- Coverage reporting as a first-class output. Without this, reconciliation is manual and will be skipped.
- Model versioning with retained run records. Without this, reproducibility checks are impossible and no earlier decision can be reconstructed.
Put these in the evaluation criteria rather than the post-purchase policy.
Why the structure of the output decides the cost of QA
Read the five controls again and notice they all depend on one property: output that is structured and grounded — a value per defined category, carrying the sentence it rests on, measured against a stated threshold.
Narrative output cannot support any of them. There is no category to sample by, no threshold to evidence, no score to track overrides against, no quote to trace. Verification means re-reading the document, and a QA framework whose every control costs a re-read is a framework that will not be followed.
This is the practical reason a scoring stage is worth more than a summarisation stage — not that numbers are wiser than prose, but that scores are the only output format in which quality assurance is cheap enough to actually happen.
The obligation underneath
ABA Formal Opinion 512 (29 July 2024) holds that lawyers using generative AI must “fully consider their applicable ethical obligations,” including competence and supervisory responsibility.[3] For regulated professionals, a QA framework is not an internal nicety layered on top of a tool — it is the mechanism by which reliance on the tool becomes defensible.
Bottom line
Design for the worst night, not the first week. Five controls, each cheap: reconcile coverage before reading findings, trace grounding on escalations, sample the tail nobody reads, capture every override, and spot-check reproducibility.
Measure four numbers and watch them move. Decide your response thresholds before you look at the data. And buy tools that make the controls affordable — because a QA framework that costs a re-read per finding is a document, not a programme.
Sources
- NIST AI Risk Management Framework (AI RMF 1.0), published 26 January 2023. nist.gov/itl/ai-risk-management-framework
- Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., & Ho, D. E. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. arXiv:2405.20362; Journal of Empirical Legal Studies (2025). Measures legal research, not document review — see the scoping note above. arxiv.org/abs/2405.20362
- ABA Standing Committee on Ethics and Professional Responsibility, Formal Opinion 512: Generative Artificial Intelligence Tools, 29 July 2024. americanbar.org
We cite only sources we have retrieved and read — see our methodology.