Building Internal Frameworks for AI Quality Assurance

Designing durable QA processes that keep AI outputs trustworthy

Updated August 2026 · 6 min read · Deal Room Intelligence Series

The QA framework that fails is the one written in calm conditions and abandoned in week three of a live deal. Every control in it has to survive 10pm on a Tuesday with a committee meeting on Thursday — which means the binding design constraint is not thoroughness. It is cost per finding. Anything expensive gets skipped exactly when volume is highest, which is exactly when it was load-bearing.

Build on a structure that already exists

You do not need to invent a taxonomy. The NIST AI Risk Management Framework 1.0 (26 January 2023) is voluntary, public, and organised around four functions — Govern, Map, Measure, Manage — which map cleanly onto what a diligence QA programme actually needs.[1]

FunctionIn practice, for diligenceTypical maturity
GovernWho owns AI risk, who approves tools, escalation pathUsually present, often nominal
MapWhere AI is actually used, on what data, in which decisionsFrequently incomplete — shadow usage is common
MeasureError rates, override rates, coverage gaps, grounding checksAlmost always absent
ManageThresholds, verification protocol, retention, response to findingsPartial

Measure is where nearly every firm is weakest, and it is the function that makes the other three credible. Without measurement you have a policy document, not a quality programme — and a policy that has never been tested against data is indistinguishable from one that does not work.

The five controls that survive a live deal

Chosen for cost, not for completeness. Each takes seconds to minutes.

1. Coverage reconciliation — before anyone reads a finding

Documents submitted, processed in full, partially processed, failed, and why. Two minutes for an entire data room, and the highest-value control available.

The reason it comes first: a document that failed extraction produces zero findings, and zero findings is indistinguishable from a clean document in almost every report format. If coverage is broken, everything downstream is auditing a sample you did not choose.

2. Grounding trace on escalations

For every finding heading to a committee, open the cited document and read the cited sentence against the claim. Not that the reference exists — that it supports the assertion. Misgrounded citations pass an existence check and fail only a reading check.

Cost: seconds, if the tool quotes source text. Prohibitive, if it only links to a document. This single property decides whether your QA framework is affordable.

3. Fixed random sample of the low-severity tail

Twenty findings nobody would otherwise read, every deal, checked properly. This is the only control that measures your error rate rather than assuming it, because every other control concentrates on findings you already suspect matter. Errors live where attention does not.

4. Override capture

What was flagged, what the reviewer concluded, why. Three payoffs from one habit: evidence that judgment was applied, calibration data for the rubric, and an answer to “why was this not escalated?” months later.

5. Reproducibility spot-check

Periodically re-run a previously processed document and compare. Divergence means either an undisclosed model change or non-determinism — both of which break your ability to reconstruct any earlier decision.

What to actually measure

Four numbers, tracked over time. Absolute values matter less than movement.

Scoping that hallucination figure. It measures open-ended legal research against case law, not review of a document you supplied — a harder retrieval problem. Use it to set the expectation that errors exist and must be measured, not as a target rate for your own sampling.

Decide the response before you run the measurement

An audit with no pre-committed response is a document-generating exercise. Define in advance:

Deciding afterwards means deciding while looking at a number you do not like, with a live transaction depending on the answer. That is not a decision, it is a negotiation with yourself.

The framework has to shape procurement

Most QA programmes are written after the tool is bought, which guarantees some controls are unimplementable. Three properties are not add-ons — if the tool lacks them, the corresponding control cannot exist at any price:

  1. Source quotes on findings. Without this, grounding traces cost a re-read and will be skipped.
  2. Coverage reporting as a first-class output. Without this, reconciliation is manual and will be skipped.
  3. Model versioning with retained run records. Without this, reproducibility checks are impossible and no earlier decision can be reconstructed.

Put these in the evaluation criteria rather than the post-purchase policy.

Why the structure of the output decides the cost of QA

Read the five controls again and notice they all depend on one property: output that is structured and grounded — a value per defined category, carrying the sentence it rests on, measured against a stated threshold.

Narrative output cannot support any of them. There is no category to sample by, no threshold to evidence, no score to track overrides against, no quote to trace. Verification means re-reading the document, and a QA framework whose every control costs a re-read is a framework that will not be followed.

This is the practical reason a scoring stage is worth more than a summarisation stage — not that numbers are wiser than prose, but that scores are the only output format in which quality assurance is cheap enough to actually happen.

The obligation underneath

ABA Formal Opinion 512 (29 July 2024) holds that lawyers using generative AI must “fully consider their applicable ethical obligations,” including competence and supervisory responsibility.[3] For regulated professionals, a QA framework is not an internal nicety layered on top of a tool — it is the mechanism by which reliance on the tool becomes defensible.

Bottom line

Design for the worst night, not the first week. Five controls, each cheap: reconcile coverage before reading findings, trace grounding on escalations, sample the tail nobody reads, capture every override, and spot-check reproducibility.

Measure four numbers and watch them move. Decide your response thresholds before you look at the data. And buy tools that make the controls affordable — because a QA framework that costs a re-read per finding is a document, not a programme.

Sources

  1. NIST AI Risk Management Framework (AI RMF 1.0), published 26 January 2023. nist.gov/itl/ai-risk-management-framework
  2. Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., & Ho, D. E. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. arXiv:2405.20362; Journal of Empirical Legal Studies (2025). Measures legal research, not document review — see the scoping note above. arxiv.org/abs/2405.20362
  3. ABA Standing Committee on Ethics and Professional Responsibility, Formal Opinion 512: Generative Artificial Intelligence Tools, 29 July 2024. americanbar.org

We cite only sources we have retrieved and read — see our methodology.

See output built for cheap verification →

Anweshna Portal
Anweshna Demo