What Documents Should AI Review in M&A Deals?

Identifying scope and highest-value use cases for AI in due diligence

Updated August 2026 · 6 min read · Deal Room Intelligence Series

The instinct is to point screening at the important documents. That is backwards. Your team already reads the important documents carefully — the top customer contracts, the credit agreement, the cap table. Screening earns its place on the material nobody was going to reach, because that is where unexamined risk actually accumulates.

The selection rule

Two questions decide whether a document belongs in a screening pass:

  1. Would a senior reviewer definitely read this in full? If yes, screening adds a consistency check but little coverage value.
  2. Could this document contain an obligation that changes the deal? If yes, and the answer to (1) is no, it is exactly what screening is for.

Everything worth knowing about scoping follows from the intersection: documents that can carry material obligations but are unlikely to be individually read. In most data rooms this is the majority of the room by count and a substantial share of the risk.

Highest value: the long tail of commercial contracts

Customer and supplier agreements ranked roughly eleven through two hundred. Individually unremarkable, collectively material, and reliably under-read because attention goes to the top ten by revenue.

What hides here: change-of-control and assignment provisions, most-favoured-nation pricing, exclusivity, auto-renewal, unusual termination rights, uncapped indemnities. Any one of these in a mid-tier contract can alter deal economics, and none of them announce themselves from a folder listing.

This is the archetypal screening case, and the mechanism is worth naming precisely: the value is not that a system reads these better than a lawyer would. It is that a lawyer was never going to read contract number 143. The comparison is against nothing, not against expert review.

High value: employment and contractor agreements at volume

Individually small, structurally repetitive, and the place where non-competes, IP assignment gaps, change-of-control bonuses, severance triggers and misclassification risk aggregate into real numbers. Consistency checking across hundreds of near-identical documents is precisely where human attention degrades and a system does not — the same clause on document 400 receives what it received on document 4.

The high-value question here is usually an aggregate rather than a clause: how many agreements lack IP assignment? That is a coverage question, and it is unanswerable without processing all of them.

High value: amendments, side letters and appendices

The single most under-read category in most data rooms, and the one most likely to contain the provision that matters. A base agreement gets read; the fourth amendment filed separately eighteen months later does not.

Two things to demand for this class specifically: the system must link amendments to their base agreements, and it must flag a base agreement whose referenced amendments are absent from the room. The second is a disclosure-gap detector and it is worth more than most finding types.

Moderate value: financial statements and schedules

Useful for locating disclosure items, related-party transactions, contingent liabilities and off-balance-sheet arrangements — and for cross-referencing what the contracts say against what the accounts recognise. Not a substitute for quality-of-earnings work, which is analytical rather than extractive.

Moderate value: corporate records and permits

Board minutes, consents, share transfers, licences. Screening is good at completeness checks against an expected list — which permits exist, which expire inside the deal horizon, which consents are missing. This is closer to a checklist function than anomaly detection, and it is undramatic but reliably useful.

Low value, and worth saying so

The scoping failure that causes real misses

Every list like this assumes documents get processed. In practice a meaningful share do not, and the failure is silent.

Each of these produces zero findings, and zero findings is indistinguishable from a clean document in almost every report format. This is why the most valuable output of a screening pass is not the findings list — it is the coverage report: how many documents were submitted, how many processed, how many failed and why, and what was truncated.

If you take one operational point from this page: reconcile the processed count against the room's document count before reading a single finding. A screening pass that silently skipped 8% of the room is worse than no screening pass, because it produces false confidence.

Scope in three passes, not one

Sequencing matters more than the list, because the value of an early finding is much higher than the value of a complete one.

PassWhenScopePurpose
1 — BlockingDay one, on whatever existsEverything availableFind the transaction-ender before the workstream runs
2 — FullAs the room fillsAll obligation-bearing documentsRank the set; direct expert hours
3 — DeltaOn each upload batchNew and amended onlyCatch late disclosure, which is where surprises cluster

The third pass is routinely skipped and routinely regretted. Documents uploaded late in a process are disproportionately the ones a seller was slow to disclose.

Configure categories before you configure scope

A wider document scope with an undefined rubric produces more noise, not more coverage. Define the blocking categories first — the handful of issues that genuinely stop a deal in your sector — and make sure a category returning nothing carries an explicit floor rather than a comfortable zero. Then widen scope.

Scope without a rubric is how teams end up with four hundred undifferentiated flags and conclude the technology does not work.

Bottom line

Point screening at the long tail: mid-tier commercial contracts, employment agreements at volume, and amendments filed separately. Skip the handful of documents your team reads properly anyway. Include everything that can carry an obligation, because the cost of examining a document is small and the cost of an unexamined one is not.

Then check the coverage report before the findings. The document that never got processed is the one that will hurt, and it looks exactly like a clean result.

Sources

This page makes no external statistical claims; its arguments are structural and testable against your own data rooms. Where we cite figures elsewhere on this site, we cite only sources we have retrieved and read — see our methodology. Nothing here is legal advice.

Run a document through and see the coverage output →

Anweshna Portal
Anweshna Demo