The percentage answers you will see — 40%, 70%, 80% — are answering a question nobody asked, because “due diligence” is not one activity with a fraction that can be mechanised. Split it by what kind of thing is being decided and the boundary becomes sharp, non-negotiable, and far more useful than any number.
The line that actually divides the work
Every diligence task falls into one of three classes. The class determines automatability completely; the technology barely matters.
| Class | The question being answered | Automatable? |
|---|---|---|
| Retrieval | What does this document say? | Largely yes |
| Allocation | Which documents deserve expert attention? | Yes — the biggest opportunity |
| Judgment | Does this matter, and what do we do? | No |
The interesting row is the middle one, and it is the one left out of almost every discussion of this topic. Allocation is not analysis and it is not reading — it is deciding what to read. It is currently performed by the most expensive people on the deal, it produces no work product beyond an ordering, and it is almost entirely mechanisable.
Most estimates of “how much can be automated” are really estimates of retrieval, which understates the opportunity, because retrieval is the smaller prize.
What genuinely automates
- Locating provisions across a document set. Change-of-control, assignment, exclusivity, termination, indemnity caps — findable at volume with source text attached.
- Consistency checking at scale. The same clause type across 900 documents, assessed identically on the last as on the first. Humans degrade here through fatigue; systems do not.
- Completeness checks against a list. Which permits exist, which consents are missing, which base agreements reference amendments that are absent from the room.
- Ranking and triage. Scoring every document against defined categories so expert hours go to the highest-risk material rather than to whatever sat in the first folder.
- Extraction into a structured record. Turning prose obligations into fields you can filter, aggregate and audit.
What does not, and will not
- Materiality. Whether a provision matters depends on the transaction, the counterparty relationship, the client's tolerance, and the negotiation ahead. None of that is in the document.
- Inference across silence. Noticing the indemnity that should be present and is not. Experienced reviewers do this constantly; it requires a model of what ought to exist rather than what does.
- Anything outside the data room. Oral side agreements, undisclosed arrangements, the amendment the seller chose not to upload. No system reads what it is not given, and this is the largest single source of real misses.
- Predicting a regulatory decision. Whether an approval will be granted is a forecast about an institution, not a fact in a filing.
- Accountability. ABA Formal Opinion 512 (29 July 2024) holds that lawyers using generative AI must “fully consider their applicable ethical obligations,” including competence and supervisory responsibility.[1] The duty attaches to the output that leaves your desk. It does not move to a vendor, and no liability cap in this market is sized to accept it.
Why the honest number is small and the honest gain is large
If you insist on a fraction, the defensible answer is that the automatable share of hours is meaningful, while the automatable share of value is small — because the hours that compress are the low-judgment ones, and the value concentrates in the judgment that does not.
That sounds like a limitation. It is actually the argument, and here is why.
A worked illustration, with invented inputs you should replace: a 1,200-document room and a team with capacity for roughly 390 documents. Without screening, those 390 are selected by arrival order and the other 810 leave no trace. With screening, all 1,200 are examined and scored, the same 390 hours-worth are spent on the top-ranked material, and the 810 are on record as scored below a threshold set in advance.
Nothing was automated away. The hours are identical. What changed is that unexamined risk became examined-and-ranked risk — which is a different position entirely when a miss is investigated later.
The permanent cost that caps the ceiling
Any realistic estimate must include a task that automation creates: verification.
Grounded commercial legal AI has been measured hallucinating between 17% and 33% of the time in an adjacent task, with providers' hallucination-free claims found to be overstated.[2] Nobody has eliminated that. So someone must check escalated findings against source text, permanently — and that work did not exist before.
This is why the cost of verifying a single finding is the variable that decides whether automation pays at all. Output that quotes its source sentence is verified in seconds. Output that produces fluent unattributed prose is verified by re-reading the document — at which point you have added a step without removing one, and the automation is negative.
A sensible target
Rather than chasing a percentage, aim at four specific conditions:
- 100% of the room examined against defined categories — not read, examined. This is achievable and it is the whole point.
- Expert hours unchanged but redirected to the top of a risk ranking.
- Verification cheap by construction — every finding carrying the sentence it rests on.
- Absence of a finding visibly different from absence of a search — blocking categories floored rather than reporting a comfortable zero.
Hit those and you have taken essentially all of the available gain. No further percentage is waiting to be captured, because what remains is judgment, and judgment is what the client is paying for.
Bottom line
Retrieval and allocation automate. Judgment does not, and accountability cannot. The honest framing is not that a percentage of diligence disappears — it is that the unexamined portion of the data room disappears, while the expert workload stays roughly where it was and moves to better-chosen material.
Anyone quoting you a large automation percentage is either measuring hours and calling it value, or has not budgeted the verification the automation itself creates.
Sources
- American Bar Association Standing Committee on Ethics and Professional Responsibility, Formal Opinion 512: Generative Artificial Intelligence Tools, 29 July 2024. americanbar.org
- Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., & Ho, D. E. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. arXiv:2405.20362; Journal of Empirical Legal Studies (2025). arxiv.org/abs/2405.20362
The worked illustration uses stated assumptions and is not a measurement. We cite only sources we have retrieved and read — see our methodology.