Anyone quoting you a percentage here is guessing, and we are not going to add to the pile. There is no credible published benchmark for time saved on M&A document review — no peer-reviewed study, no independent measurement, and no vendor in this category publishes so much as its prices, let alone its throughput data. What can be reasoned about honestly is the structure of the saving: where the hours actually go, and which of them a screening stage can remove.
Why the published numbers are worthless
You will see 40%, 60%, 80% claimed. Check the provenance of any of them and you find one of three things: a vendor case study with no methodology, a survey commissioned by a vendor, or a figure that has been repeated so often its origin has dissolved.
The tell is that these numbers are never accompanied by the two facts that would make them meaningful — what the baseline process was, and what counted as “review.” A team that previously read 100% of a data room and now reads 30% of it has not sped up review by 70%. It has changed what it reviews. Those are different claims with different risk profiles, and the percentage conceals which one happened.
Where the hours actually go
The instinct is that review time is reading time. It is not, and this is why tools that only make reading faster disappoint.
| Activity | What it is | Can screening compress it? |
|---|---|---|
| Triage | Deciding what to read, and in what order | Yes — substantially. This is the whole opportunity |
| Reading for extraction | Locating provisions and terms in a document | Yes, partially — if output is quote-backed |
| Reading for judgment | Deciding whether a provision matters for this deal | No. And it should not |
| Cross-referencing | Reading two documents together to see a conflict | Partially — depends on cross-document capability |
| Writing up | Turning findings into committee-ready material | Partially |
| Re-reading | Going back because the first pass missed something | Yes — this is pure waste and often invisible |
Two rows matter more than the rest.
Triage is the largest compressible cost, and it is almost never measured because it does not look like work. It is the senior reviewer opening documents to find out whether they are worth opening. In a large data room this is a substantial fraction of senior time, spent at senior rates, producing nothing but an ordering.
Judgment is not compressible and should not be. Any saving that comes from compressing judgment is not a saving, it is risk transfer to a system that cannot carry it.
A worked example — change every input
Assume a data room of 1,200 documents and a reviewer who works through 12 documents an hour at a level of attention you would defend.
- Full manual coverage: 1,200 ÷ 12 = 100 reviewer-hours.
- With a screening stage: the set is scored and ordered by risk. The team reads the top 25% in full (300 documents) and samples 10% of the remainder (90 documents). That is 390 documents ÷ 12 = 32.5 hours.
- Difference: 67.5 hours on the same document set.
Now the part that matters more than the number. Nothing was read faster. The per-document reading rate is identical in both cases. The entire saving comes from 810 documents never being queued for a senior reviewer at all — because a screening pass established they ranked low against every defined risk category, and that ranking is on the record.
This is why the honest claim is narrow and the dishonest one is broad. A vendor claiming “67% faster review” from this example would be describing something that did not happen. What happened is that the allocation of attention changed.
The comparison nobody runs, and should
The interesting baseline is not “100 hours versus 32.5 hours.” Almost no team reads 100% of a large data room. The real baseline is:
| What gets read | How it was chosen | What the record shows | |
|---|---|---|---|
| Without screening | ~390 documents | Folder order, upload sequence, whoever flagged something | Nothing about the 810 unread |
| With screening | ~390 documents | Rank against defined risk categories | Every one of the 1,200 scored, with source text |
Same hours. Same volume read. Completely different risk position — and the difference is not speed at all. It is that in the first row, the 810 unread documents were selected by arrival order, which correlates with nothing; and if something material was in them, there is no record that anyone considered the question.
In the second, the 810 were examined, scored against every category, and ranked below a threshold the firm set in advance. If a miss surfaces post-close, that is a documented risk-appetite decision rather than a gap nobody was managing. Those two positions are worth very different amounts in an indemnity dispute.
Where the cash saving actually sits
Three places, in descending order of size — and the biggest is the one that never appears in a time-saved calculation.
1. Seniority mix. Triage performed by a screening pass is triage not performed by the most expensive person on the deal. The hour count may fall modestly; the blended rate of the remaining hours falls more, because the hours removed are disproportionately senior-partner-scanning-to-decide-what-matters.
2. Deals abandoned earlier. The cheapest diligence is the diligence you stop. Surfacing a blocking-category issue in the first days rather than the fourth week saves the entire remainder of the workstream, plus adviser fees, plus the opportunity cost of the team. This is almost certainly the largest financial effect available and it is almost never claimed, because it is impossible to attribute cleanly.
3. Re-work avoided. Second passes triggered by “did anyone check the customer contracts for MFN?” are expensive and invisible in every ROI model, because nobody logs them as a distinct activity.
What screening cannot compress
Stated plainly, because a page that only lists benefits is marketing:
- Materiality judgment. Whether a flagged provision kills the deal is human, and a process that lets a score answer it has moved the exposure to the layer least able to carry it.
- Negotiation and structuring. Untouched.
- Anything not in the data room. Oral side agreements, undisclosed arrangements, the thing the seller decided not to upload. No system reads what it is not given.
- Verification of the screening itself. This is a real, recurring cost, not a rounding error. Grounded commercial legal AI has been measured hallucinating between 17% and 33% of the time in an adjacent task[1] — so a verification pass over escalated findings is a permanent line item, and any ROI model that omits it is wrong.
How to measure it on your own deals
If you want a real number rather than a claimed one, instrument three things across your next few transactions:
- Hours by activity, not by deal. Separate triage from reading from judgment from write-up. Without this split you cannot tell which category moved, and the aggregate hides everything interesting.
- Documents examined versus documents read. The gap between these two numbers is your actual coverage position, and most firms have never calculated it.
- Second-pass triggers. Every time someone goes back for something missed on the first pass, log it. This is the cost re-work imposes, and it is invisible otherwise.
Three or four deals of this gives you a defensible internal figure. It will be specific to your firm, your sector and your document mix — which is exactly why the published percentages are meaningless, and why we are not adding one.
Bottom line
The honest answer to “how much time can AI save” is that it depends almost entirely on how much of your current review time is spent deciding what to read. If that is a large share — and in big data rooms it usually is — the compression is real and material. If your team already reads everything at a fixed rate, the saving is much smaller than anyone will tell you.
Either way, the more valuable output is not the hours. It is that every document in the set was examined and ranked against defined categories, so the ones you did not read were a decision rather than an accident.
Sources
- Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., & Ho, D. E. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. arXiv:2405.20362; published in the Journal of Empirical Legal Studies (2025). Measures legal research against case law, not document review against a fixed data room. arxiv.org/abs/2405.20362
The worked example above is an illustration with stated assumptions, not a measurement. We cite only sources we have retrieved and read; where a figure could not be verified, we changed the figure rather than the citation — see our methodology.