What AI Tools Actually Work for M&A Due Diligence?

Finding proven solutions that deliver real results — not just demos

Updated August 2026 · 5 min read · Deal Room Intelligence Series

You came here for a list of vendors. We are not going to publish one, and the reason is the most useful thing on this page: there is no independent benchmark for this category, and essentially no vendor publishes prices. Any ranked list — ours or anyone's — would be assembled from marketing material and sales calls, then presented as research. What can be established is which approaches work, and how to find out which product does.

What “works” has to mean here

The word does most of the damage. A tool that produces impressive output on a document you chose, in a category it was tuned for, with a presenter who knows the answer, “works” in the only sense a demo can demonstrate — and that sense predicts nothing.

A workable definition, and one you can test: a tool works if, on your documents, it

  1. tells you what it did not process,
  2. surfaces the issues you already know are there,
  3. lets you confirm a finding in seconds rather than minutes,
  4. produces the same answer twice, and
  5. can still be reconstructed in eighteen months.

Notice none of these are about intelligence. They are about behaviour under error — which is the axis products actually differ on, now that the underlying models have largely converged.

Approaches that work, by job

JobApproach that worksApproach that does not
Deciding what to read in a large roomRisk screening — every document scored, ranked, quotedQ&A — only answers what you thought to ask
Building a structured recordClause extraction into filterable fieldsNarrative summaries
Negotiating your own paperPlaybook comparisonScreening — everything deviates, so everything flags
Regulatory thresholds and filingsDeterministic rules fed by extractionAsking a model whether a filing is required
Exploring one known documentQ&A

Most disappointment in this category is a job-to-approach mismatch rather than a quality problem. Buying a playbook tool for buy-side diligence, or treating Q&A as the coverage layer, produces a poor outcome from a perfectly competent product.

The failure that makes a “working” tool dangerous

A tool can score well on every visible measure and still be unsafe, if it cannot tell you what it missed.

Documents fail silently — scans without a text layer, files past a length ceiling, unsupported formats. Each produces zero findings, and zero findings is indistinguishable from a clean document in almost every report format. A tool that processes 92% of a room and reports 100% is not 92% as good; it is producing false confidence, which is worse than a visible failure.

So the first question is not “how accurate is it?” but “what does it tell me about what it could not read?” — and the first thing to do with any output is reconcile the processed count against what you submitted.

Why we cannot tell you the accuracy of anything

There is no independent published benchmark for M&A document review. The nearest rigorous work is in an adjacent task: Stanford researchers evaluated the retrieval-grounded legal AI research tools sold by LexisNexis and Thomson Reuters and found they “each hallucinate between 17% and 33% of the time,” concluding that providers' hallucination-free claims are overstated.[1]

Scoping that properly. It measures open-ended legal research against case law — a harder retrieval problem than locating a clause in a document you supplied — so it does not give you a review-accuracy figure. What it establishes is that an independent evaluation found materially worse performance than the vendors' own descriptions. Expect that gap here too, and treat any self-reported accuracy number accordingly.

The five-step test that answers the question properly

Take 40–60 documents from a closed deal where your team knows every material issue. Include, deliberately: a scan with no text layer, an over-length credit agreement, a spreadsheet, and a base agreement whose amendment you have withheld.

  1. Reconcile counts. Submitted, processed, failed. If failures are silent, stop — the rest of the test is measuring a sample you did not choose.
  2. Score recall against your known issues, not against total findings returned.
  3. Time ten verifications with a stopwatch. This number compounds across every deal you will ever run and appears in no pitch.
  4. Read quotes against claims, hunting misgrounded citations — a real clause cited for something it does not say. These pass an existence check and fail a reading check.
  5. Run one document twice and compare.

Then ask the question that outranks the test: “Tell me about a finding your system got wrong in a live deal, and how the customer found out.” Every real product has this story. A vendor without one has no live deployments, or no feedback path by which errors return to them — meaning their error rate is unmeasured rather than low.

What we do publish

For the avoidance of doubt about our own position: our prices are public at /pricing.html, which in this category is unusual and deliberate. We publish no accuracy percentage, no throughput benchmark, and no comparison against named competitors — because a figure measured on our own chosen documents would not predict yours, and the rest is not knowable.

Run the five-step test on our output alongside anyone else's. That is a better basis for a decision than any list we could write.

Bottom line

No credible vendor ranking exists for this category, and anyone publishing one is showing you marketing with a table around it. What works is determined by matching approach to job — screening for coverage, extraction for structure, playbooks for your own paper, deterministic rules for thresholds.

Then judge the specific product on behaviour under error: does it report what it could not read, can you verify a finding in seconds, does it answer the same way twice. Those are testable in an afternoon, and they are what you will actually live with.

Sources

  1. Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., & Ho, D. E. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. arXiv:2405.20362; Journal of Empirical Legal Studies (2025). Measures legal research, not document review — see the scoping note above. arxiv.org/abs/2405.20362

We name no competitors and publish no vendor ranking, because no independent benchmark exists and no vendor in this category publishes prices. We cite only sources we have retrieved and read — see our methodology.

Run the five-step test on our output →

Anweshna Portal
Anweshna Demo