There is no independent benchmark for AI due diligence tools. None. No peer-reviewed evaluation, no published methodology, no shared test set — and essentially no vendor publishes prices either. That is not a gap you can research your way around, so the only benchmark that will ever exist for your firm is the one you run yourself. This is how to run it so the result means something.
Why vendor benchmarks tell you nothing
Any accuracy figure a vendor quotes was produced on a test set the vendor designed, scored against a definition of “correct” the vendor chose. Neither is disclosed. That is not necessarily dishonest — there is no shared standard to use instead — but it makes cross-vendor comparison meaningless.
The closest rigorous work in an adjacent domain is Stanford's evaluation of commercial legal AI research tools, which found the products sold by LexisNexis and Thomson Reuters “each hallucinate between 17% and 33% of the time,” and concluded providers' claims were overstated.[1] The instructive part is not the number — it is that an independent evaluation found materially worse performance than the vendors' own descriptions. Expect the same gap here.
Design the test before you see a demo
The order matters. A demo installs the vendor's framing — their categories, their idea of what good output looks like — and a test designed afterwards tends to measure the things that vendor happens to do well.
Build ground truth from a closed deal
Take 40–60 documents from a completed transaction. Have someone who worked it list every issue they consider material, before any tool touches it. That list is your ground truth: imperfect, but calibrated to your firm's actual standard rather than a vendor's.
Seed known negatives
Include documents you are confident are unremarkable. Findings returned on those measure the false-positive rate on exactly the material most likely to be waved through.
Seed structural edge cases deliberately
This is where evaluations usually go soft. Include:
- A scanned document with no text layer
- A document past any plausible length ceiling — the credit agreement
- An amendment whose base agreement is present, and a base agreement whose amendment is absent
- A document in another language, if you do cross-border work
- Two documents that only conflict when read together
These probe the failures that cause real misses, and no vendor demo includes them.
What to measure — and what to ignore
| Metric | How | Weight |
|---|---|---|
| Recall on your ground truth | Of your known material issues, how many surfaced? | Highest |
| Coverage integrity | Were failures and truncations reported, or silent? | Highest |
| Grounding accuracy | Does the quoted text support the claim? | Highest |
| Verification time | Seconds to confirm one finding | High |
| Top-20 precision | Of the 20 highest-ranked, how many were worth reading? | High |
| Reproducibility | Same document twice — same output? | High |
| Overall precision | Across all findings | Low |
| Processing speed | Documents per hour | Ignore |
| Feature count | — | Ignore |
Three of these deserve explanation because they invert conventional evaluation.
Coverage integrity ranks with recall. A tool that silently skips 8% of the room is not 92% as good — it is producing false confidence, which is worse than a visible failure. Check the processed count against what you submitted before scoring anything else.
Top-20 precision beats overall precision. Reviewers under deadline read the top of the list. A tool with mediocre overall precision and excellent top-20 precision is more useful than the reverse, and standard metrics hide this completely.
Verification time is a first-class metric. Multiply seconds-per-finding by your expected escalation volume and it usually dominates every efficiency claim in the pitch. Two tools with identical recall can differ by an order of magnitude here, entirely on whether findings quote their source.
Run it as a blind comparison if you can
Have someone not involved in procurement strip vendor branding from the outputs and hand the reviewer a merged, shuffled list of findings to assess. Preference for the incumbent, or for the vendor with the better sales team, is real and it survives good intentions.
At minimum, have the person scoring grounding accuracy be someone who did not sit through the demos.
The questions that decide it, which no benchmark captures
After the numbers, three qualitative answers carry disproportionate weight:
- “Tell me about a finding your system got wrong in a live deal, and how the customer found out.” Every real product has this story. A vendor without one either has no live deployments or no feedback path — meaning their error rate is unmeasured rather than low.
- “How do I distinguish a category you examined and found clean from one you did not examine?” If the answer is that the report simply omits it, the tool cannot tell you what it missed.
- “When you update the model, what happens to prior results?” “We're always improving” means last quarter's report is not reproducible.
Scoring it honestly
Weight your criteria before you see results, and write the weights down. Otherwise the tool that wins is the one that happened to perform well on whatever you looked at first.
A defensible default for buy-side diligence: recall on ground truth and coverage integrity at the top, grounding accuracy and verification time next, ranking quality and reproducibility after that, and everything else as a tiebreak. If a tool fails coverage integrity, stop — the rest of the score is measuring a sample you did not choose.
What a pilot costs, and why it is cheap
A proper evaluation takes a few days of one person's time plus, ideally, a paid pilot. Vendors resist paid pilots and should be pushed, because a free demo on vendor-selected documents measures nothing.
Set against a multi-year commitment in a market with no published prices and no independent benchmarks, a few days is the cheapest risk reduction available. It is also the only way you will ever have a number that is actually about your documents.
Bottom line
No independent benchmark exists and none is coming soon, so build your own from a closed deal where you already know the answers. Design it before the demos. Seed the structural edge cases — the scan, the over-length agreement, the missing amendment — because that is where real misses originate.
Then weight recall, coverage integrity, grounding accuracy and verification time above everything else, and ignore throughput entirely. It is the metric most likely to be quoted and the least likely to matter.
Sources
- Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., & Ho, D. E. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. arXiv:2405.20362; Journal of Empirical Legal Studies (2025). Measures legal research, not document review — see the scoping note above. arxiv.org/abs/2405.20362
We publish no benchmark scores, our own or anyone else's, because a figure measured on our chosen documents would not predict yours. We cite only sources we have retrieved and read — see our methodology.