AI Tool Benchmarking and Competitive Analysis

Running fair, decision-grade evaluations of AI due diligence platforms

Updated August 2026 · 6 min read · Deal Room Intelligence Series

There is no independent benchmark for AI due diligence tools. None. No peer-reviewed evaluation, no published methodology, no shared test set — and essentially no vendor publishes prices either. That is not a gap you can research your way around, so the only benchmark that will ever exist for your firm is the one you run yourself. This is how to run it so the result means something.

Why vendor benchmarks tell you nothing

Any accuracy figure a vendor quotes was produced on a test set the vendor designed, scored against a definition of “correct” the vendor chose. Neither is disclosed. That is not necessarily dishonest — there is no shared standard to use instead — but it makes cross-vendor comparison meaningless.

The closest rigorous work in an adjacent domain is Stanford's evaluation of commercial legal AI research tools, which found the products sold by LexisNexis and Thomson Reuters “each hallucinate between 17% and 33% of the time,” and concluded providers' claims were overstated.[1] The instructive part is not the number — it is that an independent evaluation found materially worse performance than the vendors' own descriptions. Expect the same gap here.

Scoping it. That study measures open-ended legal research against case law, not review of a document you supplied — a harder retrieval problem. Do not quote it at a vendor as if it were their score. Use it to justify insisting on your own test.

Design the test before you see a demo

The order matters. A demo installs the vendor's framing — their categories, their idea of what good output looks like — and a test designed afterwards tends to measure the things that vendor happens to do well.

Build ground truth from a closed deal

Take 40–60 documents from a completed transaction. Have someone who worked it list every issue they consider material, before any tool touches it. That list is your ground truth: imperfect, but calibrated to your firm's actual standard rather than a vendor's.

Seed known negatives

Include documents you are confident are unremarkable. Findings returned on those measure the false-positive rate on exactly the material most likely to be waved through.

Seed structural edge cases deliberately

This is where evaluations usually go soft. Include:

These probe the failures that cause real misses, and no vendor demo includes them.

What to measure — and what to ignore

MetricHowWeight
Recall on your ground truthOf your known material issues, how many surfaced?Highest
Coverage integrityWere failures and truncations reported, or silent?Highest
Grounding accuracyDoes the quoted text support the claim?Highest
Verification timeSeconds to confirm one findingHigh
Top-20 precisionOf the 20 highest-ranked, how many were worth reading?High
ReproducibilitySame document twice — same output?High
Overall precisionAcross all findingsLow
Processing speedDocuments per hourIgnore
Feature countIgnore

Three of these deserve explanation because they invert conventional evaluation.

Coverage integrity ranks with recall. A tool that silently skips 8% of the room is not 92% as good — it is producing false confidence, which is worse than a visible failure. Check the processed count against what you submitted before scoring anything else.

Top-20 precision beats overall precision. Reviewers under deadline read the top of the list. A tool with mediocre overall precision and excellent top-20 precision is more useful than the reverse, and standard metrics hide this completely.

Verification time is a first-class metric. Multiply seconds-per-finding by your expected escalation volume and it usually dominates every efficiency claim in the pitch. Two tools with identical recall can differ by an order of magnitude here, entirely on whether findings quote their source.

Run it as a blind comparison if you can

Have someone not involved in procurement strip vendor branding from the outputs and hand the reviewer a merged, shuffled list of findings to assess. Preference for the incumbent, or for the vendor with the better sales team, is real and it survives good intentions.

At minimum, have the person scoring grounding accuracy be someone who did not sit through the demos.

The questions that decide it, which no benchmark captures

After the numbers, three qualitative answers carry disproportionate weight:

  1. “Tell me about a finding your system got wrong in a live deal, and how the customer found out.” Every real product has this story. A vendor without one either has no live deployments or no feedback path — meaning their error rate is unmeasured rather than low.
  2. “How do I distinguish a category you examined and found clean from one you did not examine?” If the answer is that the report simply omits it, the tool cannot tell you what it missed.
  3. “When you update the model, what happens to prior results?” “We're always improving” means last quarter's report is not reproducible.

Scoring it honestly

Weight your criteria before you see results, and write the weights down. Otherwise the tool that wins is the one that happened to perform well on whatever you looked at first.

A defensible default for buy-side diligence: recall on ground truth and coverage integrity at the top, grounding accuracy and verification time next, ranking quality and reproducibility after that, and everything else as a tiebreak. If a tool fails coverage integrity, stop — the rest of the score is measuring a sample you did not choose.

What a pilot costs, and why it is cheap

A proper evaluation takes a few days of one person's time plus, ideally, a paid pilot. Vendors resist paid pilots and should be pushed, because a free demo on vendor-selected documents measures nothing.

Set against a multi-year commitment in a market with no published prices and no independent benchmarks, a few days is the cheapest risk reduction available. It is also the only way you will ever have a number that is actually about your documents.

Bottom line

No independent benchmark exists and none is coming soon, so build your own from a closed deal where you already know the answers. Design it before the demos. Seed the structural edge cases — the scan, the over-length agreement, the missing amendment — because that is where real misses originate.

Then weight recall, coverage integrity, grounding accuracy and verification time above everything else, and ignore throughput entirely. It is the metric most likely to be quoted and the least likely to matter.

Sources

  1. Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., & Ho, D. E. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. arXiv:2405.20362; Journal of Empirical Legal Studies (2025). Measures legal research, not document review — see the scoping note above. arxiv.org/abs/2405.20362

We publish no benchmark scores, our own or anyone else's, because a figure measured on our chosen documents would not predict yours. We cite only sources we have retrieved and read — see our methodology.

Start your benchmark with one of your own documents →

Anweshna Portal
Anweshna Demo