The warning signs that matter are not the obvious ones. A vendor with a thin security page is easy to spot and easy to fix. The ones that cost you are the vendors who answer everything smoothly — because a polished answer to a hard question usually means the question has been asked before and handled, rather than solved.
Tier one: claims that are provably overstated
“Our system doesn't hallucinate”
This is the clearest disqualifying answer available, because it has been tested. Stanford researchers evaluated the retrieval-grounded legal AI products sold by LexisNexis and Thomson Reuters and found they “each hallucinate between 17% and 33% of the time,” concluding that providers' claims of hallucination-free output are overstated.[1]
Those are the two incumbents of legal information, with proprietary curated corpora and enterprise engineering. A smaller vendor asserting they have solved what those two did not is either unaware of the evidence or hoping you are.
“It's 99% accurate”
Ask what the test set was, who built it, and what counted as correct. There is no independent published benchmark for accuracy in M&A document review, so any figure comes from a set the vendor designed. A described methodology with an unflattering number is far more credible than a flattering number with no methodology.
“It reviews contracts as well as a lawyer”
Nobody who has watched a system handle a genuinely novel structure believes this. It signals a vendor selling to people who have not tested the product, which tells you something about their other customers.
Tier two: architecture and output signals
Findings that link to a document but do not quote it
The most consequential product signal on this list, and it is easy to miss in a demo because the demo document is short. If a finding says “change-of-control provision identified” and links to a 90-page agreement, verification means re-reading the agreement — which is the work you are buying the tool to avoid. Quote-backed output is verified in seconds. This single property determines whether the workflow is viable.
No way to distinguish “nothing found” from “not examined”
Ask directly: if a category returns no findings, how do I know the category was actually assessed? If the answer is that the report simply does not mention it, the tool cannot tell you what it missed — and confident omission is the failure mode that reaches closing, precisely because there is nothing on the page to check.
Silence about document length limits
Every system has a ceiling. A vendor who does not raise it is either unaware or truncating quietly. The credit agreement is both the longest document in the room and the one that matters most, so silent truncation lands exactly where it hurts.
Non-reproducible output
Ask them to run the same document twice, live. Material divergence means you cannot reconstruct a decision — and you cannot defend one you cannot reconstruct.
“We're always improving the model” as an answer about versioning
Said as a virtue, it means the system that produced last quarter's report no longer exists. When someone asks how a conclusion was reached, the honest answer becomes a description of a system that has been replaced.
Tier three: commercial and contractual signals
“We don't train on your data — not by default”
A default is a setting, and settings change without your signature. You want a contractual prohibition that flows down to subprocessors and to the underlying model provider. Watch for carve-outs covering telemetry, evaluation sets and human review of outputs — that is where content usually escapes.
No subprocessor list
Several parties see your documents: the model provider, a cloud host, probably a logging platform, possibly OCR. A vendor who cannot name them has not mapped their own data flow, which is a worse finding than any individual name would have been.
AI output specifically carved out of warranties
Liability capped at fees paid is normal and not a red flag. What matters is whether the agreement warrants the software while excluding what the software produces. A vendor who does that has told you where they think the risk sits, and it is with you.
Annual-only commitment with no pilot
Deal flow is lumpy and this category cannot be evaluated from a demo. Refusal to run a paid pilot on your own closed deal — where you already know every material issue — is a refusal to be measured.
Tier four: the behavioural signals
They cannot name a time the system got it wrong
The single most revealing question available: tell me about a finding your system got wrong in a live deal, and how the customer found out. Every real product has this story. A vendor without one either has no live deployments, or no feedback path by which errors return to them — which means their error rate is unmeasured rather than low.
The demo is on their documents
Insist on yours. A demo on vendor-selected documents in a vendor-tuned category with a presenter who knows the answer measures nothing.
Every question gets a smooth answer
Diligence AI has genuinely hard open problems — cross-document reasoning, aggregate risks, severity calibration across sectors, absence detection. A vendor who is fluent on all of them is describing a roadmap. The better sign is a vendor who says “that one we don't do well, here's why, here's the workaround.”
Pressure on timing
Discounts expiring before you can finish a pilot is a sales tactic that survives only where evaluation is skipped.
What good actually looks like
- States a real error rate and describes how it was measured
- Quotes source text on every finding, without being asked
- Reports coverage — processed, failed, truncated — as a first-class output
- Distinguishes an examined-and-clean category from an unassessed one
- Versions models and retains run records
- Puts the training prohibition in the contract, with flow-down
- Names subprocessors
- Volunteers what it is bad at
- Agrees to a paid pilot on your closed deal
The reframe that makes these questions easy
Every signal above reduces to the same underlying test. You are not looking for a system that is never wrong — none exists, and the measured evidence says the best-funded attempts still are. You are looking for one whose errors are cheap to find and whose gaps are visible rather than silent.
That is a property of output design, not of model choice. Which is why the highest-value questions on this page are about quotes, coverage reporting and reproducibility — and why the model underneath, the thing buyers spend the most evaluation time on, barely appears.
Bottom line
Disqualify on the zero-error claim, because it is contradicted by published evidence. Then test the three properties that decide whether the tool works in practice: does every finding quote its source, can you tell a gap from a clean result, and does the same document score the same way twice.
Finally, ask for the story about the time it was wrong. The answer to that question tells you more than the rest of the evaluation combined.
Sources
- Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., & Ho, D. E. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. arXiv:2405.20362; Journal of Empirical Legal Studies (2025). Measures legal research, not document review — see the scoping note above. arxiv.org/abs/2405.20362
We name no competitors on this page. We cite only sources we have retrieved and read; where a figure could not be verified, we changed the figure rather than the citation — see our methodology.