Red Flags to Watch When Selecting AI Vendors

Warning signs that often predict under-performance or elevated risk

Updated August 2026 · 7 min read · Deal Room Intelligence Series

The warning signs that matter are not the obvious ones. A vendor with a thin security page is easy to spot and easy to fix. The ones that cost you are the vendors who answer everything smoothly — because a polished answer to a hard question usually means the question has been asked before and handled, rather than solved.

Tier one: claims that are provably overstated

“Our system doesn't hallucinate”

This is the clearest disqualifying answer available, because it has been tested. Stanford researchers evaluated the retrieval-grounded legal AI products sold by LexisNexis and Thomson Reuters and found they “each hallucinate between 17% and 33% of the time,” concluding that providers' claims of hallucination-free output are overstated.[1]

Those are the two incumbents of legal information, with proprietary curated corpora and enterprise engineering. A smaller vendor asserting they have solved what those two did not is either unaware of the evidence or hoping you are.

Use this fairly. The study measures open-ended legal research against case law, not review of a document you supplied — a harder retrieval problem. A sharp vendor will point that out and they will be right. The red flag is not a non-zero error rate, which every honest vendor has. It is the claim of zero.

“It's 99% accurate”

Ask what the test set was, who built it, and what counted as correct. There is no independent published benchmark for accuracy in M&A document review, so any figure comes from a set the vendor designed. A described methodology with an unflattering number is far more credible than a flattering number with no methodology.

“It reviews contracts as well as a lawyer”

Nobody who has watched a system handle a genuinely novel structure believes this. It signals a vendor selling to people who have not tested the product, which tells you something about their other customers.

Tier two: architecture and output signals

Findings that link to a document but do not quote it

The most consequential product signal on this list, and it is easy to miss in a demo because the demo document is short. If a finding says “change-of-control provision identified” and links to a 90-page agreement, verification means re-reading the agreement — which is the work you are buying the tool to avoid. Quote-backed output is verified in seconds. This single property determines whether the workflow is viable.

No way to distinguish “nothing found” from “not examined”

Ask directly: if a category returns no findings, how do I know the category was actually assessed? If the answer is that the report simply does not mention it, the tool cannot tell you what it missed — and confident omission is the failure mode that reaches closing, precisely because there is nothing on the page to check.

Silence about document length limits

Every system has a ceiling. A vendor who does not raise it is either unaware or truncating quietly. The credit agreement is both the longest document in the room and the one that matters most, so silent truncation lands exactly where it hurts.

Non-reproducible output

Ask them to run the same document twice, live. Material divergence means you cannot reconstruct a decision — and you cannot defend one you cannot reconstruct.

“We're always improving the model” as an answer about versioning

Said as a virtue, it means the system that produced last quarter's report no longer exists. When someone asks how a conclusion was reached, the honest answer becomes a description of a system that has been replaced.

Tier three: commercial and contractual signals

“We don't train on your data — not by default”

A default is a setting, and settings change without your signature. You want a contractual prohibition that flows down to subprocessors and to the underlying model provider. Watch for carve-outs covering telemetry, evaluation sets and human review of outputs — that is where content usually escapes.

No subprocessor list

Several parties see your documents: the model provider, a cloud host, probably a logging platform, possibly OCR. A vendor who cannot name them has not mapped their own data flow, which is a worse finding than any individual name would have been.

AI output specifically carved out of warranties

Liability capped at fees paid is normal and not a red flag. What matters is whether the agreement warrants the software while excluding what the software produces. A vendor who does that has told you where they think the risk sits, and it is with you.

Annual-only commitment with no pilot

Deal flow is lumpy and this category cannot be evaluated from a demo. Refusal to run a paid pilot on your own closed deal — where you already know every material issue — is a refusal to be measured.

Tier four: the behavioural signals

They cannot name a time the system got it wrong

The single most revealing question available: tell me about a finding your system got wrong in a live deal, and how the customer found out. Every real product has this story. A vendor without one either has no live deployments, or no feedback path by which errors return to them — which means their error rate is unmeasured rather than low.

The demo is on their documents

Insist on yours. A demo on vendor-selected documents in a vendor-tuned category with a presenter who knows the answer measures nothing.

Every question gets a smooth answer

Diligence AI has genuinely hard open problems — cross-document reasoning, aggregate risks, severity calibration across sectors, absence detection. A vendor who is fluent on all of them is describing a roadmap. The better sign is a vendor who says “that one we don't do well, here's why, here's the workaround.”

Pressure on timing

Discounts expiring before you can finish a pilot is a sales tactic that survives only where evaluation is skipped.

What good actually looks like

The reframe that makes these questions easy

Every signal above reduces to the same underlying test. You are not looking for a system that is never wrong — none exists, and the measured evidence says the best-funded attempts still are. You are looking for one whose errors are cheap to find and whose gaps are visible rather than silent.

That is a property of output design, not of model choice. Which is why the highest-value questions on this page are about quotes, coverage reporting and reproducibility — and why the model underneath, the thing buyers spend the most evaluation time on, barely appears.

Bottom line

Disqualify on the zero-error claim, because it is contradicted by published evidence. Then test the three properties that decide whether the tool works in practice: does every finding quote its source, can you tell a gap from a clean result, and does the same document score the same way twice.

Finally, ask for the story about the time it was wrong. The answer to that question tells you more than the rest of the evaluation combined.

Sources

  1. Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., & Ho, D. E. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. arXiv:2405.20362; Journal of Empirical Legal Studies (2025). Measures legal research, not document review — see the scoping note above. arxiv.org/abs/2405.20362

We name no competitors on this page. We cite only sources we have retrieved and read; where a figure could not be verified, we changed the figure rather than the citation — see our methodology.

Test these questions against our own output →

Anweshna Portal
Anweshna Demo