What Questions Should I Ask Vendors About AI Hallucinations?

Learning about quality assurance, grounding, and reliability controls

Updated August 2026 · 7 min read · Deal Room Intelligence Series

Every vendor has a prepared answer to “does it hallucinate?” None of them have a prepared answer to “show me a finding your system got wrong last month, and tell me how you found out.” The difference between those two questions is the whole of a useful diligence conversation.

Why the obvious question fails

Ask a vendor whether their system hallucinates and you will hear some version of: we use retrieval-augmented generation, the model only answers from your documents, so it cannot invent.

That answer describes an architecture that has been measured. Stanford researchers tested the retrieval-grounded legal AI products sold by LexisNexis and Thomson Reuters and reported that they “each hallucinate between 17% and 33% of the time” — concluding that providers' claims of hallucination-free output are overstated.[1]

So grounding is not a defence, it is table stakes that demonstrably does not close the gap. The productive questions are not about whether errors occur. They are about what the system does when it is uncertain, and how anyone would ever know it was wrong.

Cite this accurately or it will be turned on you. That study measured legal research against case law, not document review against a fixed data room — a meaningfully easier retrieval problem. A sharp vendor will say so, and they will be right. Use the finding to establish that the category has a real error rate nobody has eliminated, not as a number that applies to their product.

Twelve questions, and what each one is really testing

Grounding and output shape

1. Does every assertion carry the source sentence, or only a document reference?
A link to a 90-page agreement is not a citation, it is a suggestion to go and read it. You are testing whether verification costs seconds or costs re-doing the work.

2. What does the system output when it finds supporting text it cannot fully interpret?
The honest answer describes a lower-confidence state. The concerning answer is that output looks identical regardless — because a well-grounded finding and a shaky inference rendered in the same typeface are indistinguishable to a reviewer at 11pm.

3. Can it produce a finding with no quote attached? What happens to it?
Systems that permit ungrounded findings are not necessarily worse — sometimes the right answer genuinely is inferential. What matters is whether those findings are visibly demoted or silently mixed in with the grounded ones.

Coverage — the failure nobody asks about

4. How do I distinguish “we examined this and found nothing” from “we did not examine this”?
This is the highest-value question on the list and the one vendors are least ready for. Confident omission is the hardest hallucination class to catch, because there is nothing on the page to verify. If the answer is “the report just doesn't mention it,” the tool cannot tell you what it missed.

5. What is the document length limit, and what happens above it?
You are listening for silent truncation. An explicit error or a documented routing path to a different processing mode is fine. “It handles any length” is not an answer any system can honestly give.

6. If a category returns nothing, does that show as a zero or as a gap?
A clean zero on a high-stakes category is a claim, and a strong one. Systems that assign an explicit floor when nothing was found — rather than reporting comfortable silence — are making a deliberate choice not to let absence read as safety.

Reproducibility and audit

7. Run the same document twice. Do I get the same output?
Ask them to do it live. Material divergence means you cannot reconstruct a decision, which means you cannot defend one.

8. When you update the model, what happens to prior results?
“We're always improving” is a warning. If the system that produced last quarter's report no longer exists, that report is no longer reproducible, and a regulator or claimant asking how a conclusion was reached will get an answer about a system that has been replaced.

9. What record survives of why a specific document was flagged?
The defensible minimum is: the source text, the rule or category applied, the resulting value, the threshold it was measured against, and the model version. Anything less and “why did you escalate this?” has no answer.

Data handling

10. Is our content used for training — contractually, not by policy?
“Not by default” means there is a default that could change. You want a prohibition in the agreement.

11. What is retained after a deal closes, where, and for how long?
Deal documents are among the most sensitive material a firm handles, and they stop being needed the moment the transaction completes.

12. Which subprocessors see the content?
The model provider is one. Ask about the rest — logging, monitoring, support tooling — because that is where deal data usually leaks in ways nobody intended.

The single best question, if you only get one

“Tell me about a finding your system got wrong in a live deal. How did the customer find out?”

Every real product has this story. A vendor who cannot produce one either has no live deployments, or is not being straight with you, or does not have a feedback path by which errors ever come back to them — and the third possibility is the most alarming, because it means their error rate is unmeasured rather than low.

What you want to hear is specific and slightly uncomfortable: the class of error, how it surfaced, what changed as a result. What you do not want is a smooth pivot to a case study.

What a good answer to all of this looks like structurally

The pattern across every question above is the same: you are not trying to find a system that does not err. You are trying to find one whose errors are cheap to catch. That is a property of output design, not model selection.

Concretely, it means preferring output that is scored against defined categories over output that is written as prose. A value per category, carrying the source text behind it, checked against a stated threshold, gives you four things narrative cannot:

To be precise about what this does and does not do: structured scoring does not lower the hallucination rate. It lowers the cost of detecting hallucinations, and it makes the omissions visible. Given that no vendor in this category has eliminated the underlying error rate, that is the property worth buying for.

The obligation that sits behind the procurement

ABA Formal Opinion 512 (29 July 2024) — the first formal ABA ethics guidance on generative AI — holds that lawyers using these tools must “fully consider their applicable ethical obligations,” including competence, confidentiality, client communication, supervision, and reasonable fees.[2]

Note where that duty sits. It attaches to the output you rely on, not to the vendor you selected. No answer in this document transfers it. The questions are worth asking because they determine how hard the duty is to discharge, not whether you still carry it.

Bottom line

Stop asking whether it hallucinates. Ask what it does when uncertain, how you would tell a gap from a clean result, whether the same input produces the same output, and what record survives. Then ask for the story about the time it was wrong.

A vendor who answers those precisely — including the unflattering parts — is describing a system you can build a defensible process on. A vendor who answers them smoothly and generally is describing a demo.

Sources

  1. Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., & Ho, D. E. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. arXiv:2405.20362; published in the Journal of Empirical Legal Studies (2025). arxiv.org/abs/2405.20362
  2. American Bar Association Standing Committee on Ethics and Professional Responsibility, Formal Opinion 512: Generative Artificial Intelligence Tools, 29 July 2024. americanbar.org

We cite only sources we have retrieved and read. Where a figure could not be verified, we changed the figure rather than the citation — see our methodology.

See what quote-backed, scored output looks like →

Anweshna Portal
Anweshna Demo