We are not going to invent case studies. Deal work is confidential, the named-firm anecdotes circulating in this market are almost entirely vendor marketing, and a fabricated example would be worse than none — it would read as authoritative and be unfalsifiable. What can be described honestly are the deployment patterns: the recurring shapes an AI diligence rollout takes, and which ones survive past the first deal.
Why the case studies you find elsewhere are thin
Three structural reasons, worth naming so you can discount accordingly:
- Confidentiality. The interesting detail — which provision was found, in which document, on which transaction — is exactly what cannot be published. What survives publication is usually the part that carries no information.
- Selection. Vendors publish the deployments that worked. The ones that stalled after two deals do not become case studies, and they are the more instructive population.
- Attribution. “The tool found an issue that saved the deal” is unfalsifiable. Nobody can show the issue would have been missed, because the counterfactual never ran.
The same problem applies to adoption statistics generally: the figures in circulation trace to vendor-sponsored fieldwork, and we could not locate an independent measurement of how firms use these tools.
Pattern 1 — The long-tail sweep
Shape: screening is pointed at commercial contracts ranked roughly eleven through two hundred — the ones below the threshold anyone was going to read individually — while the top ten and the credit agreement continue to get full human review.
Why it persists: it does not ask anyone to trust the system with the documents they care most about. The tool operates entirely in territory that was previously unexamined, so the comparison is against nothing rather than against expert review. That makes it easy to adopt and hard to argue with.
Where it fails: nowhere interesting. This is the most durable pattern and the one to start with.
Pattern 2 — The day-one blocking pass
Shape: a narrow scan for a short list of transaction-ending issues, run on whatever documents exist at kickoff, re-run as the room fills.
Why it persists: the economics of diligence are front-loaded. An issue found on day two stops the workstream before advisers bill; the same issue in week four saves nothing. This pattern targets the largest financial effect available, and it does so with a deliberately small rubric — a handful of categories rather than a full analysis.
Where it fails: when teams widen the blocking list until everything blocks, at which point nothing does. The discipline is keeping it short.
Pattern 3 — The consistency sweep on employment documents
Shape: hundreds of near-identical agreements checked for the same handful of provisions — IP assignment, non-compete, change-of-control triggers, severance.
Why it persists: this is where human attention genuinely degrades, and everyone knows it. The interesting output is usually an aggregate rather than a finding: how many agreements lack IP assignment. That is a coverage question, unanswerable without processing all of them, and it is the kind of number that changes a purchase price.
Where it fails: tools that only report what they find, rather than checking against an expected list, cannot answer the absence question at all.
Pattern 4 — Cross-border, routing local counsel
Shape: screening runs per-jurisdiction and produces a ranked set for local counsel, rather than local counsel receiving whatever the deal team could tell was relevant.
Why it persists: local counsel time is the bottleneck on international deals — expensive, engaged late, in another time zone. Ranking changes what reaches them. More valuably, their corrections (“that clause is routine here” / “the one you ranked low is a blocker”) become the calibration for the next deal in that jurisdiction.
Where it fails: when jurisdiction is set per-deal rather than per-document. A target's contracts may sit under four governing laws, and a single deal-level setting applies the wrong baseline to most of the set.
Pattern 5 — The pattern that does not survive
Shape: conversational Q&A over the data room, used as the primary coverage layer. Someone asks the system questions and treats the answers as the review.
Why it fails: it only tells you about things you thought to ask, and the expensive misses in diligence are the provisions nobody knew to ask about, in documents nobody prioritised. It demos superbly and produces false confidence, because the answers it does give are fluent and specific.
Q&A is a genuinely useful supplement over a document you already know matters. It is not a coverage instrument, and deployments that treat it as one tend to be abandoned after something is missed.
What separates deployments that stick
| Sticks | Stalls |
|---|---|
| Runs parallel to the existing process on deal one | Changes what people read immediately |
| Thresholds set in writing before a live deal | Thresholds decided case by case |
| Overrides captured and fed back | Reviewers approve everything |
| Named owner for calibration and sampling | Available to everyone, nobody's job |
| Findings quote source text | Verification costs a re-read, so it lapses |
| Coverage reconciled before findings are read | Findings trusted without checking what processed |
The last two rows are the ones that decide it, and both are properties of the tool rather than of the team. Grounded commercial legal AI has been measured hallucinating between 17% and 33% of the time in an adjacent task, with providers' hallucination-free claims judged overstated.[1] Verification is therefore permanent — and any deployment where verification costs four minutes per finding will quietly stop verifying in week three of a live deal, which is when it mattered.
The failure nobody publishes
The most common unsuccessful deployment is not a dramatic miss. It is a licence that renews while the tool is used by one person on two deals a year — the champion who ran the evaluation.
It is invisible in every case-study collection because nothing went wrong. Nothing went right either. The distinguishing feature, consistently, is that the operating model was never changed: the tool was added alongside the existing process and nobody's job was to make it part of the process.
How to build your own case study
Since the published ones are of little use, instrument your own across three deals:
- Documents in the room versus documents anyone opened. Most firms have never calculated this. It is usually uncomfortable and entirely defensible, because it is your own data.
- Override rate by category. Routinely downgraded means the rubric is mis-tuned; routinely upgraded means under-detection.
- Verification minutes per escalated finding. Multiply by escalation volume — this is your real bottleneck, and it is the number that decides whether any of it pays.
Three deals of that produces something specific to your firm, your sector and your document mix — which is exactly why someone else's case study was never going to tell you much.
Bottom line
The published case studies in this market are selected, unfalsifiable, and stripped of the detail that would make them useful. The patterns are more informative: sweep the long tail, run a short blocking pass on day one, use consistency checks where human attention degrades, and route local counsel by rank on cross-border work.
Then watch the two things that actually determine survival — whether verification is cheap enough to keep happening, and whether anyone owns the calibration.
Sources
- Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., & Ho, D. E. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. arXiv:2405.20362; Journal of Empirical Legal Studies (2025). Measures legal research, not document review — see the scoping note above. arxiv.org/abs/2405.20362
This page deliberately contains no named-firm case studies. We have none we can verify and publish, and we will not invent them — see our methodology.