Case Studies: How Firms Actually Use AI in Diligence

Patterns from real deployments rather than vendor slide decks

Updated August 2026 · 7 min read · Deal Room Intelligence Series

We are not going to invent case studies. Deal work is confidential, the named-firm anecdotes circulating in this market are almost entirely vendor marketing, and a fabricated example would be worse than none — it would read as authoritative and be unfalsifiable. What can be described honestly are the deployment patterns: the recurring shapes an AI diligence rollout takes, and which ones survive past the first deal.

Why the case studies you find elsewhere are thin

Three structural reasons, worth naming so you can discount accordingly:

The same problem applies to adoption statistics generally: the figures in circulation trace to vendor-sponsored fieldwork, and we could not locate an independent measurement of how firms use these tools.

Pattern 1 — The long-tail sweep

Shape: screening is pointed at commercial contracts ranked roughly eleven through two hundred — the ones below the threshold anyone was going to read individually — while the top ten and the credit agreement continue to get full human review.

Why it persists: it does not ask anyone to trust the system with the documents they care most about. The tool operates entirely in territory that was previously unexamined, so the comparison is against nothing rather than against expert review. That makes it easy to adopt and hard to argue with.

Where it fails: nowhere interesting. This is the most durable pattern and the one to start with.

Pattern 2 — The day-one blocking pass

Shape: a narrow scan for a short list of transaction-ending issues, run on whatever documents exist at kickoff, re-run as the room fills.

Why it persists: the economics of diligence are front-loaded. An issue found on day two stops the workstream before advisers bill; the same issue in week four saves nothing. This pattern targets the largest financial effect available, and it does so with a deliberately small rubric — a handful of categories rather than a full analysis.

Where it fails: when teams widen the blocking list until everything blocks, at which point nothing does. The discipline is keeping it short.

Pattern 3 — The consistency sweep on employment documents

Shape: hundreds of near-identical agreements checked for the same handful of provisions — IP assignment, non-compete, change-of-control triggers, severance.

Why it persists: this is where human attention genuinely degrades, and everyone knows it. The interesting output is usually an aggregate rather than a finding: how many agreements lack IP assignment. That is a coverage question, unanswerable without processing all of them, and it is the kind of number that changes a purchase price.

Where it fails: tools that only report what they find, rather than checking against an expected list, cannot answer the absence question at all.

Pattern 4 — Cross-border, routing local counsel

Shape: screening runs per-jurisdiction and produces a ranked set for local counsel, rather than local counsel receiving whatever the deal team could tell was relevant.

Why it persists: local counsel time is the bottleneck on international deals — expensive, engaged late, in another time zone. Ranking changes what reaches them. More valuably, their corrections (“that clause is routine here” / “the one you ranked low is a blocker”) become the calibration for the next deal in that jurisdiction.

Where it fails: when jurisdiction is set per-deal rather than per-document. A target's contracts may sit under four governing laws, and a single deal-level setting applies the wrong baseline to most of the set.

Pattern 5 — The pattern that does not survive

Shape: conversational Q&A over the data room, used as the primary coverage layer. Someone asks the system questions and treats the answers as the review.

Why it fails: it only tells you about things you thought to ask, and the expensive misses in diligence are the provisions nobody knew to ask about, in documents nobody prioritised. It demos superbly and produces false confidence, because the answers it does give are fluent and specific.

Q&A is a genuinely useful supplement over a document you already know matters. It is not a coverage instrument, and deployments that treat it as one tend to be abandoned after something is missed.

What separates deployments that stick

SticksStalls
Runs parallel to the existing process on deal oneChanges what people read immediately
Thresholds set in writing before a live dealThresholds decided case by case
Overrides captured and fed backReviewers approve everything
Named owner for calibration and samplingAvailable to everyone, nobody's job
Findings quote source textVerification costs a re-read, so it lapses
Coverage reconciled before findings are readFindings trusted without checking what processed

The last two rows are the ones that decide it, and both are properties of the tool rather than of the team. Grounded commercial legal AI has been measured hallucinating between 17% and 33% of the time in an adjacent task, with providers' hallucination-free claims judged overstated.[1] Verification is therefore permanent — and any deployment where verification costs four minutes per finding will quietly stop verifying in week three of a live deal, which is when it mattered.

Scoping that figure. It measures open-ended legal research against case law, not review of a document you supplied — a harder retrieval problem. It establishes that errors are a permanent design assumption, not a diligence-accuracy rate.

The failure nobody publishes

The most common unsuccessful deployment is not a dramatic miss. It is a licence that renews while the tool is used by one person on two deals a year — the champion who ran the evaluation.

It is invisible in every case-study collection because nothing went wrong. Nothing went right either. The distinguishing feature, consistently, is that the operating model was never changed: the tool was added alongside the existing process and nobody's job was to make it part of the process.

How to build your own case study

Since the published ones are of little use, instrument your own across three deals:

  1. Documents in the room versus documents anyone opened. Most firms have never calculated this. It is usually uncomfortable and entirely defensible, because it is your own data.
  2. Override rate by category. Routinely downgraded means the rubric is mis-tuned; routinely upgraded means under-detection.
  3. Verification minutes per escalated finding. Multiply by escalation volume — this is your real bottleneck, and it is the number that decides whether any of it pays.

Three deals of that produces something specific to your firm, your sector and your document mix — which is exactly why someone else's case study was never going to tell you much.

Bottom line

The published case studies in this market are selected, unfalsifiable, and stripped of the detail that would make them useful. The patterns are more informative: sweep the long tail, run a short blocking pass on day one, use consistency checks where human attention degrades, and route local counsel by rank on cross-border work.

Then watch the two things that actually determine survival — whether verification is cheap enough to keep happening, and whether anyone owns the calibration.

Sources

  1. Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., & Ho, D. E. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. arXiv:2405.20362; Journal of Empirical Legal Studies (2025). Measures legal research, not document review — see the scoping note above. arxiv.org/abs/2405.20362

This page deliberately contains no named-firm case studies. We have none we can verify and publish, and we will not invent them — see our methodology.

Start your own measurement with one document →

Anweshna Portal
Anweshna Demo