Red flags are rarely a single dramatic sentence. They are patterns — a provision that is unremarkable alone and material at scale, an inconsistency between two documents drafted eighteen months apart, an obligation attached to the wrong entity. AI surfaces these when it is given a clear definition of “unusual” and the whole document set. It fails, predictably, when either is missing.
The five red-flag classes, and which are findable
| Class | Example | Findable by screening? |
|---|---|---|
| Clause-level | Change-of-control in a mid-tier customer contract | Yes — the strongest case |
| Cross-document | Cross-default linking two separate facilities | Only with genuine cross-document analysis |
| Aggregate | Customer concentration above tolerance | Requires summing across the set, not extraction |
| Absence | The IP assignment that should exist and does not | Only if the system checks against an expected list |
| External | Litigation or enforcement not disclosed in the room | No — not in the documents at all |
Most discussion of this topic covers only the first row. The other four are where expensive surprises actually live, and each needs a different capability — which is why “can it find red flags?” is not a question with one answer.
Why humans miss clause-level flags, and it isn't skill
A senior reviewer reading a change-of-control provision will assess it correctly every time. The problem is arithmetic. In a 1,200-document room with capacity for perhaps 400, the 800 unread documents are selected by folder structure, upload order, and whoever mentioned something on a call — none of which correlate with risk.
That is a smaller claim than the category usually makes, and it is the one that survives a deal where something was missed — because it was never a promise of detection, only of coverage and ranking, and the record shows both.
The absence class, which almost nothing handles well
The most dangerous red flag is a document that says nothing where it should say something. An employment agreement with no IP assignment. A supply contract with no force majeure. A base agreement whose referenced fourth amendment is not in the room.
Extraction-based systems are structurally blind here — they report what they find, and finding nothing produces no finding. The output is a clean-looking report on a genuine gap.
Two design choices fix it, and they are worth asking any vendor about directly:
- Check against an expected list rather than only reporting what exists. Which permits should this business hold? Which consents does this structure require?
- Floor the blocking categories. A category with zero findings should not score zero. It should carry an explicit non-zero floor plus a stated diligence request, so “we examined this and found no disclosure” reads differently from “this looks fine.” This single choice converts the hardest failure mode into a visible one.
Aggregate flags need summing, not reading
Customer concentration is written in no contract. It emerges from reading forty and adding up. The same is true of aggregate indemnity exposure, total change-of-control-triggered payments, and how many employment agreements lack a non-compete.
Ask specifically whether anything in the product operates across the set rather than within a document. Many do not, and it is not obvious from a demo, because a demo runs on one document at a time.
Ranking matters more than detection
A system tuned for high recall will find a great deal. That is only useful if the output is ordered, because a flat list of 400 flags and a ranked list of the same 400 have identical accuracy and completely different value.
The reason is behavioural rather than technical. By the fortieth “critical” finding that turns out to be market-standard, reviewers stop reading carefully — and the forty-first, which is real, receives the same skim. False positives are not independently cheap; they compound, and what they consume is the attention the tool exists to direct.
So the useful configuration is not maximum sensitivity. It is high recall with severity ranked against a rubric your firm actually agrees with, calibrated against your own precedent — the clause types your committee has waved through for a decade should not be arriving as critical.
Verify before you escalate
Every red flag that reaches a committee will be challenged, by the seller and internally. It needs to survive being read.
This is not optional caution. Grounded commercial legal AI has been measured hallucinating between 17% and 33% of the time in an adjacent task, with providers' hallucination-free claims found to be overstated.[1] The failure mode that matters most here is the misgrounded citation — a real clause cited for a proposition it does not support. It passes a click-the-link check and fails a read-the-clause check.
Practically: for every finding that will appear in the pack, open the cited document and read the cited sentence against the claim. If the tool quotes source text this takes seconds. If it only links to a document, it means re-reading — and at that point the escalation process costs more than it saves.
A workable structure
- Define blocking categories before the deal. A short list of what genuinely stops a transaction in your sector, with a score threshold that escalates immediately. If everything blocks, nothing does.
- Run on day one, on whatever exists. A red flag found before the workstream runs saves the workstream. Found in week four, it saves nothing.
- Score the entire set, so unread documents are ranked rather than ignored.
- Verify escalations against source text.
- Calibrate severity against your own history, not vendor defaults.
- Sample the low-severity tail — errors live where attention does not, and this is the only way to measure your real miss rate rather than assume it.
- Record overrides. They are the evidence judgment was applied, and the calibration set for next time.
Limits to state to your committee
- Screening cannot see oral side agreements or documents never uploaded — and a clean result on an incomplete data room is a confident answer to the wrong question.
- Whether a flag is genuinely deal-killing is a human call, and should stay one.
- A clause read correctly but scored against the wrong sector or jurisdiction produces a low-severity finding indistinguishable from a correct one.
- Coverage reporting matters more than finding quality. A document that failed to parse produces zero findings, and zero findings looks exactly like a clean document.
Bottom line
AI is a coverage instrument, not a judgment instrument. It finds the ordinary provision in the document nobody was going to reach — which is where most expensive surprises actually come from, because they are rarely exotic.
Define what blocks before you start, make absence of a finding visible rather than comfortable, rank rather than merely detect, verify what you escalate, and keep the materiality call with the people who can be accountable for it.
Sources
- Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., & Ho, D. E. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. arXiv:2405.20362; Journal of Empirical Legal Studies (2025). Measures legal research, not red-flag detection — see the scoping note above. arxiv.org/abs/2405.20362
An earlier version of this page carried an unsourced percentage claim about AI identifying more risk provisions than manual review, presented under a “documented impact” heading. We could not trace it to any source and have removed it. We are not restating the figure here, because an invented number does not become harmless by being disavowed. We cite only sources we have retrieved and read — see our methodology.