The comparison is usually framed as accuracy: which reads a contract better? That framing produces a bad answer, because on any single document a competent lawyer wins and it is not close. The comparison that decides real outcomes is different — what happens to the documents nobody has time to read — and on that question manual review has no answer at all.
The comparison that misleads
Put one contract in front of an experienced M&A lawyer and one in front of a screening system, and ask who identifies the material issues better. The lawyer wins. They understand the transaction context, the client's risk appetite, what is market for this sector, and what the clause will mean in a negotiation eighteen months out. No current system approaches that.
This comparison is also irrelevant to almost every real decision, because it holds document count constant at one. Live diligence does not hold document count constant. It holds reviewer hours constant, and lets document count vary — usually upward.
The comparison that matters
| Manual only | Screening + human review | |
|---|---|---|
| Documents examined | Whatever hours allow | All of them |
| Documents read closely | Whatever hours allow | Roughly the same number |
| How the read set was chosen | Folder order, upload sequence, instinct | Rank against defined risk categories |
| Quality of judgment on what is read | Higher — undivided expert attention | Same, if verification is cheap |
| Consistency across document 3 and document 900 | Degrades — fatigue is real | Constant |
| Record of the unread material | None | Scored, with source text |
| Error mode | Omission through non-coverage | Misranking, plus model error on what it did read |
Two rows deserve attention.
“Documents read closely” is roughly the same in both columns. This surprises people who expect screening to mean reading less. It usually does not. The hours are similar; what changes is which documents receive them.
The error modes are different in kind, not degree. Manual review fails by never reaching a document. Screening fails by reaching it and ranking it wrongly. The second failure is recoverable — the finding exists, the source text is attached, a reviewer can disagree with it. The first leaves no trace of anything.
Where manual review is strictly better
A page that only lists advantages of the new thing is marketing. These are real:
- Contextual judgment. Whether a provision matters depends on the transaction, the counterparty relationship, the client's tolerance and the negotiation ahead. Screening has none of that context and should not pretend to.
- Novel structures. An unusual arrangement with no analogue in the rubric will be scored against categories that do not fit. A human notices something strange; a system scores it against what it knows.
- The inference across silence. Experienced reviewers notice what is missing from a document — the indemnity that should be there and is not. This is inferential, contextual, and largely beyond current systems.
- Small document sets. Below a few dozen documents, screening adds process overhead for little benefit. Read them.
Where manual review fails predictably
- Coverage at volume. The binding constraint, and it is arithmetic rather than skill. Hours are fixed; documents are not.
- Consistency across a long set. The same clause type assessed on document 900 does not receive the attention it received on document 3. This is fatigue, not incompetence, and no professional is exempt.
- Cross-document connections. A cross-default linking two facilities reviewed by two different people on two different days is a structural blind spot in any divided workstream.
- Reconstructing the decision later. Ask why a specific document was not escalated, six months on. Manual review usually has no answer, because the reason was that nobody opened it.
Why the hybrid is not a compromise
“Hybrid” sounds like splitting the difference. It is not — it is assigning two genuinely different capabilities to the tasks each is actually good at.
The division is clean once you name it correctly. Screening allocates attention. Humans make judgments. Those are separate jobs, and the reason manual-only review struggles at volume is that it forces the second capability to perform the first — expert judgment spent on triage, which is expensive, slow, and not what the expertise is for.
Sequenced properly:
- Screen the full set against defined categories. Every document scored, source text attached, blocking categories floored so absence of a finding is visible rather than comfortable.
- Verify the escalations. Read the quote against the claim. Cheap if output is quote-backed; prohibitive if it is not — which is the property that decides whether the whole model works.
- Apply judgment to what survives, with full expert attention, on a set small enough to deserve it.
- Sample the tail nobody would otherwise read, to measure the real error rate rather than assume it.
- Record overrides — they are the evidence judgment was applied, and the calibration data for next time.
The cost that decides it
Whether a hybrid workflow beats manual review comes down to one variable that rarely appears in evaluations: the cost of verifying a finding.
If a system quotes the source sentence beside its claim, verification takes seconds and the model works. If it produces fluent unattributed prose, verification means re-reading the document — and the hybrid collapses, because you have added a step without removing one.
This matters more given what is known about error rates. Grounded commercial legal AI has been measured hallucinating between 17% and 33% of the time in an adjacent task,[1] and no vendor has eliminated it. Verification is therefore permanent, not transitional. Its per-finding cost is not a detail of the tool; it is the thing that determines whether the workflow is viable at all.
What does not change
Materiality stays human. So does the negotiation, the structuring, and the decision to walk. And neither approach reads what was never uploaded — the oral side agreement and the withheld amendment are invisible to both columns of the table, which is why coverage reporting against an expected document list matters more than either.
Bottom line
Stop comparing on single-document accuracy. On one contract, the lawyer wins. On twelve hundred, the question is which documents the lawyer reads, and manual-only review answers that with folder order.
The durable operating model is not a compromise between the two. It is screening to allocate attention and humans to exercise judgment — with the caveat that it only works if verifying a finding costs seconds. Buy for that property, because it is the one the whole workflow rests on.
Sources
- Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., & Ho, D. E. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. arXiv:2405.20362; Journal of Empirical Legal Studies (2025). Measures legal research, not document review — see the scoping note above. arxiv.org/abs/2405.20362
We cite only sources we have retrieved and read. Where a figure could not be verified, we changed the figure rather than the citation — see our methodology.