On one contract, no — and it is not close. A competent M&A lawyer brings transaction context, client risk appetite, sector norms and a sense of how a clause will behave in a negotiation eighteen months out. No current system approaches that. But the comparison holds document count at one, and live diligence never does. It holds reviewer hours fixed and lets document count rise — and on that comparison the answer changes completely.
Where the lawyer wins decisively
- Contextual materiality. Whether a provision matters depends on the deal, the counterparty relationship, the client's tolerance and the leverage available. None of that is in the document.
- Inference across silence. Noticing the indemnity that should be there and is not. This requires a model of what ought to exist, not what does — and it is where experienced reviewers add the most value.
- Novel structures. An unusual arrangement with no analogue in the rubric gets scored against categories that do not fit. A human notices something strange; a system scores it against what it knows.
- Reading the negotiation. Which provisions the other side will actually defend, and what they signal about the counterparty's own concerns.
- Being accountable. ABA Formal Opinion 512 (29 July 2024) holds that lawyers using generative AI must “fully consider their applicable ethical obligations,” including competence and supervisory responsibility.[1] That duty does not transfer, and no vendor accepts it.
Where the lawyer loses, and it is not about skill
- Coverage at volume. Hours are fixed; documents are not. A team with capacity for 400 of 1,200 documents reads a 400 selected by folder order and upload sequence — which correlate with nothing.
- Consistency across a long set. The same clause type on document 900 does not receive the attention it received on document 3. This is fatigue, not incompetence, and no professional is exempt from it.
- Cross-document connections. A cross-default linking two facilities reviewed by two people on two days is a structural blind spot in any divided workstream.
- Aggregates. Customer concentration is written in no contract. It emerges from reading forty and adding up.
- Reconstructing the decision later. Ask why a specific document was not escalated, six months on. Usually there is no answer, because nobody opened it.
The comparison that decides real outcomes
| Lawyer alone | Screening + lawyer | |
|---|---|---|
| Quality of judgment on what is read | Higher | Same, if verification is cheap |
| Documents examined | Whatever hours allow | All of them |
| How the read set was chosen | Folder order, instinct | Rank against defined categories |
| Record of the unread material | None | Scored, with source text |
| Error mode | Omission through non-coverage | Misranking — recoverable, leaves a trace |
The bottom row matters more than the top one. Manual review fails by never reaching a document, which leaves no trace of anything. Screening fails by reaching it and ranking it wrongly — and that failure is recoverable, because the finding exists, the source text is attached, and a reviewer can disagree with it on the record.
What the measured evidence says about the machine side
Stanford researchers evaluated the retrieval-grounded legal AI research tools sold by LexisNexis and Thomson Reuters and found they “each hallucinate between 17% and 33% of the time,” concluding that providers' hallucination-free claims are overstated.[2]
Read against the question: a system with that error profile cannot replace a lawyer's judgment. It can, reliably, tell you which forty of twelve hundred documents deserve one — which is a different job that nobody was doing well.
The framing that survives contact with a bad outcome
Do not tell a client or committee that AI reviews contracts as well as a lawyer. The first time a finding is wrong, that claim collapses and takes the rest of your credibility with it.
That is a smaller claim. It is also true, verifiable, and it holds up when something is missed — because it never promised detection, only coverage and ranking, and the record evidences both.
The variable that decides whether the pairing works
One thing determines whether lawyer-plus-screening beats lawyer-alone: the cost of verifying a single finding.
If a finding quotes its source sentence, verification takes seconds and the arrangement works. If it links to a 90-page agreement without quoting, verification means re-reading — and you have added a step without removing one, making the pairing strictly worse than manual review.
This is why output format matters more than model quality when selecting a tool. The model is the layer you cannot inspect or control. The output shape determines your lawyer's actual workload, permanently.
Bottom line
On one contract the lawyer wins and will keep winning, because the things they bring are not in the document. On twelve hundred contracts the question is which ones the lawyer reads — and answering that with folder order is the status quo, not a standard worth defending.
So the honest answer is that AI does not review contracts as well as a lawyer, and does not need to. It needs to read everything well enough to rank it, quote its source so the lawyer can check in seconds, and leave a record of what it ranked low. The judgment stays where it was.
Sources
- ABA Standing Committee on Ethics and Professional Responsibility, Formal Opinion 512: Generative Artificial Intelligence Tools, 29 July 2024. americanbar.org
- Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., & Ho, D. E. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. arXiv:2405.20362; Journal of Empirical Legal Studies (2025). arxiv.org/abs/2405.20362
Nothing here is legal advice. We cite only sources we have retrieved and read — see our methodology.