Can AI Review Contracts as Well as a Lawyer?

Comparing AI capabilities to human expertise — without the marketing spin

Updated August 2026 · 5 min read · Deal Room Intelligence Series

On one contract, no — and it is not close. A competent M&A lawyer brings transaction context, client risk appetite, sector norms and a sense of how a clause will behave in a negotiation eighteen months out. No current system approaches that. But the comparison holds document count at one, and live diligence never does. It holds reviewer hours fixed and lets document count rise — and on that comparison the answer changes completely.

Where the lawyer wins decisively

Where the lawyer loses, and it is not about skill

The comparison that decides real outcomes

Lawyer aloneScreening + lawyer
Quality of judgment on what is readHigherSame, if verification is cheap
Documents examinedWhatever hours allowAll of them
How the read set was chosenFolder order, instinctRank against defined categories
Record of the unread materialNoneScored, with source text
Error modeOmission through non-coverageMisranking — recoverable, leaves a trace

The bottom row matters more than the top one. Manual review fails by never reaching a document, which leaves no trace of anything. Screening fails by reaching it and ranking it wrongly — and that failure is recoverable, because the finding exists, the source text is attached, and a reviewer can disagree with it on the record.

What the measured evidence says about the machine side

Stanford researchers evaluated the retrieval-grounded legal AI research tools sold by LexisNexis and Thomson Reuters and found they “each hallucinate between 17% and 33% of the time,” concluding that providers' hallucination-free claims are overstated.[2]

Scope this in both directions. It measures open-ended legal research against case law — a harder retrieval problem than locating a clause in a document you supplied — so it overstates the difficulty of the review task. But it is the best independent evidence available, and it establishes that the error rate in this category is real, measurable, and unsolved by the best-funded attempts. No comparable independent measurement of M&A document review appears to exist.

Read against the question: a system with that error profile cannot replace a lawyer's judgment. It can, reliably, tell you which forty of twelve hundred documents deserve one — which is a different job that nobody was doing well.

The framing that survives contact with a bad outcome

Do not tell a client or committee that AI reviews contracts as well as a lawyer. The first time a finding is wrong, that claim collapses and takes the rest of your credibility with it.

Say instead: every document was examined and scored against every defined category, so our reviewers spent their hours on the highest-ranked material rather than on whatever was in the first folder — and the documents we did not read are on record as ranked below a threshold we set in advance.

That is a smaller claim. It is also true, verifiable, and it holds up when something is missed — because it never promised detection, only coverage and ranking, and the record evidences both.

The variable that decides whether the pairing works

One thing determines whether lawyer-plus-screening beats lawyer-alone: the cost of verifying a single finding.

If a finding quotes its source sentence, verification takes seconds and the arrangement works. If it links to a 90-page agreement without quoting, verification means re-reading — and you have added a step without removing one, making the pairing strictly worse than manual review.

This is why output format matters more than model quality when selecting a tool. The model is the layer you cannot inspect or control. The output shape determines your lawyer's actual workload, permanently.

Bottom line

On one contract the lawyer wins and will keep winning, because the things they bring are not in the document. On twelve hundred contracts the question is which ones the lawyer reads — and answering that with folder order is the status quo, not a standard worth defending.

So the honest answer is that AI does not review contracts as well as a lawyer, and does not need to. It needs to read everything well enough to rank it, quote its source so the lawyer can check in seconds, and leave a record of what it ranked low. The judgment stays where it was.

Sources

  1. ABA Standing Committee on Ethics and Professional Responsibility, Formal Opinion 512: Generative Artificial Intelligence Tools, 29 July 2024. americanbar.org
  2. Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., & Ho, D. E. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. arXiv:2405.20362; Journal of Empirical Legal Studies (2025). arxiv.org/abs/2405.20362

Nothing here is legal advice. We cite only sources we have retrieved and read — see our methodology.

See how fast a quote-backed finding checks →

Anweshna Portal
Anweshna Demo