Should We Replace Lawyers with AI for Document Review?

Understanding complementary versus replacement use cases in 2026

Updated August 2026 · 6 min read · Deal Room Intelligence Series

No — and the reason is more specific than “AI isn't good enough yet.” The professional duty attached to reviewing a document is non-delegable. You cannot transfer it to a vendor, and no contract in this market attempts to accept it. So the question is not whether a system could do the reading. It is who is accountable for the result, and that answer does not change with model capability.

The duty is the constraint, not the capability

In July 2024 the ABA Standing Committee on Ethics and Professional Responsibility issued Formal Opinion 512, its first formal guidance on generative AI. It holds that lawyers and firms using these tools must “fully consider their applicable ethical obligations,” naming duties of competence, confidentiality, communication with clients, candor toward tribunals, supervisory responsibility, and charging reasonable fees consistent with time actually spent.[1]

Read that against the replacement question and it resolves immediately. Competence and supervision attach to the output that leaves your desk. Selecting a capable vendor does not discharge them; it only changes how the work gets done underneath. There is no configuration of tooling in which the duty moves.

The same logic operates outside regulated practice. A deal principal answering to an investment committee, or a fund manager answering to limited partners, holds an accountability that no software licence absorbs. The liability caps available — fees paid, typically — are not sized for deal-scale loss even where a vendor accepts them.

The practical formulation: you can delegate the reading. You cannot delegate the responsibility for having read. Every sensible operating model in this category is a consequence of that sentence.

What actually happens to the work

Document review is not one activity, and the replacement question only looks interesting because the phrase bundles several jobs together. Separate them and the picture is unambiguous.

TaskWhat it involvesWhere it should sit
TriageDeciding which documents deserve attentionMove it. Expensive, slow, produces only an ordering
ExtractionLocating provisions and termsMove it, with source quotes attached
Consistency checkingSame clause type across 900 documentsMove it — humans degrade here, systems do not
Materiality judgmentDoes this matter for this deal?Keep. Contextual, accountable
Inference across silenceNoticing the indemnity that should be there and is notKeep. Largely beyond current systems
AdviceWhat to do about itKeep. This is the professional act

The rows that move share a property: they are allocation and mechanical work, not judgment. The rows that stay are the ones a client is actually paying for. Framed this way, the interesting observation is that triage — currently performed by the most expensive people on the deal, producing nothing but an ordering — has never been a good use of a lawyer, and was only ever done by one because there was no alternative.

Why the reading itself is not the bottleneck

The replacement framing assumes reading speed is the constraint. In large data rooms it is not. The constraint is that nobody reads most of the room at all.

A team with capacity for 30% of a 1,200-document set reads a 30% selected by folder structure, upload order, and whoever flagged something on a call — none of which correlate with risk. The other 70% is not judged unimportant. It is never seen, and nothing in the record indicates anyone considered it.

A screening stage does not make anyone read faster. It changes which 30% gets the expert hours, and it leaves a scored record of the other 70%. That is a coverage change, not a replacement — and it is where essentially all of the value sits.

Two things that go wrong when firms try replacement anyway

Verification cost is discovered late

Grounded commercial legal AI has been measured hallucinating between 17% and 33% of the time in an adjacent task, with providers' hallucination-free claims found to be overstated.[2] That means checking escalated findings against source text is permanent, not transitional.

Firms that plan for replacement do not budget this, then discover that the reviewer they removed was the person who would have caught it. Firms that plan for reallocation budget it from the start and buy tools whose output makes it cheap — the source sentence beside the claim, so verification takes seconds rather than a re-read.

Scoping that figure honestly. It measures open-ended legal research against case law, not review of a document you have supplied — a meaningfully harder retrieval problem. Treat it as evidence that the category has a real unsolved error rate, not as a measurement of contract review accuracy.

The record disappears

A replacement model tends to produce a report and no reasoning. A reallocation model produces scores, thresholds, source quotes, and — critically — recorded human overrides. That override log is the evidence that judgment was applied, and it is what makes reliance defensible when a miss is later investigated. It is also the difference between “we ranked this below our documented escalation line” and “the system didn't flag it,” which are very different sentences in a dispute.

What a defensible operating model looks like

  1. Screen the full set against defined categories, with blocking categories floored so that absence of a finding is visibly different from absence of a search.
  2. Verify escalations against source text. Cheap by design, or the model fails.
  3. Apply expert judgment to what survives — on a set small enough to deserve real attention.
  4. Sample the tail to measure the actual error rate instead of assuming it.
  5. Record every override, because that is the audit trail of professional judgment.

Note that steps 3 and 5 are lawyers doing lawyer work, and steps 1 and 4 are the parts nobody should have been doing by hand. Nothing was replaced. The allocation changed.

The fee question, which is real

Opinion 512 addresses fees directly: charges must be reasonable and consistent with time actually spent.[1] If a tool compresses triage from twenty hours to two, billing twenty is not defensible.

This is genuinely uncomfortable for hourly models, and it deserves a straight answer rather than avoidance. The work that compresses is the low-judgment work, which is also the low-rate work; what remains is disproportionately the analysis clients actually value. Firms that handle this well tend to reprice around the judgment rather than defending hours on the triage — and they are better positioned when a client eventually asks what the tool changed.

Bottom line

The replacement question has a fixed answer because it is a question about accountability, not capability. The duty is non-delegable, no vendor accepts it, and no liability cap is sized for the loss.

So move the triage, the extraction and the consistency checking — the work that was never a good use of professional time. Keep the materiality calls, the inference across silence, and the advice. Budget the verification permanently. And record the overrides, because in the end that log is the thing that demonstrates a professional was actually in the loop.

Sources

  1. American Bar Association Standing Committee on Ethics and Professional Responsibility, Formal Opinion 512: Generative Artificial Intelligence Tools, 29 July 2024. americanbar.org
  2. Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., & Ho, D. E. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. arXiv:2405.20362; Journal of Empirical Legal Studies (2025). arxiv.org/abs/2405.20362

Nothing here is legal or ethics advice; consult your own regulator and counsel. We cite only sources we have retrieved and read — see our methodology.

See what moves and what stays →

Anweshna Portal
Anweshna Demo