The legal risk of using AI in diligence is not that a model invents something. It is that the professional duty attached to the work is non-delegable, no vendor accepts it, and no liability cap in this market is sized for deal-scale loss. Everything below follows from that, and none of it is changed by choosing a better tool.
Risk 1 — The duty does not move
In July 2024 the ABA Standing Committee on Ethics and Professional Responsibility issued Formal Opinion 512, its first formal guidance on generative AI. It holds that lawyers and firms using these tools must “fully consider their applicable ethical obligations,” naming competence, confidentiality, communication with clients, candor toward tribunals, supervisory responsibility, and charging reasonable fees consistent with time actually spent.[1]
Two consequences are worth stating bluntly:
- Vendor selection does not discharge competence. The duty attaches to the work product that leaves your desk, not to the procurement decision behind it.
- Supervision applies to output you did not personally generate. This is the same principle that governs work delegated to a junior — except the junior can explain their reasoning and the model cannot.
Outside regulated practice the same structure holds commercially. A principal answering to an investment committee, or a manager answering to limited partners, carries accountability no licence absorbs.
Risk 2 — Reliance on unverified output
Grounded commercial legal AI has been measured hallucinating between 17% and 33% of the time in an adjacent task, and the researchers concluded that providers' hallucination-free claims are overstated.[2]
The legally dangerous failure is not the invented clause — that one is caught easily. It is the misgrounded citation: a real provision cited for a proposition it does not support. It survives a click-the-link check and fails only a read-the-clause check, which means it survives exactly the level of review most teams actually perform under deadline.
Risk 3 — Confidentiality and privilege
Deal documents are covered by NDAs, by professional confidentiality duties, and often by privilege. Sending them to a third-party system raises three distinct exposures:
- Training use. “Not by default” is a setting, not a term. You want a contractual prohibition flowing down to subprocessors and the underlying model provider.
- Undisclosed subprocessors. The model provider, cloud host, logging platform and OCR service are all parties your NDA never contemplated.
- Privilege. Whether disclosure to a vendor waives privilege is jurisdiction-specific and not settled uniformly. Treat it as a question for your own counsel, and note that a vendor's reassurance on this point is not advice you can rely on.
Risk 4 — The record you cannot reconstruct
When a miss is investigated — by an insurer, a claimant, or your own committee — the question resolves to one of two findings:
Not defensible: the document was never processed, or the category never covered, and nothing in the record shows anyone knew.
The first is an ordinary professional judgment that turned out wrong — the kind of risk processes and insurance are built to absorb. The second is an unmanaged gap and very hard to characterise otherwise.
What makes the difference is not accuracy. It is whether the system produces a reconstructible record: source text, category, score, threshold, model version, timestamp, and every human override. A vendor who updates models continuously without versioning has removed your ability to answer the question at all.
Risk 5 — Fees and disclosure to clients
Opinion 512 addresses fees directly: charges must be reasonable and consistent with time actually spent.[1] If a tool compresses twenty hours of triage to two, billing twenty is not defensible.
The opinion also addresses when use must be disclosed to clients. This is fact-specific rather than a blanket rule, and it is worth resolving with your own risk function before a client asks rather than after.
Risk 6 — The regulatory surface, which now has teeth
The governance landscape hardened quickly and on a documented timeline:
| Instrument | Date | Relevance |
|---|---|---|
| NIST AI Risk Management Framework 1.0[3] | 26 Jan 2023 | Voluntary; a workable structure for internal assessment (Govern, Map, Measure, Manage) |
| ISO/IEC 42001 | 2023 | Certifiable AI management system standard — worth asking vendors about |
| ABA Formal Opinion 512[1] | 29 Jul 2024 | Directly binding on professional conduct |
| EU AI Act, Art. 99[4] | Applies 2 Aug 2025 | Fines to €35,000,000 or 7% of worldwide annual turnover for prohibited practices; €15,000,000 or 3% for provider/deployer obligation breaches |
Whether the AI Act reaches your specific use depends on classification and on where you operate, which is a question for counsel. The general point stands regardless: the expectation that you can show how a conclusion was reached is now written into instruments with penalties attached.
Risk 7 — Insurance
Check your professional indemnity policy for an AI exclusion or a notification condition, before a deal rather than after a claim. This is a short call to your broker and it occasionally produces a very unwelcome answer. Firms discover this at the worst possible moment with some regularity.
What actually mitigates
- Verify everything you escalate against source text — the quote read against the claim, not merely the link followed.
- Set escalation thresholds in writing before the deal. Pre-committed, a miss below the line is a documented risk-appetite decision. Set afterwards, it is a rationalisation.
- Make absence of a finding visibly different from absence of a search. Blocking categories floored rather than reporting a comfortable zero.
- Retain the run record — document set, categories, scores, thresholds, quotes, model version, overrides.
- Get the training prohibition and subprocessor list in the contract, not on a policy page.
- Push on warranty carve-outs, not on the liability cap. A cap at fees paid is normal; excluding AI output from warranties that otherwise cover the product tells you where the vendor thinks the risk sits.
- Keep materiality with a named human, and record when they disagree with the system. That override log is the evidence of professional judgment.
The risk of not using it
Worth stating, because this page is otherwise one-sided. If your team reads 400 of 1,200 documents and the other 800 were selected by folder order, you are already carrying unexamined risk with no record that anyone considered it. That is not a safe default; it is an undocumented one.
The genuine choice is not between risk and no risk. It is between an unexamined portion of the data room and an examined-but-ranked one — and only the second produces the record that makes a miss defensible.
Bottom line
You keep the duty, you keep the liability, and the caps on offer will not cover a deal-scale loss. So optimise for the position you want to be in when something is missed: document examined, category scored, provision ranked below a threshold you set in advance and wrote down, named human on the call, record retained.
Get there and a miss is a judgment. Fail to get there and it is a gap nobody was managing.
Sources
- ABA Standing Committee on Ethics and Professional Responsibility, Formal Opinion 512: Generative Artificial Intelligence Tools, 29 July 2024. americanbar.org
- Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., & Ho, D. E. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. arXiv:2405.20362; Journal of Empirical Legal Studies (2025). arxiv.org/abs/2405.20362
- NIST AI Risk Management Framework (AI RMF 1.0), 26 January 2023. nist.gov
- Regulation (EU) 2024/1689 (EU AI Act), Article 99; applies from 2 August 2025. artificialintelligenceact.eu/article/99
Nothing here is legal advice; consult your own counsel and regulator. We cite only sources we have retrieved and read — see our methodology.