What’s the Timeline for Implementing AI in Our Deal Process?

Understanding realistic rollout phases and adoption milestones

Updated August 2026 · 6 min read · Deal Room Intelligence Series

Technical setup takes days. That is not the timeline that matters, and quoting it is how implementations get scheduled badly. The real schedule is governed by three things that run on their own clocks and mostly cannot be compressed: your security review, your calibration data, and the arrival of a deal to actually run it on.

The four workstreams, and which one gates

WorkstreamDriven byCompressible?
Technical setupAccount provisioning, access, document routingYes — days
Security and procurementYour infosec team, DPA negotiation, training-prohibition termsNo — usually the gate
CalibrationReal deals producing real overridesNo — needs elapsed deals
Behaviour changePeople trusting the output enough to change what they readPartially

Vendors quote the first row. Your actual go-live is set by the second, and your actual value arrives on the schedule of the third — which is measured in deals, not weeks.

Phase 0 — Security review, before anything else

Start here, in parallel with evaluation rather than after selecting. For a firm handling deal documents this is the longest pole and it is frequently discovered late.

What has to be settled: a contractual training prohibition flowing down to subprocessors and the underlying model provider; a named subprocessor list; stated retention with client-triggered deletion; region pinning where NDAs require it; and a DPA where personal data is in scope.

This is not a formality. ABA Formal Opinion 512 (29 July 2024) names confidentiality among the duties engaged when lawyers use generative AI.[1] Putting client material into a system whose terms permit training is a conduct problem before it is a security one — which means your risk function, not just IT, has to sign.

Phase 1 — Calibrate on a closed deal

Before any live transaction, run 40–60 documents from a deal that has already completed and where your team knows every material issue.

Two purposes, and the second matters more:

Also seed the structural edge cases: a scanned document with no text layer, an over-length credit agreement, a base agreement whose amendment is deliberately absent. These probe the failures that cause real misses and never appear in a demo.

Phase 2 — Run it alongside, not instead

On the first live deal, run screening in parallel with your existing process. Do not let it change what anyone reads yet.

This costs a little duplicated effort and buys the thing you cannot get any other way: a direct comparison on a deal where the human process is still fully intact. You learn whether the ranking would have pointed your reviewers somewhere useful, without betting a transaction on the answer.

Expect the first parallel run to be uncomfortable. Over-flagging is normal at default settings, and the temptation is to conclude the tool does not work. It is more likely that severity is calibrated to a generic rubric rather than your sector — which is what Phase 3 fixes.

Phase 3 — Calibrate against your own overrides

This phase runs on deals, not weeks, and it is the one no schedule can compress.

Each time a reviewer disagrees with a finding, capture what was flagged, what they concluded, and why. After three or four deals you have enough to see patterns: a category routinely downgraded is mis-tuned and its scores are noise; a category routinely upgraded means systematic under-detection.

Adjust weights and thresholds against that data. This is where a screening stage stops producing generic output and starts reflecting how your firm actually thinks about risk — and it is the difference between a tool people use and a tool people ignore.

Phase 4 — Let it change what gets read

Only now does the operating model actually change: reviewers work down a ranked list rather than through folder order, and the documents below the threshold are on record as examined and ranked rather than unopened.

Two disciplines to establish before this becomes routine, because they are what keep it defensible:

Scoping that figure. It measures open-ended legal research against case law, not review of a document you supplied. Use it to set the expectation that errors exist and must be sampled for, not as a target rate.

What the schedule actually looks like

Stated as dependencies rather than dates, because dates depend on your security function and your deal flow:

  1. Security and procurement — start immediately, in parallel with evaluation. Typically the gate.
  2. Calibration on a closed deal — a few days of one person's time, once terms are signed.
  3. First parallel live deal — whenever a deal arrives.
  4. Three to four deals of override capture — the real clock. Cannot be shortened by adding people.
  5. Operating-model change — after calibration, not before.

The honest summary: you can be technically live in days and genuinely calibrated after three or four deals. A firm doing two transactions a quarter should plan for the second half of the year, not the second week.

Three ways the timeline slips

Security review starts after vendor selection. The most common cause of a multi-month delay, and entirely avoidable by running it in parallel.

The operating model changes before calibration. Reviewers are told to trust the ranking while it is still generic, they see obviously wrong severity, and confidence is lost in a way that is hard to recover. Sequence matters more than speed here.

Nobody owns it. Implementations without a named owner responsible for thresholds, override review and the sampling protocol tend to stall after the first deal — the tool remains available and nobody's job is to make it work.

Bottom line

Technical setup is days and irrelevant. Start the security review immediately, calibrate on a closed deal before touching a live one, run parallel for at least one transaction, and expect real calibration to take three or four deals because it runs on override data you cannot manufacture.

Change what people read last, not first. The sequence is what determines whether the thing gets used a year from now.

Sources

  1. ABA Standing Committee on Ethics and Professional Responsibility, Formal Opinion 512: Generative Artificial Intelligence Tools, 29 July 2024. americanbar.org
  2. Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., & Ho, D. E. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. arXiv:2405.20362; Journal of Empirical Legal Studies (2025). Measures legal research, not document review. arxiv.org/abs/2405.20362

We publish no implementation-duration statistic because it depends on your security function and deal flow, not on the software. We cite only sources we have retrieved and read — see our methodology.

Start with a document you already know →

Anweshna Portal
Anweshna Demo