Operating the software takes an hour. That is not the learning curve, and quoting it is how rollouts get planned badly. The real curve is learning what to trust — and that runs on deals, not on training sessions, because it is calibration rather than instruction.
Three curves, on three different clocks
| Curve | What is learned | Measured in |
|---|---|---|
| Mechanical | Upload, run, read output, export | Hours |
| Interpretive | What a score means, when to trust a finding | One or two deals |
| Calibration | What this firm considers serious | Three to four deals |
Vendors quote the first. Your team's confidence is set by the second. The tool's actual usefulness is set by the third — and none of the three can be compressed by more training, because the last two require live deals to generate the data.
The interpretive curve: what reviewers are really learning
A reviewer looking at a score of 82 in a category does not initially know what it means. Not because the number is unclear, but because they have no basis for knowing whether 82 in this system corresponds to something they would have flagged.
What builds that judgment is repetition against known answers. Which is why the first exercise should be a closed deal where the team already knows every material issue — the reviewer sees the score, sees what was actually there, and starts forming a mapping. On a live deal there is no answer key, so nothing calibrates.
Expect two specific reactions in this phase, and pre-empt both:
- “It flagged something obviously routine.” Default severity is tuned to a generic commercial baseline, not your sector. This is expected and it is what the calibration phase fixes.
- “It missed something I would have caught.” Sometimes real, sometimes a threshold question — the finding may exist, ranked below the escalation line. Check before concluding, because the two have completely different remedies.
The skill that actually needs teaching
Not operating the tool. Verification. And specifically, one non-obvious thing: reading the quoted source sentence against the claim, rather than confirming the reference exists.
The failure mode this catches is the misgrounded citation — a real clause cited for a proposition it does not support. It passes an existence check and fails only a reading check, which means it survives exactly the level of review most people perform by default.
This matters because errors are permanent. Grounded commercial legal AI has been measured hallucinating between 17% and 33% of the time in an adjacent task, with providers' hallucination-free claims judged overstated.[1]
Ten minutes of instruction, and it is the only genuinely new professional skill in the rollout.
By role
- Analysts — fastest on the mechanical curve, slowest on interpretive, because they have less precedent to compare a score against. Pair them with a reviewer for the first deal.
- Senior reviewers — often slower mechanically, much faster interpretively; they can tell immediately whether a ranking is sensible. They are also the ones whose overrides carry the most calibration value, so their engagement matters disproportionately.
- Deal leads — mainly need to read a coverage report and understand what a threshold commits them to. Twenty minutes.
- Risk and compliance — need the audit surface: what is recorded, how a decision is reconstructed, what the retention terms are.
What makes the curve steeper than it needs to be
Starting on a live deal. No answer key, high pressure, and any early error becomes evidence against the tool.
Changing the operating model before calibration. Telling reviewers to trust a ranking that is still generic produces exactly the loss of confidence that is hardest to recover.
Output that is expensive to verify. If checking a finding costs four minutes rather than seconds, verification lapses in week three of a deal — and a team that has stopped verifying is not learning what to trust, it is guessing.
No override capture. Without it there is no calibration data, so the third curve never completes and the tool stays generic indefinitely.
A rollout that respects the curves
- One hour — mechanical training for whoever runs scans.
- Ten minutes — verification technique: read the quote against the claim.
- Half a day — closed-deal exercise where the team knows the answers. Set thresholds here, in writing, before any deal depends on them.
- One live deal in parallel — screening runs alongside the existing process without changing what anyone reads.
- Three to four deals of override capture — the real clock, and it cannot be shortened by adding people.
- Then change what gets read.
What never becomes automatic
Two things stay deliberate no matter how experienced the team becomes:
- Coverage reconciliation. Submitted versus processed versus failed, before reading findings, every deal. Two minutes, and it never stops being necessary — a document that failed extraction produces zero findings, which looks exactly like a clean one.
- Sampling the low-severity tail. A fixed random sample every deal. Familiarity actively works against this: the more comfortable a team gets, the less they check the boring middle, which is exactly where errors accumulate.
Bottom line
An hour to operate it, a deal or two to interpret it, three or four deals to calibrate it. The mechanical curve is trivial and is the only one anyone quotes.
Start on a closed deal where the answers are known, teach verification as reading the quote against the claim, capture every override, and do not change what people read until calibration has actually happened. Rushing the sequence is what produces a licence that renews while nobody uses it.
Sources
- Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., & Ho, D. E. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. arXiv:2405.20362; Journal of Empirical Legal Studies (2025). See the scoping note above. arxiv.org/abs/2405.20362
We publish no time-to-proficiency statistic because it depends on your deal cadence, not on the software. We cite only sources we have retrieved and read — see our methodology.