Comparing solutions on licence fee gets the ranking wrong, reliably. The dominant cost in this category is not what you pay the vendor — it is the labour their output format imposes on your team, forever. Two tools with identical prices can differ several-fold in total cost, and the difference never appears in a quote.
The four solution shapes
| Shape | Licence cost | Labour it imposes | Best when |
|---|---|---|---|
| General-purpose LLM | Lowest | Highest — no coverage record, no ranking, manual everything | Ad-hoc questions on one document |
| Clause extraction | Moderate | Moderate — output unranked, triage still manual | Building a structured record |
| Risk screening | Moderate to high | Lowest, if findings quote source | Triage across a large unread set |
| Full platform / managed | Highest | Low, plus vendor dependency | Large teams with procurement capacity |
Read the first row carefully, because it is the most common false economy. A general-purpose model appears nearly free and imposes the highest ongoing cost: no coverage accounting, no ranking, no persistent record, and verification by re-reading. It is genuinely useful for a targeted question on a document you already know matters. As a diligence layer it transfers the entire workload to people.
The five cost lines, only one of which is quoted
- Licence. The only one in the proposal. Note the model matters more than the number — per-seat pricing taxes exactly the reviewers who generate value, so the tool ends up used by two people and judged on that basis.
- Volume overage. Deal flow is lumpy. Ask the per-unit rate above the included volume and whether unused capacity carries forward.
- Verification labour. Permanent, and usually the largest line. See below.
- Configuration and calibration. Mostly one-off, borne by your most expensive people.
- Security and procurement review. For a first AI vendor, often the longest part of the process.
Why verification labour dominates
Grounded commercial legal AI has been measured hallucinating between 17% and 33% of the time in an adjacent task, with providers' hallucination-free claims judged overstated.[1] No vendor has eliminated the error rate, so checking escalated findings against source text is a standing cost, not a transitional one.
Its magnitude is set almost entirely by output format:
- Finding quotes its source sentence → verification is reading two things side by side. Seconds.
- Finding links to a document → verification means locating and reading the relevant passage. Minutes.
Multiply by escalations per deal and deals per year. A tool that is cheaper by a few thousand and costs three extra minutes per finding is not cheaper — and the gap widens every year you own it.
A worked comparison — replace every input
Assume 12 deals a year, 40 escalated findings per deal, and a reviewer at $300/hour.
- Tool A — quotes source text. Verification: 20 seconds per finding.
40 × 12 × 20s = 2.7 hours/year × $300 ≈ $800/year - Tool B — links to documents only. Verification: 3 minutes per finding.
40 × 12 × 3min = 24 hours/year × $300 = $7,200/year
A $6,400 annual difference in labour alone, before considering that Tool B's verification is far more likely to be skipped under deadline — which converts a cost line into a risk.
Change the escalation count and the gap moves proportionally. At 100 escalations per deal it exceeds $16,000 a year, which is more than many licence fees in this category.
The costs of getting it wrong
Three failure modes that turn a cheap solution expensive, none of them on a spreadsheet:
No coverage reporting. You cannot distinguish an examined-and-clean document from one that failed extraction. Either you re-check manually — expensive — or you carry an unmeasured gap, which is the actual cost.
Silent truncation. The credit agreement is the longest document in the room and the one that matters most.
Non-reproducible output. If the same file scores differently across runs, or the model was replaced without versioning, you cannot reconstruct a decision. The cost lands as indefensibility rather than as hours.
What to compare instead of price
- Seconds to verify one finding. Time it with a stopwatch during the pilot. This is the single most predictive number.
- Whether coverage is reported — submitted, processed, failed, truncated.
- Whether a category returning nothing is distinguishable from one never assessed.
- Whether output is reproducible and models are versioned.
- Whether severity is configurable to your sector, and whether overrides feed calibration.
- Then the licence fee and its structure.
On benchmarking against the market
You cannot. Essentially no vendor in this category publishes prices, so there is no reference rate — and any “market rate” someone cites is anecdote. Neither is there an independent published accuracy benchmark to compare against.
Which means the comparison has to be constructed from your own numbers: your escalation volume, your reviewer rate, your deal count, timed on your own documents. That is more work than reading a comparison table, and it is the only version that predicts anything.
Bottom line
The cheapest licence is frequently the most expensive solution, because the dominant cost is the verification labour the output format imposes — permanently, and growing with your deal volume.
So compare on seconds-to-verify, coverage reporting, and reproducibility first. Then look at the licence. A tool that costs more and makes every finding checkable in twenty seconds will usually be cheaper by year two, and safer throughout.
Sources
- Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., & Ho, D. E. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. arXiv:2405.20362; Journal of Empirical Legal Studies (2025). See the scoping note above. arxiv.org/abs/2405.20362
The worked comparison is an illustration with stated assumptions, not a measurement, and Tool A and Tool B are not real products. We publish our own prices and no competitor's, because they do not publish them — see our methodology.