Put two vendors' feature lists side by side and they are nearly identical — both extract clauses, flag risks, support the major formats, integrate with data rooms. The features have converged. What has not converged is a set of properties that never appear on a feature list, because they describe how a tool behaves when it is wrong, and that is the axis on which these products actually differ.
Why feature parity is real
The capabilities that differentiated tools three years ago are now table stakes, for a straightforward reason: most vendors build on the same handful of foundation models. The model layer is largely commoditised. What a vendor adds on top — the rubric, the output format, the coverage accounting, the record — is where the actual product lives, and almost none of it shows up as a bullet point.
This has a practical consequence for evaluation. Time spent comparing feature matrices is time spent comparing the commoditised layer. The differentiators are all in the second column below.
| On the feature list | What actually differs |
|---|---|
| “Extracts key provisions” | Whether the extraction quotes its source sentence |
| “Identifies risks” | Whether risks are ranked, and against whose rubric |
| “Processes any document type” | What happens to the ones it cannot process |
| “Comprehensive analysis” | Whether you can tell a clean result from an unexamined one |
| “Enterprise security” | Whether the training prohibition is contractual and flows down |
| “Continuously improving” | Whether last quarter's output is still reproducible |
The five properties that separate tools
1. Cost of verifying a single finding
The highest-leverage difference, and the one buyers most often fail to measure.
Grounded commercial legal AI has been measured hallucinating between 17% and 33% of the time in an adjacent task, with providers' hallucination-free claims judged overstated.[1] Nobody has eliminated the error rate, so verification is permanent. Its per-finding cost therefore multiplies across every finding, every deal, forever.
A finding that carries its source sentence verifies in seconds. A finding that links to a 90-page agreement verifies by re-reading it. Two tools with identical recall can differ by an order of magnitude in total cost of ownership on this one property alone — and it never appears in a pitch.
2. Whether silence is distinguishable from coverage
Ask: if a category returns no findings, how do I know it was assessed?
Most reports render “examined and clean” identically to “never examined” — as blank space or a comfortable zero. Tools that assign blocking categories an explicit floor when nothing is found, and report coverage as a first-class output, are in a different reliability class. This is the failure mode that reaches closing, precisely because there is nothing on the page to check.
3. Behaviour at the edges
Every system has limits. The difference is whether it tells you when it hits one.
- Document past the length ceiling — explicit error, documented routing, or silent truncation?
- Scanned document with no text layer — reported as a failure, or absent from the output?
- Unsupported format — flagged, or quietly dropped from the count?
Silent degradation is the single most common source of real misses, and it is invisible in a demo because demos use clean documents.
4. Reproducibility and versioning
Run the same document twice, live, during evaluation. If findings move materially, you cannot reconstruct a decision — and a decision you cannot reconstruct is one you cannot defend.
Related: “we're always improving the model” is offered as a virtue and means the system that produced last quarter's report no longer exists. Versioned models with retained run records are a genuine differentiator and cost the vendor almost nothing, which is why their absence tells you about engineering priorities.
5. Whose rubric the severity reflects
A score is only meaningful relative to a definition of serious. If severity is a fixed global property of a clause type, the tool cannot represent that a provision is routine in your sector and material in another — or vice versa.
Ask whether categories, weights and thresholds are configurable, and whether your reviewers' overrides feed back into calibration. A tool that improves through your use is worth substantially more over three years than one that improves only when the vendor ships.
Two things that matter less than buyers think
Which model is underneath. It is the layer you cannot inspect, control or audit, and it is largely shared across the category. Evaluation time spent here is close to wasted.
Processing speed. Your reviewers were never the bottleneck at reading speed — they were the bottleneck at deciding what to read, and now at verifying what comes back. A tool that processes ten times faster delivers results into a queue that moves at exactly the same pace.
The test that separates them in an afternoon
Take 40–60 documents from a closed deal where your team already knows every material issue. Then:
- Reconcile counts first. Submitted versus processed versus failed. A tool that silently drops documents fails here and the rest of the evaluation is moot.
- Score recall against your known issues — not against total findings returned.
- Time the verification of ten findings with a stopwatch. This number will surprise you and it is the one that compounds.
- Read the quotes against the claims, hunting misgrounded citations — a real clause cited for something it does not say.
- Run one document twice.
- Submit the credit agreement and check whether findings from the final third appear at all.
The question that outranks the test
Every real product has this story. A vendor who cannot produce one has no live deployments, is not being straight, or has no feedback path by which errors return to them — and the third is the most concerning, because it means their error rate is unmeasured rather than low.
What you want is specific and slightly uncomfortable: the class of error, how it surfaced, what changed. What you do not want is a pivot to a case study.
Bottom line
Features have converged; behaviour under error has not. The better tool is the one whose mistakes are cheap to find, whose gaps are visible rather than silent, whose edges fail loudly, and whose output you can still reconstruct in eighteen months.
None of that is on a feature list, all of it is testable in an afternoon, and it is what you will actually live with.
Sources
- Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., & Ho, D. E. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. arXiv:2405.20362; Journal of Empirical Legal Studies (2025). Measures legal research, not document review — see the scoping note above. arxiv.org/abs/2405.20362
We name no competitors and publish no comparison scores, because no independent benchmark exists and we will not invent one. We cite only sources we have retrieved and read — see our methodology.