Tools in this space look interchangeable in a demo and behave very differently in a data room, because they were built to answer different questions. Feature lists hide this — everything “extracts clauses” and “identifies risk.” What separates them is architectural, and once you can name the four families, most of the confusion in a vendor comparison disappears.
The four families
| Family | Built to answer | Strong at | Weak at |
|---|---|---|---|
| Playbook comparison | Does this deviate from our standard position? | Repeat contracting, your own paper | Buy-side diligence on someone else's paper |
| Clause extraction | Where is provision X across these documents? | Building a structured record at volume | Telling you which provisions matter |
| Conversational Q&A | What does this document say about Y? | Ad-hoc exploration of a known document | Coverage — it only answers what you ask |
| Risk screening | Which documents carry the most risk? | Triage across a large unread set | Deep single-document analysis |
These are not marketing segments; they are different jobs, and a tool optimised for one is genuinely poor at another. Most disappointment in this category comes from buying a tool from the wrong family and concluding the technology does not work.
Playbook comparison — the CLM heritage
Descended from contract lifecycle management. You define standard positions — 90-day termination, mutual indemnity, no MFN — and the system flags deviations in incoming paper.
Excellent for a recurring counterparty negotiation where you control the template. The structural mismatch with M&A diligence is that you are reading a target's contracts with a hundred different counterparties, none of which were drafted against your playbook. Everything deviates, so everything flags, and the output degrades into noise.
Test for it: ask what happens with no playbook configured. If the answer is that it falls back to general market norms, ask what corpus defines “market” and whether it covers your sector.
Clause extraction — the eDiscovery heritage
Locates and pulls specified provision types across a set: every change-of-control clause, every assignment restriction, every governing-law designation. Output is typically a structured table.
Genuinely useful, and it composes well with everything else — extraction into a filterable record is a real capability. The limitation is that it answers “where is X?” and not “does X matter?” You get 340 change-of-control provisions and no ordering, which relocates the triage problem rather than solving it.
Test for it: ask whether the output is ranked, and by what. If the answer is that ranking is left to the reviewer, you have bought extraction, not screening — which may be exactly right, as long as you know it.
Conversational Q&A — the chat heritage
Ask questions in natural language against a document or a set. Demos superbly, because a good question produces an impressive answer.
Its weakness is structural and easy to miss: it only tells you about things you thought to ask. In diligence the expensive failures are the provisions nobody knew to ask about, in documents nobody prioritised. A Q&A interface over a data room is a better search box, not a coverage instrument, and it can create real false confidence because the answers it does give are fluent and specific.
Test for it: ask how you would discover an issue you did not anticipate. If the honest answer is that you would have to think of the question, that defines the tool's ceiling.
Risk screening — the triage function
Scores every document against a defined set of risk categories, attaches the source text behind each score, and ranks the set so reviewer attention goes to the top.
The trade-off is real and worth stating: screening is shallower per document than a focused Q&A session or a playbook comparison. It is not trying to tell you everything about one contract. It is trying to tell you which forty of twelve hundred deserve a lawyer's afternoon, and to leave a record of the other eleven hundred and sixty.
This is the right family for buy-side diligence specifically, because the binding constraint there is coverage rather than depth. Where a team can read 400 documents of 1,200, the 800 unread are currently selected by folder order — which correlates with nothing.
Architectural differences that outrank family
Within any family, four properties matter more than the feature list:
Does a finding quote its source sentence?
The single most consequential difference, and it decides your team's real workload. A finding with the sentence attached verifies in seconds. A finding that links to a 90-page agreement verifies by re-reading it. Given that grounded commercial legal AI has been measured hallucinating between 17% and 33% of the time in an adjacent task[1], verification is permanent — so its per-finding cost compounds across every deal you ever run.
Can it tell you what it did not examine?
A document that failed extraction produces zero findings, which looks exactly like a clean document. Tools that report coverage as a first-class output — submitted, processed, failed, truncated — are in a different reliability class from those that only report findings.
Does anything operate across the set?
Customer concentration, aggregate indemnity exposure, and how many employment agreements lack IP assignment are not properties of any single document. Many tools are strictly per-document and it is not obvious from a demo, because demos run one document at a time.
Is output reproducible?
Ask them to run the same file twice, live. If findings move, you cannot reconstruct a decision — and you cannot defend one you cannot reconstruct.
Why there is no comparison table on this page
You will find plenty elsewhere with named vendors, prices and accuracy scores side by side. We checked, and the underlying data does not exist: essentially no vendor in this category publishes prices, and there is no independent published benchmark for accuracy in M&A document review.
A table built from that would be assembled from sales conversations and vendor marketing, presented as research. We would rather explain the architecture and let you test the four properties above on your own documents, which is the only comparison that predicts anything.
Matching the family to the job
- Negotiating your own paper repeatedly → playbook comparison
- Building a structured record of specific provisions → clause extraction
- Exploring a document you already know matters → conversational Q&A
- Deciding what to read in a large unread set → risk screening
- Buy-side M&A diligence → screening first, with extraction alongside; Q&A is a supplement, never the coverage layer
Bottom line
The meaningful differences are architectural, not featural. Work out which question the tool was built to answer, then check whether it is your question — most buyer disappointment is a family mismatch rather than a quality problem.
Then test the four properties that cut across all families: does it quote its source, can it tell you what it missed, does anything work across the set, and does the same document score the same way twice.
Sources
- Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., & Ho, D. E. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. arXiv:2405.20362; Journal of Empirical Legal Studies (2025). Measures legal research, not contract review — see the scoping note above. arxiv.org/abs/2405.20362
We name no competitors and publish no comparison table, because the underlying pricing and accuracy data is not public and we will not invent it. We cite only sources we have retrieved and read — see our methodology.