Screening thousands of companies and screening one company's data room are opposite problems, and the instinct that works for the first is dangerous in the second. At the top of the funnel you optimise for throughput and tolerate error, because a wrong candidate costs a phone call. At the bottom you optimise for coverage and verifiability, because a wrong conclusion costs the deal.
The funnel, and where the economics invert
| Stage | Volume | Data | Cost of a false positive | Cost of a false negative |
|---|---|---|---|---|
| Universe filter | Thousands | Databases, filings | A wasted screen | A missed opportunity — tolerable |
| Shortlist | Dozens | Public filings, news | An hour of analyst time | Tolerable |
| Engaged target | One | Disclosed documents | A consultant's hour | The deal |
Read the last two columns down the table. The asymmetry flips completely between the first row and the last, and with it the correct posture toward automation.
At the top, aggressive filtering is rational — you are trying to reduce a universe, and missing one candidate among thousands is survivable. At the bottom, the same tolerance means a provision reaching closing unexamined. Teams that carry top-of-funnel habits downstream — light verification, comfort with noise, trust in an untraceable score — are the ones that get hurt.
Top of funnel: what actually constrains it
Not model quality. Data quality. Filtering a universe of companies runs on registries, filings, market databases and news, and the output is bounded by the coverage, freshness and accuracy of those sources.
Two consequences worth planning around:
- Private companies are thinly covered. If your thesis targets founder-owned businesses below public-reporting thresholds, the data simply is not there in most markets, and no amount of processing creates it.
- Staleness is invisible. A database record from eighteen months ago looks identical to one from last week. Ask any sourcing vendor how recency is represented, and whether it is exposed in the output.
We build for the bottom of the funnel, not the top, so treat this section as orientation rather than expertise.
Bottom of funnel: where advisers actually lose money
Once a target engages and a data room opens, the constraint becomes arithmetic. An adviser running several processes at once cannot read a 1,200-document room in full. They read what hours allow — and without a screening stage, that subset is chosen by folder structure, upload order, and whoever flagged something on a call.
For an adviser this matters twice over: once for the client's outcome, and once for your own position if something surfaces post-close and the question becomes what your process covered.
Running several mandates at once
The operational reality that distinguishes advisers from corporate acquirers is concurrency — three or four live processes with overlapping deadlines. Three implications:
- Per-engagement isolation is not optional. A screening system that indexes across mandates can quietly become the fastest way to breach an information barrier. Require per-room isolation, server-side role enforcement, and read-level access logging — most systems log writes and not reads, which is backwards for this work.
- Verification cost multiplies by mandate count. If confirming a finding takes four minutes rather than seconds, that cost lands three times over in the same week. This single property decides whether screening is net-positive for a multi-mandate team.
- Calibration is harder than for a repeat acquirer. A generalist adviser sees different sectors each quarter, so overrides accumulate more slowly against any one rubric. Capture them by sector rather than in aggregate, or the signal averages out to nothing.
The two numbers worth instrumenting
Advisers are unusually well placed to measure this, because you run many processes and can compare across them.
Documents in the room versus documents anyone opened. Calculate it for your last three mandates. Most firms have never done this. It is usually uncomfortable, entirely defensible because it is your own data, and it quantifies the risk you are already carrying — which is a stronger basis for a decision than any vendor claim.
Verification minutes per escalated finding. Multiply by escalation volume and mandate count. This is your real bottleneck, and it will not appear in any pitch.
What does not change
- Materiality judgment. Whether a flagged provision changes the price is a judgment about a specific transaction and a specific client's appetite.
- Anything undisclosed. A clean result on an incomplete room is a confident answer to the wrong question.
- Verification. Grounded commercial legal AI has been measured hallucinating between 17% and 33% of the time in an adjacent task, with providers' hallucination-free claims judged overstated.[1] Checking escalations is permanent.
- Accountability. It stays with the adviser, and no vendor liability cap is sized for a deal-scale loss.
Bottom line
Keep the funnel's two ends apart. Top-of-funnel screening runs on purchased data with high error tolerance, and its ceiling is data coverage rather than model capability. Bottom-of-funnel assessment runs on disclosed documents where a false negative is the expensive error.
For a multi-mandate adviser the decisive properties are per-engagement isolation, cheap verification, and coverage reporting you can point at later. Instrument your examined-versus-read gap across three mandates first — that number will make the decision for you, and it is yours rather than a vendor's.
Sources
- Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., & Ho, D. E. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. arXiv:2405.20362; Journal of Empirical Legal Studies (2025). See the scoping note above. arxiv.org/abs/2405.20362
We build for document assessment, not top-of-funnel sourcing, and describe the latter only in general terms. We publish no adoption or accuracy statistics — see our methodology.