A generic screening rubric applied to a specialist target produces confident, well-formatted, wrong answers. The clauses get read correctly. The severity is assigned against a general commercial baseline that does not describe the sector — so genuine blockers rank low and routine provisions rank high, with nothing in the output to indicate which is which.
What actually differs by sector
Three things, and they need different fixes:
| Difference | Example | Fix |
|---|---|---|
| Document types | Aircraft lease schedules, clinical trial protocols, loan tapes | Extraction coverage |
| Risk categories | Airworthiness directives; CRE concentration; reimbursement exposure | Rubric — new categories |
| Materiality baseline | A provision routine in one sector, a blocker in another | Severity calibration — the hard one |
Vendors compete on the first. The third is where deals go wrong, and it is the one that cannot be bought — only accumulated, through your own reviewers' corrections.
Regulated sectors have thresholds that are rules, not judgments
In banking, some of what diligence must check is written into regulation with specific numbers. These belong in a deterministic check against maintained thresholds, with the model's role limited to extracting the inputs and quoting where each came from.
Three examples, each verified against primary sources:
- Commercial real estate concentration. Supervisory guidance identifies institutions with CRE loans representing 300% or more of total capital where the portfolio has grown 50% or more over the prior 36 months, and construction and land development loans at 100% or more of total capital, as warranting heightened scrutiny.[1]
- Capital adequacy. A leverage ratio below 4.0% places an institution in the undercapitalized category under the prompt corrective action framework.[2]
- Lending limits. The single-borrower limit is 15% of capital and surplus, with an additional 10% available for loans fully secured by readily marketable collateral.[3]
Asking a language model to decide whether a threshold is breached is a category error. Extraction is probabilistic; the threshold comparison should not be. Where a figure sits close to a limit, the correct output escalates rather than resolves.
The baseline problem, stated precisely
Severity encodes a judgment about what counts as serious, formed on some corpus of deals. When your target does not match that corpus, the model reads the clause correctly and ranks it wrongly.
It also explains why sector expertise cannot be replaced by better models. The information required — what is customary in this industry, at this size, in this decade — is not in the document being read.
What to require for specialist work
- Sector-specific categories, not just keywords. A new risk category with its own weight and threshold, not a general category with sector terms bolted on. If the rubric cannot express “airworthiness directive compliance” as a first-class category, it cannot rank it properly.
- Per-sector severity configuration. The same clause type must be able to score differently in different industries. If severity is a fixed global property, the tool cannot represent the problem.
- Deterministic threshold checks for regulatory limits, with extraction feeding them rather than judging them.
- Absence detection against a sector-specific expected list. Which permits should this business hold? Which certifications, filings, consents? The missing licence is a bigger finding than any clause, and extraction-only systems are structurally blind to it.
- Override capture that feeds calibration. Your specialist reviewers correcting severity is the mechanism by which the baseline becomes yours. A system that improves through your use is worth far more over three years than one that improves when the vendor ships.
Where generic screening still earns its place
Worth saying, because the argument above could be read as “wait for a sector product.”
Most of a specialist target's data room is not specialist. Employment agreements, customer contracts, leases, supplier terms, corporate records — these are ordinary commercial documents, they are the bulk of the room by count, and generic screening handles them well. The sector-specific material is usually a minority of documents carrying a majority of the sector risk.
So a workable split: generic screening across the whole room for coverage and ranking, with the specialist categories configured on top and specialist documents routed to human experts regardless of score. That is achievable now and does not require waiting for a vertical product that may never exist for your industry.
Evaluating a sector claim
Vendors advertise industry versions. Four questions separate a real one from a marketing page:
- Show me the category list for this sector. If it is the generic list with different examples, there is no sector rubric.
- Can we change the weights and thresholds ourselves? If calibration requires a vendor engagement, it will not happen at the pace deals require.
- How are regulatory thresholds handled — model or rule? Listen for whether numeric limits are checked deterministically.
- What does it do about a missing permit? Absence detection is the clearest test of whether a sector rubric is real, because it requires knowing what should exist.
The verification point, sharper here
Grounded commercial legal AI has been measured hallucinating between 17% and 33% of the time in an adjacent task, with providers' hallucination-free claims judged overstated.[4] In specialist work, verification is harder — the person who can confirm a clinical-trial or airworthiness finding is a scarce expert, not whoever ran the scan.
That raises the value of quoted source text considerably. A finding that reproduces the source sentence can be checked by the specialist asynchronously in seconds. One that merely asserts a conclusion requires them to read the document, which is the bottleneck you were trying to relieve.
Bottom line
Document types are the easy difference and the one vendors sell against. Severity calibration is the hard one, and it decides whether specialist output is useful or confidently wrong.
Require sector categories as first-class rubric entries, deterministic checks for regulatory thresholds, absence detection against an expected list, and override capture so your experts' corrections accumulate. Then run generic screening across the ordinary majority of the room, because that is where the coverage gap actually is.
Sources
- Concentrations in Commercial Real Estate Lending, Sound Risk Management Practices, 71 Fed. Reg. 74580 (12 December 2006) — CRE and construction/land development concentration criteria.
- 12 C.F.R. § 324.403 — prompt corrective action capital categories, including the 4.0% leverage ratio threshold for undercapitalized institutions.
- 12 C.F.R. § 32.3(a) — single-borrower lending limit of 15% of capital and surplus, plus an additional 10% for loans fully secured by readily marketable collateral.
- Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., & Ho, D. E. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. arXiv:2405.20362; Journal of Empirical Legal Studies (2025). Measures legal research, not document review. arxiv.org/abs/2405.20362
Regulatory thresholds change; verify against the primary source before relying on any figure. Nothing here is legal or regulatory advice. We cite only sources we have retrieved and read — see our methodology.