In 2023 the market filled up with "chat with your PDF" tools. The pitch was simple: drop a 500-page merger agreement into a chat window and ask, "any risks?" For institutional due diligence — M&A lawyers, PE associates, investment bankers — that approach doesn't hold up. A general-purpose model is built to be helpful and fluent. A due diligence read needs to be skeptical, exhaustive, and repeatable, and those are different jobs.
What goes wrong with a bare model
Ask a generic AI wrapper to find risks in a deal document and three failure modes show up quickly:
- Boilerplate read as risk. Standard mechanics like a stockholder-approval clause or appraisal-rights language get flagged as material because they sound legally dense, with nothing separating routine transaction plumbing from an actual problem.
- Confident invention. Pressed for a specific clause, a model will sometimes produce one that reads plausibly and isn't in the document.
- Instability. The same prompt against the same document can come back with a different severity on a different run.
Anweshna doesn't ask the model to make the final call. It uses the model to read and describe, and puts every one of those outputs through deterministic code before it reaches a person. The six checks below are the ones doing that work today, verified directly against the shipped codebase rather than described in the abstract.
1. An industry-specific ruleset, not an open-ended prompt
The model isn't asked to "find risks" in general. It's given a fixed category taxonomy for the document's industry — 15 categories for M&A, a different set for banking or pharma rulesets — and extracts findings against that structure. What counts as a category, and which of those categories can block a deal outright, is defined in a versioned ruleset file, not left to the model's judgment on the day.
2. Two scoring paths, not one context window
A large data room can exceed what a single model call can hold. Rather than lean on retrieval that can miss a needle-in-the-haystack clause, the Deep Scan path chunks a long document, scores each chunk independently, and reconciles the findings back into one report. A faster Precision Scan path handles shorter filings in a single call. Which path runs is a function of document size and plan tier, not a guess.
3. Absence is treated as a finding, not a clean bill of health
The hardest thing for a model to notice is what a document doesn't say. For each blocking category — financial, legal, HR & labour, anti-bribery & sanctions, IP & licensing, regulatory & merger control, and AI & tech governance — a document set that comes back with zero findings for that category doesn't score as clean. It's floored to a small non-zero score and stamped with an explicit synthetic finding: "No disclosure found for [category] — treat as a diligence gap requiring seller confirmation, not evidence the category is clean." Silence on anti-bribery or merger control is itself a gap worth flagging, not evidence there's nothing there.
4. Every finding is checked against the source text it claims to quote
This is the part that most separates the platform from a chat wrapper. Every finding the model produces is run back against the original extracted document text, searching for the exact span or figure it claims to be quoting. A finding that can't be located in the source is tagged [UNVERIFIED]; a finding that never offered anything checkable in the first place is tagged separately as [UNVERIFIED-NO-EVIDENCE]. Neither tag is optional or model-controlled — the check runs on everything the model returns.
5. A second pass reclassifies what the first pass found
A separate deterministic pass re-evaluates every finding's evidence type and materiality after the fact, independent of whatever severity the model assigned it. If the model flags a "no-shop" provision as a risk, this pass recognizes it as standard transaction mechanics and reclassifies it accordingly, rather than letting routine deal language inflate the findings list. The goal is that only findings backed by something that actually happened — not boilerplate, not a hedge, not a hypothetical — reach a reviewer as a live issue.
6. Negation and hedging are checked, and known deal-killers can't be argued down
A document that explicitly says "the Company is not subject to any material litigation" shouldn't trip a litigation flag — the negation-detection layer checks the clause around a risk phrase, both before and after it, before letting it register as a finding. In the other direction, a short list of confirmed, terminal phrases — going concern, environmental remediation order, an active criminal investigation, a confirmed sanctions breach — force a category into a blocked state regardless of what score the model itself assigned. The model can soften language; it can't talk its way out of one of these.
Where that leaves the numbers
Accuracy is measured against a public, human-annotated corpus of SEC filings — real documents anyone can re-fetch and check, labelled under fixed rules, hash-pinned so the measurement can't quietly score a different revision. No headline precision or recall figure is published yet: the sample is not yet confirmed balanced, and a number drawn from a single reviewer's labels would read as more precise than it is. When it is published it will come with its sample size, and the categories the system does worst on, attached. The method is set out in full on the Accuracy & Benchmarks page.
None of this makes the model disappear. It's still doing the reading. What changes is that nothing it says reaches a reviewer without first passing through code that doesn't get tired, doesn't get talked into a softer read, and checks its work the same way on every document.