Why generic AI wrappers fail at due diligence
Using a general-purpose AI chatbot to scan a 300-page merger agreement presents four immediate liability risks for a deal team:
- Hallucination: Chat models will invent clauses that sound plausible when asked leading questions.
- Instability: Asking the exact same prompt twice can yield different severities or entirely different findings.
- Boilerplate blindness: Generic AI struggles to separate a standard "no-shop" provision from an actual transaction risk.
- Extraction blindness: A model reads whatever the PDF parser handed it. If the parser split a currency symbol from its digits, or emitted a typographic apostrophe, a naive citation check silently fails against text that is verbatim correct.
Anweshna solves this by treating the LLM strictly as a text-extraction module within a larger, deterministic software pipeline. Severity assessment, citation checking, category gating and the deal-blocking decision all happen in code that behaves identically every run.
The 7-Layer Architecture
1. Multi-Industry Ruleset Architecture
Before scanning begins, documents are routed through industry-specific taxonomies. A pharmaceutical IP agreement is evaluated under vastly different materiality thresholds than a commercial real estate lease. Our taxonomy governs what the AI is allowed to look for.
2. Dual Scoring Engines
Documents pass through either our Precision Scan (for targeted agreements) or our Map-Reduce Deep Scan (for massive, 500+ page data room dumps). The Deep Scan chunks, analyzes, and synthesizes findings across a document hierarchy while preserving global context. An oversized or unusually dense filing is re-routed to Deep Scan automatically rather than being silently truncated. Both engines are held to identical published verdicts on identical findings by automated parity checks — the engine you are entitled to never changes the answer.
3. Missing-Information Detection
A risk doesn't just exist in what is written; it exists in what is omitted. Our scoring_policy engine maintains baselines of required disclosures per industry and deterministically flags when critical schedules, appendices, or financial clauses are missing from the data room.
4. Negation-Aware Risk Detection
Contracts frequently list risks they have explicitly mitigated. Our two-tier negation detection engine ensures that "Company is NOT subject to any litigation" does not trigger a false-positive litigation flag.
5. Evidence Verification (Citation Integrity)
Every finding surfaced by the AI must carry a quoted span from the source document. Our proprietary evidence_verifier engine strips away the LLM's response and deterministically searches the original extracted PDF/Word text. If the exact quote or figure cannot be located, the finding is tagged [UNVERIFIED].
6. Extraction-Artefact Normalisation
Real filings do not arrive as clean text. Across the 30 SEC filings in our benchmark corpus we measured 6,664 typographic apostrophes and 10,785 currency figures where extraction split the symbol from its digits — both more common than the tidy forms a model naturally writes. Layer 5 compares quotes against the source, so an unhandled artefact means a truthful citation fails its own check. We normalise these before verification, and deliberately never touch accented characters, which are real letters inside real company names.
7. Risk Taxonomy & Blocking Gates
The final layer computes the composite deal score across 15 risk categories, seven of which are deal-blocking. We enforce two-tier "critical-phrase floors" — absolute deal-killers such as an FCPA violation or a going-concern qualification floor their category outright, a second band floors slightly lower — unconditionally overriding the model's subjective score and forcing a "Blocked" state until a human reviewer intervenes. A blocking category that discloses nothing does not read as clean either: it earns an explicit "no disclosure" finding and a diligence request.
Where the pipeline runs
The seven layers above describe how one document is scored. Deals are not one document.
- Data-room scale. Bulk ingestion accepts up to 25 files per request and scores them through the identical single-document path — the same gates, the same auto-routing, the same verification. Batch work is never a weaker code path. Reports can be pulled back as a single archive.
- Clean rooms enforced in the database, not in policy. The Clean Rooms add-on stores clean-room documents and reports in a separate Postgres schema, readable only by a dedicated database role and encrypted with a key of their own; the role that serves ordinary deal rooms has no permission on that schema. Inside a room, access follows each member's room role, output reaches the deal team only through a release a lead or counsel approves, and the room's audit log is hash-chained and append-only at the database level.
- Agent-reachable. A hosted Model Context Protocol server at
/mcp/exposes screening as tools, so Anweshna can be driven from Claude, ChatGPT, or your own agentic workflow — see the API & MCP integration guide. Every call re-enters the platform through its own API, so billing, quota, role and clean-room gates apply exactly as they do to a human. A failed scan surfaces as an explicit error, never as an empty — and therefore clean-looking — result. - Provenance on every export. Markdown, JSON, PDF and DOCX downloads carry who pulled the report, in what role, and when — stamped in the reader's own timezone.
The Value Moat
The moat is not the model. Anyone can call the same foundation models we do. The moat is everything that happens after the model answers: severity floors it cannot argue with, citation checks that read the source document rather than the model's own summary, category gates that fail closed, and two engines held to the same verdict.
Every one of those is deterministic code. Run the same document twice and you get the same answer — which is the minimum bar for a document a deal team is going to rely on, and the bar a chat interface cannot clear.
How we measure this — the public annotated corpus, the labelling rules and the measured result — is set out on the Accuracy & Benchmarks page. The current figure is 94.1% precision and 84.2% recall on the deal-blocking gate, published with its sample size, its confidence intervals, the categories it scores worst and the one transaction-critical finding it failed to surface, named.