Methodology

How Anweshna Scores Risk

Anweshna assigns every uploaded document a 0–100 score in each of 15 M&A risk categories, plus 8 additional categories for institutional banking transactions. This page sets out where those categories come from, how they are weighted, what makes a category “blocking,” and — equally important — what the resulting score does not claim to establish.

What the score is, and what it is not

Anweshna performs document-level risk screening, not due diligence. A score reflects what the uploaded documents disclose and how that language reads against a defined rubric. It involves no independent verification, no primary research, no site visits, no management interviews, and no legal judgement. It does not replace counsel, an accountant's quality-of-earnings work, or a regulator's own review.

The intended use is triage: taking a data room too large to read sequentially and producing a ranked queue, so that human reviewers start with the documents most likely to contain material findings. Every output is designed to be checked, not trusted blindly — which is why findings carry the quoted source text they were derived from.

Technology & Differentiation

Raw foundation models lack the precision required for institutional due diligence. Discover how our proprietary 7-layer pipeline strictly separates natural language understanding from deterministic risk scoring.

Accuracy & Benchmarks

Does the methodology work? We hold ourselves to a strict standard of independent validation. Our accuracy is measured continually against human-annotated real-world filings using an open benchmark.

Security & Compliance

Security isn't an afterthought — it's baked into the architecture from the ground up. Every client's documents and findings are encrypted at rest under a key derived uniquely for your engagement.

Enterprise FAQ

Evaluating AI for institutional use requires answering hard questions about data privacy, model training, and vendor lock-in. We answer the top 30 questions directly.

The 15 M&A risk categories and their weights

Each category is scored independently on a 0–100 scale. The composite deal score is the weighted sum of the category scores; the weights below total 1.00.

M&A ruleset category weights and blocking status.
CategoryWeightBlocking
Financial0.18Yes
Legal0.13Yes
HR & Labour Compliance0.12Yes
Regulatory / Merger Control0.10Yes
Anti-Bribery, Sanctions & AML0.08Yes
AI & Tech Governance0.07Yes
Cybersecurity0.07
IP & Licensing0.06Yes
Operational0.04
Supply Chain & Geopolitical0.04
Tax0.03
Related Party0.03
ESG & Environmental0.03
Reputational0.01
Data Privacy0.01

Weights encode deal-materiality, not frequency. Financial carries the largest weight because financial findings most often change price or kill a transaction outright; reputational carries the smallest because reputational signals in documents are typically corroborative rather than decisive. See the full category glossary for what each one screens for.

Blocking categories and the 70 threshold

Seven of the fifteen categories are blocking: Financial, Legal, HR & Labour, Regulatory / Merger Control, Anti-Bribery Sanctions & AML, AI & Tech Governance, and IP & Licensing. A score of 70 or above in any blocking category flags the deal for mandatory human review before it proceeds, regardless of how low the composite score is.

This is deliberate: a composite score can average away a single transaction-ending finding. A sanctions exposure does not become tolerable because the tax position is clean, so blocking categories are evaluated on their own terms rather than through the weighted total.

Absence of disclosure is not evidence of absence

Every blocking category carries a baseline floor on absence. If a document set contains no findings at all in a blocking category, that category does not score zero — it receives a small non-zero floor plus an explicit “no disclosure found” finding directing the reviewer to request the missing material. A silent data room is a diligence question, not a clean result.

Critical-phrase floors

Language models can soften genuinely severe findings. To stop that, a defined list of phrases forces a minimum score in the relevant category regardless of what the model itself assigned, in two tiers:

  • Always-block phrases floor the category to 95. These are confirmed, terminal findings — an FCPA violation, a sanctions breach, an SDN-list match, a criminal investigation, an environmental remediation order, a going-concern qualification.
  • Remaining critical phrases floor the category to 85. These indicate a serious but unresolved state rather than a confirmed outcome — for example a second request or Phase II merger review, which signals scrutiny rather than a decided result.

Floors are checked against negation. Explicit disclaimers (“no litigation pending”), hypothetical framing (“stress-test scenario”), and forward-looking diligence recommendations (“must be confirmed”) are detected so that a document stating a risk is absent is not scored as though the risk were present.

How sources are handled in our guides

The written guides on this site follow three rules, which are worth stating because they constrain what appears on the page:

  1. Statutory and regulatory figures are cited to primary sources — the United States Code, the Code of Federal Regulations, or the issuing agency — not to secondary summaries.
  2. Where a figure cannot be traced to a primary source, the figure is changed rather than the citation. Claims are softened to qualitative statements or removed. An approximate source is never attached to a precise claim.
  3. Ranges drawn from practice rather than statute are labelled as such. Timelines and cost bands reflect commonly observed patterns, vary substantially by transaction, and are not represented as authoritative.

Guides that cite statutory figures carry a Sources section listing those primary references. Thresholds and timelines change; verify current requirements with the relevant agency before relying on them.

Emerging risks outside the rubric

A fixed rubric cannot anticipate everything. On the deeper scanning tier, a separate detection pass looks for material risks that fall outside the defined categories and would not trip an existing critical-phrase trigger. These are logged as candidates for human review only.

A candidate never affects a score. It cannot change a category result, a composite score, or a blocking decision, and it never edits a ruleset automatically. Where a pattern recurs across several independent documents, it is clustered into a written proposal that a person must approve, after which any actual rule change is made through normal engineering review. Discovery proposes; humans decide.

How we measure accuracy

Scoring claims are worth what the evidence behind them is worth, so this section describes how accuracy is measured here — and states plainly that no accuracy figure is published yet, because the sample is not large enough to support one.

The ground truth is public

Measurement runs against real annual reports filed with the U.S. Securities and Exchange Commission, retrieved from EDGAR. These are public, free to re-fetch, and redistributable, which means anyone can pull the same filing and check our work. A benchmark built on documents nobody else can see is a marketing claim, not a benchmark.

Each filing is pinned by a cryptographic hash of its extracted text. If the text changes, the measurement refuses to run rather than quietly scoring a different revision against annotations written for the original.

What is labelled, and what is deliberately not

For each filing, a reviewer records which risk categories a competent reader would refuse to close without resolving, and which they read and found clean. Three rules constrain that labelling:

  1. Only categories actually checked are labelled. Anything unlabelled is excluded from the arithmetic entirely rather than counted as a clean result. Treating "not examined" as "nothing found" would manufacture correct answers by the hundred and inflate any accuracy figure for free.
  2. A category is called clean only on positive grounds — an explicit disclosure that nothing material exists, or a section read and found clear. "I did not look" is not the same as "there is nothing there."
  3. Every annotation quote must appear verbatim in the filing. A paraphrased annotation becomes a target that cannot be found, which the scoring engine is then blamed for missing.

Where a reviewer is unsure, the category is left unlabelled. An honest gap makes the sample smaller; a guess makes it wrong.

Measured against the decision, not the number

Accuracy is measured against the flag — whether a category crossed the review threshold — not against the composite score. That is the decision a deal team actually acts on, and the only one a reviewer can label without inventing a numeric ground truth.

Detection of specific transaction-critical findings is measured separately. For each filing, the individual disclosures a deal team would stop on are recorded by name, and whether each was surfaced is tracked independently of the category flag. A report can trip the right category while missing one of the reasons it should have.

Filings are chosen in matched pairs. For any given signal, the sample holds both a filing where it is genuinely present and one where the same vocabulary appears as routine legal boilerplate. A system can score perfectly on either alone by being uniformly cautious or uniformly alarmed; only the pair distinguishes judgement from bias.

Evidence checks are reported in three parts, not one

Every finding is checked against the source document. Findings are sorted into those whose quoted text and figures were located, those carrying something checkable that could not be located, and those offering nothing checkable at all. The second and third are reported separately because they mean different things: unsupported is not the same as contradicted.

The proportion that fails an automated check is not a hallucination rate and will never be published as one. It is an upper bound. Known causes flag correct findings — for example, a finding that restates two figures from a document and computes the difference states a number that is legitimately absent from the source. A genuine hallucination rate requires reading the flagged findings individually, and we have not finished doing that.

Current status

The measurement harness is built and running. The annotated sample now covers thirty filings — the target it was set — and every one of the fifteen risk categories carries at least one label. Thirteen of the fifteen are matched in both directions; two are represented in one direction only, because the filing needed for the opposite case is genuinely hard to find rather than merely unwritten. Preliminary results exist internally and are deliberately not published here, because a figure drawn from a sample labelled by a single reviewer would carry an air of precision it has not earned — and it would be quoted without its caveats.

Reaching a filing count is not the same as the sample being balanced, and it is treated here as the former. The results will be published in full once that is settled, including the categories where the system performs worst and every transaction-critical finding it failed to surface, named individually. That commitment is stated before the numbers are known, which is the only point at which stating it means anything.

Known limitations

  • Scanned documents without a text layer cannot be scored. Pre-flight checks reject unreadable files rather than scoring them as clean.
  • Very large documents are chunked, and where a document exceeds the single-pass window it is either routed to the chunked engine or, with explicit consent, truncated — in which case the stored report is marked as truncated.
  • Scores are not comparable across different rulesets. A 62 under the banking ruleset does not mean the same thing as a 62 under the general M&A ruleset.
  • The system reads what it is given. It cannot detect what a seller chose not to place in the data room, which is precisely why absent disclosure in a blocking category raises a flag rather than passing silently.
Anweshna Demo