Methodology

AI Accuracy & Benchmarks

How scoring accuracy is measured here — against real SEC filings anyone can re-fetch, labelled by hand under fixed rules. The most recent completed run: 94.1% precision and 84.2% recall on the deal-blocking gate, across 20 annotated SEC filings and 35 labelled category-cases.

What is being measured

Accuracy here means one thing: does the deal-blocking gate fire when it should, and stay quiet when it should not? A risk category either crossed the review threshold or it did not. That is the decision the product forces on a deal team, and the only one a reviewer can label without inventing a numeric ground truth. Continuous 0–100 scores are not part of the measurement.

Only categories a reviewer explicitly examined and labelled on a given filing are counted. Anything unlabelled is excluded from the arithmetic entirely, rather than counted as a clean result — treating “not examined” as “nothing found” would manufacture correct answers by the hundred and inflate any figure for free.

The ground truth is public

Measurement runs against real annual reports filed with the U.S. Securities and Exchange Commission, retrieved from EDGAR. These are public, free to re-fetch, and redistributable, which means anyone can pull the same filing and check the work. A benchmark built on documents nobody else can see is a marketing claim, not a benchmark.

Each filing is pinned by a cryptographic hash of its extracted text. If the text changes, the measurement refuses to run rather than quietly scoring a different revision against annotations written for the original.

What is labelled, and what is deliberately not

For each filing, a reviewer records which risk categories a competent reader would refuse to close without resolving, and which they read and found clean. Three rules constrain that labelling:

  1. Only categories actually checked are labelled. Anything unlabelled is excluded from the arithmetic rather than counted as a clean result.
  2. A category is called clean only on positive grounds — an explicit disclosure that nothing material exists, or a section read and found clear. “I did not look” is not the same as “there is nothing there.”
  3. Every annotation quote must appear verbatim in the filing. A paraphrased annotation becomes a target that cannot be found, which the scoring engine is then blamed for missing.

Detection of specific transaction-critical findings is tracked separately from the category flag. For each filing, the individual disclosures a deal team would stop on are recorded by name, and whether each was surfaced is measured on its own. A report can trip the right category while missing one of the reasons it should have.

Filings are chosen in matched pairs. For any given signal, the sample aims to hold both a filing where it is genuinely present and one where the same vocabulary appears as routine legal boilerplate. A system can score perfectly on either alone by being uniformly cautious or uniformly alarmed; only the pair distinguishes judgement from bias.

Evidence checks are reported in three parts, not one

Every finding is checked against the source document. Findings are sorted into those whose quoted text and figures were located, those carrying something checkable that could not be located, and those offering nothing checkable at all. The second and third are reported separately because they mean different things: unsupported is not the same as contradicted.

The proportion that fails an automated check is not a hallucination rate and will never be published as one. It is an upper bound. Known causes flag correct findings — for example, a finding that restates two figures from a document and computes the difference states a number that is legitimately absent from the source. A genuine hallucination rate requires reading the flagged findings individually.

Second-stage verification

A deterministic second pass runs over every AI finding to separate actual events from boilerplate language, and potential severity from established materiality.

Results, 31 August 2026

Measured on the deal-blocking gate — whether a risk category scored at or above 70 — across twenty annotated SEC filings, comprising thirty-five labelled category-cases.

These are the figures from the most recent completed run. The benchmark is re-run against the corpus from time to time as the scoring engine changes, and this page is updated with each measurement and the date it was taken.

94.1% Precision 16 of 17 flagged categories were correct · 95% CI 73.0–99.0%
84.2% Recall 16 of 19 expected blockers were caught · 95% CI 62.4–94.5%
1 of 24 Deal-killers missed 4.2% false-negative rate on findings a reviewer called transaction-critical
MetricValue95% CIBasis
Precision94.1%73.0–99.0%16 TP / 17 flagged
Recall84.2%62.4–94.5%16 TP / 19 expected
F188.9%——
False-positive rate6.2%1.1–28.3%1 FP / 16 labelled clear
Deal-killer miss rate4.2%—1 of 24
Sample20 filings—35 labelled category-cases

Read the interval, not the point estimate. Thirty-five labelled cases is a small sample: one category flipping either way moves precision by several points. Two measurements whose intervals overlap are not distinguishable, and reporting the gap between them as an improvement would be reading noise. The labels are the work of a single reviewer, which is a real limit on what any figure here can carry.

The second-stage verifier produced 5,285 individual records across the same twenty filings: 3,530 marked high-priority verification, 1,754 screening signal, one potential blocking condition, and zero confirmed blocking conditions. The confirmed rung is the strongest statement the system can make, and it did not fire on a single real filing. That is the intended behaviour — confirming a blocking condition requires affirmative evidence in the document plus human review, and a screen is not entitled to make that call on its own.

Where it performs worst

Publishing the weakest categories was committed to on this page before the numbers were known. All three failures below are the same shape: the finding was scored, but not scored high enough to clear the 70 gate.

CategoryRecallWhat was missedScore against the 70 gate
HR & labour0% (0 of 1)Bird Global FY202260
IP & licensing0% (0 of 1)Nutanix FY2022 10-K/A35
ESG & environmental50% (1 of 2)Devon Energy FY202362

The single false positive in the run is in financial: Mitek Systems FY2025 scored 75 in a category a reviewer labelled clear. That filing discloses a material weakness seventeen times over, every instance of it remediated, with an unqualified auditor opinion on internal control. The screen read the disclosure and did not fully credit the remediation.

Two categories — HR & labour and IP & licensing — are represented by a single labelled positive each in this run, so their 0% recall is one miss, not a rate. It is reported as a named failure rather than a statistic for that reason.

The finding it missed

One finding out of twenty-four that a reviewer marked transaction-critical was not surfaced. Naming each one individually was committed to on this page in advance; this is the whole list.

Ocean Thermal Energy Corp, FY2025 — related-party payments to a CEO-controlled entity. The filing states: “In May 2023, we entered into a month-to-month agreement with a company controlled by our chief executive officer for shared use of an office and facilities in Lancaster, Pennsylvania.” The screen did not raise this as a distinct finding.

The category still gated correctly, because a second related-party finding in the same filing — accrued interest on related-party notes payable — was surfaced and carried related_party above the threshold. A deal team reading the report would have been directed to related-party exposure in this filing. They would not have been pointed at this particular arrangement, and would have had to find it themselves.

Five filings, and what the screen did with them

Each of these is a real SEC filing in the corpus, pinned by hash. Every quote below is the verbatim passage a reviewer read when labelling it, checked against the pinned text automatically. Follow the link to read the source.

Clorox FY2024 — a real incident next to the boilerplate

Every large filer carries ransomware and phishing language in its risk factors. This filing carries that too, and also an incident that actually happened: “On Monday, August 14, 2023, the Company disclosed it had identified unauthorized activity on some of its IT systems.” The screen gated cybersecurity and surfaced the realised incident rather than the surrounding boilerplate. Separating the two is most of the job. Read the filing.

iRobot FY2022 — a second request and going-concern doubt

“On September 19, 2022, the Company and Amazon each received a request for additional information and documentary material (the ‘Second Request’) from the Federal Trade Commission.” Both categories a reviewer labelled blocking — regulatory_merger_control and financial — were gated. Read the filing.

ChampionX FY2024 — the same pattern, a different regulator

“On July 2, 2024, SLB announced that ChampionX and SLB had each received a Request for Additional Information and Documentary Material (collectively, the ‘Second Request’) from the United States Department of Justice.” regulatory_merger_control gated. Paired with iRobot, this is the category working on two filings, two regulators and two transactions rather than once by luck. Read the filing.

Nutanix FY2022 10-K/A — found, but not gated

This is the instructive failure. The filing discloses “our controls were not designed and operating effectively to provide the information necessary for our risk assessment process to identify non-compliant use of third-party software.” The screen did surface that text. It scored ip_licensing at 35, well under the gate, so the report did not raise it as a blocking category. Surfacing a finding and gating on it are different operations, and this filing is where the difference shows. Read the filing.

Ocean Thermal Energy FY2025 — the one it missed

The related-party arrangement named in the section above. The category gated on a sibling finding; this particular one was not surfaced. Read the filing.

A low overall score does not mean a filing is clean. Clorox scored 46.4 overall and Nutanix 23.6, and both were flagged for mandatory review. Blocking is decided per category, not by the overall number: one blocking category over 70 flags the document however calm the weighted average looks. The overall score is a summary, not the decision.

Frequently asked questions

How accurate is Anweshna?

On the most recent measurement, the deal-blocking gate reached 94.1% precision and 84.2% recall across twenty annotated SEC filings and thirty-five labelled category-cases. The 95% confidence intervals are 73.0-99.0% and 62.4-94.5% respectively, which are wide because the sample is small. The labels were produced by a single reviewer. Both facts are limits on how much weight the headline figures can carry.

What is the benchmark corpus?

Thirty United States SEC annual filings, 10-K and 10-K/A, annotated by hand for M&A deal-blocking risk. Every filing is pinned by SHA-256 and every annotated finding carries a verbatim contiguous quote from the source, checked automatically against the pinned text. The filings themselves are public, so anyone can re-fetch them and read the same text. The labelled corpus is not published.

Can I reproduce the results?

Not independently, today. The filings are public SEC documents, several are linked on this page, and the method and gating rule are described here. The full labelled corpus and the model output for each filing are not published, so the headline figures cannot yet be re-derived without them.

What does the system get wrong?

In this run it missed three blocking categories and produced one false positive. The misses were HR and labour on Bird Global, scored 60 against a gate of 70; IP and licensing on Nutanix, scored 35; and ESG and environmental on Devon Energy, scored 62. The false positive was financial on Mitek Systems, scored 75 in a category a reviewer labelled clear because the disclosed material weakness had been remediated.

Does a high score mean a deal is blocked?

No. Blocking is decided per category, not by the overall score: a single blocking category at or above 70 flags the document for mandatory review however low the overall number looks. The overall score is a summary of the document, not the decision about it. Nothing the system produces blocks a deal on its own, and confirming a blocking condition requires human review.

Why publish the failures?

This page committed to naming the weakest categories and every transaction-critical finding the system failed to surface before those numbers were known, which is the only point at which such a commitment means anything. A benchmark that reports only its wins tells a buyer nothing they can rely on.

Anweshna Demo