Independent Evidence

AI Accuracy & Benchmarks

Does Anweshna actually detect material risks better than a skilled analyst? We publish our exact performance metrics against human-annotated ground truth.

What These Numbers Measure

Precision and recall are computed strictly on our deal-blocking gate—whether a risk category scored at or above the critical blocking threshold (70+)—not on continuous 0-100 scores. This is the decision the product actually forces on a deal team.

Only categories that a reviewer explicitly labeled on a given filing are counted. Unlabeled categories are excluded rather than assumed clean, ensuring we do not artificially inflate our true-negative count to boast a 99% precision rate.

Headline Metrics

Based on our latest 11-filing benchmark run containing 2,473 total findings across 15 risk categories:

Metric Value Description
Precision 84.6% When Anweshna blocks a deal, it is correct 84.6% of the time.
Recall 91.7% When a material risk exists, Anweshna detects it 91.7% of the time.
F1 Score 88.0% Harmonic mean of precision and recall.
False-Positive Rate 15.4% Share of "clean" categories mistakenly flagged as blocking.
Deal-Killer FN Rate 0.0% False Negative rate on designated "deal-killers". Anweshna surfaced 14 out of 14 known deal-killers.

Evidence & Citation Integrity

Every finding surfaced by Anweshna is verified deterministically against the source text. If the AI claims a specific quote or figure exists, our evidence_verifier engine hunts for it in the extracted text.

Metric Value Description
Citation Accuracy 88.5% Share of checkable findings that successfully verified against the source text word-for-word.
Unverified Rate 11.5% Upper bound on hallucination. Findings whose quoted span or figure was not located in the text.

Note: The Unverified Rate is not a pure hallucination rate. It often flags derived arithmetic (where the AI correctly calculated a delta not explicitly printed in the text). It remains an honest upper bound on potential AI invention.

Second-Stage Verification

Anweshna runs a proprietary deterministic second pass over all AI findings to separate actual events from boilerplate language, and to separate potential severity from established materiality.

Anweshna Demo