What These Numbers Measure
Precision and recall are computed strictly on our deal-blocking gate—whether a risk category scored at or above the critical blocking threshold (70+)—not on continuous 0-100 scores. This is the decision the product actually forces on a deal team.
Only categories that a reviewer explicitly labeled on a given filing are counted. Unlabeled categories are excluded rather than assumed clean, ensuring we do not artificially inflate our true-negative count to boast a 99% precision rate.
Headline Metrics
Based on our latest 11-filing benchmark run containing 2,473 total findings across 15 risk categories:
| Metric | Value | Description |
|---|---|---|
| Precision | 84.6% | When Anweshna blocks a deal, it is correct 84.6% of the time. |
| Recall | 91.7% | When a material risk exists, Anweshna detects it 91.7% of the time. |
| F1 Score | 88.0% | Harmonic mean of precision and recall. |
| False-Positive Rate | 15.4% | Share of "clean" categories mistakenly flagged as blocking. |
| Deal-Killer FN Rate | 0.0% | False Negative rate on designated "deal-killers". Anweshna surfaced 14 out of 14 known deal-killers. |
Evidence & Citation Integrity
Every finding surfaced by Anweshna is verified deterministically against the source text. If the AI claims a specific quote or figure exists, our evidence_verifier engine hunts for it in the extracted text.
| Metric | Value | Description |
|---|---|---|
| Citation Accuracy | 88.5% | Share of checkable findings that successfully verified against the source text word-for-word. |
| Unverified Rate | 11.5% | Upper bound on hallucination. Findings whose quoted span or figure was not located in the text. |
Note: The Unverified Rate is not a pure hallucination rate. It often flags derived arithmetic (where the AI correctly calculated a delta not explicitly printed in the text). It remains an honest upper bound on potential AI invention.
Second-Stage Verification
Anweshna runs a proprietary deterministic second pass over all AI findings to separate actual events from boilerplate language, and to separate potential severity from established materiality.
- Boilerplate Demotion: Standard merger clauses (e.g. "no-shop" provisions, appraisal rights) are identified and demoted from "Risk" to "Transaction Mechanic".
- Inference Flagging: Speculative or forward-looking AI statements are labeled as Model Inference to ensure human review.
- Diligence Gaps: Missing information is surfaced explicitly rather than hidden.