FinProof v1 spans 7 attack categories — investment advice, KYC bypass, regulatory misrepresentation, document hallucination, data rights, transaction integrity, and account bypass — across three deployment registers: professional compliance, retail customer mobile, and RM internal.
Medium-difficulty attacks are generated with a Quantum Circuit Born Machine (QCBM) on PennyLane, producing diverse adversarial coverage that static datasets tend to miss — the method and dataset are published on the public benchmark.
Across all tiers and registers.
BFSI-specific threat taxonomy.
Compliance, retail, internal.
Augmented attack diversity.
Eval harness + 782 benign FPR-calibration examples.
1,606 direct-difficulty adversarial prompts.
2,036 medium-difficulty attacks from a QCBM on PennyLane.
1,747 hard attacks — evaluated by Zytra only.
Eliciting unlicensed or non-compliant financial recommendations.
Attempts to circumvent identity verification controls.
Inducing false claims about products, terms or compliance.
Fabricating statements, figures or official documentation.
Privacy violations and unauthorized data disclosure.
Manipulating payments, transfers or transaction logic.
Unauthorized access to accounts or privileged actions.
Full category definitions are published in the open attack taxonomy on Hugging Face.
How leading safety models perform. Lower false-positive rate means fewer legitimate customers wrongly blocked.
| Rank | Model | HackaPrompt R | AgentHarm FPR | WildGuard F1 | Latency |
|---|---|---|---|---|---|
| 1 | Aval v1.5 Zytra | 0.994 | 0.5% | 0.303 | 11.6ms |
| 2 | PromptGuard-86M Meta | 1.000 | 96.9% | 0.095 | 8ms |
| 3 | LlamaGuard-3-1B Meta | 0.0% | 0% | 0.0 | ~60ms |
| 4 | Granite Guardian IBM | 0.0% | 45% | 0.0 | ~100ms |
HackaPrompt R (recall) and AgentHarm FPR come from different benchmark suites, so a high recall score does not imply a low false-positive rate — Meta’s PromptGuard leads on raw recall (1.000) and latency (8ms) but flags 96.9% of legitimate queries. Official evaluation on the withheld Tier 4 set is conducted by Zytra. Public self-evaluation (Tier 1 + 2) is available now.
The first BFSI-specific AI governance maturity model that scores People, Process, and Technology as separate evidence-gated capability axes in every governance domain — verified against a review of ~35 published frameworks.
Ad Hoc to Optimised, each with a verifiable gate test.
Risk to Agentic AI — the full governance lifecycle.
People, Process, Technology — scored separately.
The weakest cell sets the level — no averaging.
What each level looks like inside a bank, who sits where, why organisations stall, and the gate to the next level.
The 27-cell grid, evidence criteria per level, gate tests, and the imbalance diagnostics — shelfware, paper governance, hero-dependence.
Self-assess all nine domains across People, Process, and Technology. Get your level, grid, imbalance flags, and priority gaps.
AGMM workshops run from a 90-minute maturity sprint to a full-day agentic AI governance certification. Talk to us about a facilitated assessment.
Run against the FinProof withheld test set and see where you rank.