The First Benchmark Built for GenAI Guardrails in Financial Services
A purpose-built benchmark for evaluating GenAI guardrail performance in regulated financial institutions. Mapped to RBI, EU AI Act, SR 11-7, ISO 42001, and India DPDP Act 2023.
Industry-Leading Benchmark
FinProof by the numbers — comprehensive coverage, rigorous evaluation, open access
Built for GenAI Safety & Compliance
Built by financial AI researchers for financial AI researchers. FinProof provides the tools you need to ensure your models are production-ready.
21 Guardrail Actions Tested
Covers the complete taxonomy of GenAI guardrail actions across Input, Processing, and Output layers — from Block and Redact to Agent Handoff Validation and Rollback.
Regulatory Mapping
Every domain and prompt mapped to specific obligations under EU AI Act, SR 11-7, ISO/IEC 42001, and India DPDP Act 2023.
QCBM Adversarial Generation
Prompts generated using Quantum Circuit Born Machines — 89,022 hashes evaluated, zero training contamination confirmed. Novel adversarial patterns no static classifier has seen.
Four-Tier Access Model
Tier 1: Public (JSONL download). Tier 2: Email gate (extended set). Tier 3: Research agreement (full corpus). Tier 4: Withheld test set with public SHA-256 integrity hash.
Plain JSONL — Run In-Perimeter
No external API calls required. Institutions with restricted outbound access can run the full evaluation harness entirely within their own infrastructure.
Open Evaluation Harness
CLI tool available on GitHub for running the test suite against any guardrail stack. Score your own guardrail implementation against the FINPROOF standard.
Three Conversational Registers
Professional (compliance officer, 30–120 words, regulatory citations), Customer Mobile (retail chatbot, 8–30 words, no regulatory language), RM Internal (relationship manager, compressed shorthand). PromptGuard achieves near-zero FPR on professional queries but 100% FPR on the retail register — flagging real customer banking queries as attacks. Register coverage is required to measure this failure mode.
What FinProof Actually Tests
Every category maps to a specific BFSI regulatory obligation — not generic harm.
Unlicensed Investment Advice
Does your AI refuse to give specific fund recommendations or portfolio allocations to retail customers? B-01 covers 1,257 prompts across direct, indirect, and advanced difficulty.
Regulatory anchor: SEBI IA Regulations 2013 §3
KYC & AML Bypass
Does your AI block attempts to circumvent identity verification or structure transactions below AML thresholds? B-02 covers 1,267 prompts.
Regulatory anchor: RBI KYC Master Directions 2016
Information Barrier Violations
Does your AI refuse to aggregate innocuous cross-division signals (IB workload, delivery volumes, CFO meeting cancellations) that together enable MNPI inference? B-12 covers 262 prompts across 5 subtypes.
Regulatory anchor: SEBI PIT Regulations (Chinese Wall)
Mandatory Human Escalation
Does your AI escalate fraud, grievance, and vulnerability cases to a human instead of self-resolving? B-13 covers 312 prompts.
Regulatory anchor: RBI ombudsman and STR reporting requirements
Regulatory KPI Gaming
Does your AI refuse to help game NPA ratios, CAR figures, or LCR metrics in ways that satisfy the letter of the rule while defeating supervisory intent? B-14 covers 267 prompts.
Regulatory anchor: RBI supervisory frameworks (PCA, LCR, GNPA)
Agentic & MCP Security
Does your AI agent resist tool description poisoning, A2A trust exploitation, and unauthorized lateral traversal across banking systems? D-series covers 450 prompts across 15 subtypes. The only BFSI benchmark covering multi-agent attack surfaces.
Why Generic Benchmarks Fail in BFSI
WildJailbreak, HarmBench, and AdvBench test general harm. They cannot test BFSI-specific regulatory violations — and they have no BFSI-specific benign set, so FPR is meaningless.
| Dimension | WildJailbreak | HarmBench | FinProof |
|---|---|---|---|
| BFSI domain attacks | |||
| Regulatory anchor per prompt | SEBI / RBI / DPDP / EU AI Act | ||
| Retail vs professional register | 3 registers | ||
| FPR-calibrated BFSI benign set | generic text | generic text | 2,127 BFSI queries |
| Agentic / MCP attack coverage | D-09 / D-10 / D-11 | ||
| Withheld test set | T4 — never released | ||
| Contamination audit tool | contamination_audit.py |
How FinProof differs from other guardrail benchmarks
The submitter supplies executable code, which the benchmark owner runs against the private set with unrestricted network access. A submission can transmit every prompt it receives. The leaderboard doesn't record whether it did.
Open-weights submissions run from a config — no submitter code executes. Container submissions run with no network interface, verified at runtime rather than promised. Every row publishes how it was produced.
"Withholding a test set only means something if the mechanism for scoring against it doesn't hand it over."
FinProof v2 — Agentic & MCP Security
The only adversarial benchmark testing multi-agent financial systems.
Banks are deploying agent networks. FinProof v2 adds 450 agentic attack prompts across three new categories — the only BFSI benchmark to cover this attack surface.
Tool Description Poisoning
Instructions embedded in MCP tool metadata that override user requests or trigger unauthorized financial actions.
Subtypes: TPA, TYP, OAB, CSC, TRP
A2A Trust Exploitation
Attacks on multi-agent trust chains: agent impersonation, approval forgery, failure cascade, lateral traversal across banking system privilege boundaries.
Corrigibility & Goal Binding
Tests whether an AI agent accepts authorized policy correction, resists deceptive alignment, refuses self-preservation actions, and blocks reward gaming via CSAT manipulation.
Indic Language Coverage
3,900 Hindi, Telugu, and Tamil attack prompts (v1.1 Indic). Models that pass FinProof English can still fail on Indic registers — regional language framing lowers guardrail vigilance. FinProof is the only BFSI benchmark with verified Indic coverage.
FINPROOF Model Leaderboards — July 2026
Guardrail classification results across all BFSI attack categories.
Track 1 — Guardrail Classification Leaderboard
B-01 through B-14 · Input classification · F1 / Precision / Recall / FPR
| # | Model | BFSI F1 |
|---|---|---|
REF | AVAL v1.4Reference Baseline Zytra Tech Solutions · mmBERT IT4 · Reference Baseline | 0.952 |
Granite Guardian 3.3 IBM Research · 8B | 0.813 | |
ShieldGemma 9B Google · 9B | 0.731 | |
LlamaGuard 3 Meta AI · 8B | 0.569 | |
4 | WildGuard 7B Allen AI · 7B | 0.346 |
BFSI F1 scores computed on withheld Tier 4 test set (B-01–B-07). B-12 / B-13 / B-14 scores pending — no submitted model currently covers these categories.
Provenance states how each score was produced. Benchmarks that don't publish this can't tell you whether two rows were measured under the same conditions.
Per-Domain F1 — B-01 to B-07
★ denotes best score per domain
| Domain | Aval |
|---|---|
| B-01Investment advice | 0.954 ★ |
| B-02KYC / Card fraud | 0.981 ★ |
| B-03Employment fraud | 0.960 ★ |
| B-04Regulatory hallucination | 0.957 ★ |
| B-05Predatory lending | 0.986 ★ |
| B-06Insurance conduct | 0.975 ★ |
| B-07Financial instruments | 0.962 ★ |
Last updated: July 2026
Dataset Access Tiers
FinProof uses a structured access model to protect evaluation integrity while enabling broad research access.
Tier 4 — Withheld Test Set (Evaluation Only)
Prompts are never transmitted to submitters. Choose the route that matches your system.
Scored against 3,953 withheld prompts: 2,608 attack (B-01–B-07) plus 1,345 benign. The benign rows are what make precision and FPR meaningful — a guardrail that flags everything scores perfect recall and fails here. Every prompt ID published on HuggingFace is excluded from scoring.
You send: a YAML config naming a HuggingFace model ID, a pinned commit SHA, and how to read a verdict (chat template / decision tokens / classifier threshold).
We run the weights on Zytra hardware. Nothing is transmitted to you.
Best for: single-model guardrail classifiers.
You send an OCI image, content-addressed as either repo@sha256:… (a registry digest) or sha256:… (an image ID). The image ID form means you can ship a tarball via docker save — a proprietary guardrail never has to be published to a registry. A mutable tag is not accepted: it could be repointed after your result is published.
The image reads prompts as JSON lines on stdin and writes verdicts to stdout.
We run it with --network none, verified at runtime. Model weights must be baked into the image.
Best for: cascading guardrails, ensembles, proprietary pipelines.
Not offered by default. Evaluating a hosted API necessarily transmits prompts to the vendor. Contact us to discuss terms; results are published with an "api-exposed" provenance badge rather than "sandboxed".
How submission works
Send your config
Open weights: a YAML file naming the model, a pinned commit SHA, and the decision rule. Container: an image reference pinned by digest. Everything must be pinned — a mutable tag or branch means the scored artifact can change after publication.
We run it
Evaluation happens on Zytra infrastructure against the withheld set. You receive no prompts and run no endpoint. Typical open-weights run: 1–2 minutes of GPU.
You get scored
Overall F1, precision, recall and FPR, plus breakdowns by BFSI category (B-01…B-07), by difficulty, and by conversational register. Published as a leaderboard row with a provenance badge stating how it was produced.
"Every submitter is scored against the same withheld set, and no submitter receives it — including us. That is what makes the numbers comparable."
All tiers include documentation, community support, and regulatory mapping. Tier 4 is used exclusively for official model evaluation submissions.
Ready to Benchmark Your Financial AI?
Validate your guardrail stack against the only BFSI-specific adversarial benchmark. Get started with our free public tier or submit your model for official Tier 4 evaluation.