FraudBench Benchmarks Adversarial Robustness in Financial Risk Assessment

Xitong Zeng, Zhaoge Bi, Yitian Yang, Huaming Chen, Quan Z. Sheng· August 26, 2026 View original

Key takeaways

  • Adversarial robustness in finance is highly dependent on the evaluation protocol.
  • FraudBench provides a protocol-sensitive benchmark for financial risk assessment.
  • Integrating domain constraints into attack generation is crucial for realistic evaluation.
  • Robustness evaluations should jointly report predictive degradation and attack feasibility.

Who benefits

BFSIFintechRisk ManagementCybersecurity

Summary

This paper introduces FraudBench, a protocol-sensitive benchmark for evaluating the adversarial robustness of machine learning models in financial fraud and credit-risk detection. It demonstrates that robustness conclusions are highly dependent on how domain-specific constraints and attacker capabilities are incorporated into the evaluation protocol.

This research highlights a critical gap in evaluating the adversarial robustness of machine learning models used in financial fraud and credit-risk detection. Traditional robustness assessments often fail to account for the unique characteristics of financial tabular data, such as domain-specific constraints, severe class imbalance, and asymmetric attacker capabilities. The paper argues that robustness is not solely a model attribute but also heavily influenced by the evaluation protocol itself. To address this, the authors present FraudBench, a novel benchmark designed for protocol-sensitive adversarial robustness evaluation. FraudBench assesses models under three distinct protocols: unconstrained attacks, post-hoc feasibility filtering, and deployment-aware constraint-integrated attacks. Experiments across four public financial datasets and various model types (neural, tree-based, ensemble) reveal that robustness conclusions vary significantly based on the protocol. For instance, integrating constraints directly into attack generation yields vastly more feasible adversarial examples than post-hoc filtering, underscoring the necessity of protocol-aware evaluation for accurate risk assessment.

Why it matters

Professionals in financial services and risk management must adopt protocol-sensitive benchmarking to accurately assess the adversarial robustness of their ML models, ensuring more reliable fraud detection and credit risk assessment against sophisticated attacks.

How to implement this in your domain

  1. 1Review current adversarial robustness testing methodologies for financial ML models.
  2. 2Adopt a protocol-sensitive approach to benchmarking, considering domain constraints and attacker capabilities.
  3. 3Implement constraint-integrated attack generation rather than relying solely on post-hoc filtering for adversarial examples.
  4. 4Evaluate model robustness across different attack settings (white-box, black-box) and model families.
  5. 5Report both predictive degradation and attack feasibility jointly to provide a comprehensive view of robustness.

Original post by Xitong Zeng, Zhaoge Bi, Yitian Yang, Huaming Chen, Quan Z. Sheng

"arXiv:2608.24551v1 Announce Type: new Abstract: Machine learning models are widely used in financial fraud and credit-risk detection, yet their adversarial robustness remains difficult to evaluate because financial tabular data involve domain-specific constraints, severe class im…"

View on X

Originally posted by Xitong Zeng, Zhaoge Bi, Yitian Yang, Huaming Chen, Quan Z. Sheng on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses

More in AI Research

AI ResearchAI Engineering & DevTools

Persistent Cross Entropy Extends Topological Data Analysis

This paper introduces Persistent Cross Entropy (PCE), a novel extension of cross-entropy to persistence diagrams, which are used in topological data analysis. PCE bridges different event spaces of diagrams using an induced probability, enabling new applications like distinguishing diagrams with similar persistent entropy and separating causal directions in dynamical systems.

Sijin Yeom, Jae-Hun JungAug 26, 2026
AI ResearchAI Engineering & DevTools

Bridging Numerical PDE Solvers and Neural Emulators for Faster Simulation

This thesis explores the deep connections between traditional numerical solvers for Partial Differential Equations (PDEs) and neural emulators, arguing that they are more alike than different. It proposes that insights can flow profitably in both directions, leading to faster and more efficient scientific and engineering simulations.

Felix KoehlerAug 26, 2026
AI ResearchAI Engineering & DevTools

New UQ Method Integrates Joint Aleatoric and Epistemic Uncertainties

This paper introduces a novel approach for Uncertainty Quantification (UQ) that jointly integrates aleatoric (data noise) and epistemic (model confidence) uncertainties in high-dimensional regression tasks. The method uses a low-rank plus diagonal covariance structure to capture essential output correlations efficiently, leading to more reliable deep learning predictions.

Leonhard F. Feiner, Manuel Nickel, Martin Menten, Laurin Lux, Rickmer Braren, Daniel Rueckert, Georgios Kaissis, Raphael Rehms, Johannes PaetzoldAug 26, 2026