English

Benchmarking Generalization in Financial Statement Fraud Detection: robust evaluation and novel tasks

Machine Learning 2026-07-21 v1 Artificial Intelligence

Abstract

Financial statement fraud detection (FSFD) is crucial for market integrity but faces challenges from increasingly sophisticated schemes and under-utilized textual data in financial reports. Existing methods often rely on random data splits, leading to overoptimistic performance estimates that do not reflect real-world generalization to new companies or future periods. To address this recurring problem with the state of the art, we propose a robust FSFD framework leveraging Large Language Models (LLMs) to integrate both structured financial data and unstructured textual information from financial reports. We provide a more realistic evaluation through a novel and challenging benchmark task called Company-Isolated FSFD (CI-FSFD). We construct and make publicly available a comprehensive U.S. company dataset combining financial statements, summarized MD&A text, and fraud labels. Our approach achieves the best performance on the challenging CI-FSFD task, demonstrating the critical value of textual data and robust evaluation for reliable financial fraud detection.

Cite

@article{arxiv.2607.19259,
  title  = {Benchmarking Generalization in Financial Statement Fraud Detection: robust evaluation and novel tasks},
  author = {Guy Stephane Waffo Dzuyo and Gaël Guibon and Christophe Cerisara and Luis Belmar-Letelier},
  journal= {arXiv preprint arXiv:2607.19259},
  year   = {2026}
}

Comments

Accepted at FinLLM@IJCAI2026

R2 v1 2026-07-22T20:50:15.812Z