English

FinDABench: Benchmarking Financial Data Analysis Ability of Large Language Models

Computation and Language 2024-06-17 v4 Artificial Intelligence

Abstract

Large Language Models (LLMs) have demonstrated impressive capabilities across a wide range of tasks. However, their proficiency and reliability in the specialized domain of financial data analysis, particularly focusing on data-driven thinking, remain uncertain. To bridge this gap, we introduce \texttt{FinDABench}, a comprehensive benchmark designed to evaluate the financial data analysis capabilities of LLMs within this context. \texttt{FinDABench} assesses LLMs across three dimensions: 1) \textbf{Foundational Ability}, evaluating the models' ability to perform financial numerical calculation and corporate sentiment risk assessment; 2) \textbf{Reasoning Ability}, determining the models' ability to quickly comprehend textual information and analyze abnormal financial reports; and 3) \textbf{Technical Skill}, examining the models' use of technical knowledge to address real-world data analysis challenges involving analysis generation and charts visualization from multiple perspectives. We will release \texttt{FinDABench}, and the evaluation scripts at \url{https://github.com/cubenlp/BIBench}. \texttt{FinDABench} aims to provide a measure for in-depth analysis of LLM abilities and foster the advancement of LLMs in the field of financial data analysis.

Keywords

Cite

@article{arxiv.2401.02982,
  title  = {FinDABench: Benchmarking Financial Data Analysis Ability of Large Language Models},
  author = {Shu Liu and Shangqing Zhao and Chenghao Jia and Xinlin Zhuang and Zhaoguang Long and Jie Zhou and Aimin Zhou and Man Lan and Qingquan Wu and Chong Yang},
  journal= {arXiv preprint arXiv:2401.02982},
  year   = {2024}
}
R2 v1 2026-06-28T14:09:47.119Z