English
Related papers

Related papers: FLAME: Financial Large-Language Model Assessment a…

200 papers

As large language models (LLMs) increasingly permeate the financial sector, there is a pressing need for a standardized method to comprehensively assess their performance. Existing financial benchmarks often suffer from limited language and…

Computation and Language · Computer Science 2025-12-09 Xiaojun Wu , Junxi Liu , Huanyi Su , Zhouchi Lin , Yiyan Qi , Chengjin Xu , Jiajun Su , Jiajie Zhong , Fuwei Wang , Saizhuo Wang , Fengrui Hua , Jia Li , Jian Guo

Recent LLMs have demonstrated promising ability in solving finance related problems. However, applying LLMs in real-world finance application remains challenging due to its high risk and high stakes property. This paper introduces FinTrust,…

Machine Learning · Computer Science 2025-10-20 Tiansheng Hu , Tongyan Hu , Liuyang Bai , Yilun Zhao , Arman Cohan , Chen Zhao

The rapid advancement of large language models presents significant opportunities for financial applications, yet systematic evaluation in specialized financial contexts remains limited. This study presents the first comprehensive…

Computation and Language · Computer Science 2025-09-08 Xuan Yao , Qianteng Wang , Xinbo Liu , Ke-Wei Huang

Artificial intelligence is making significant strides in the finance industry, revolutionizing how data is processed and interpreted. Among these technologies, large language models (LLMs) have demonstrated substantial potential to…

Computation and Language · Computer Science 2024-07-02 Cehao Yang , Chengjin Xu , Yiyan Qi

We introduce Fraud-R1, a benchmark designed to evaluate LLMs' ability to defend against internet fraud and phishing in dynamic, real-world scenarios. Fraud-R1 comprises 8,564 fraud cases sourced from phishing scams, fake job postings,…

Computation and Language · Computer Science 2025-05-27 Shu Yang , Shenzhe Zhu , Zeyu Wu , Keyu Wang , Junchi Yao , Junchao Wu , Lijie Hu , Mengdi Li , Derek F. Wong , Di Wang

This paper introduces the UCFE: User-Centric Financial Expertise benchmark, an innovative framework designed to evaluate the ability of large language models (LLMs) to handle complex real-world financial tasks. UCFE benchmark adopts a…

Computational Finance · Quantitative Finance 2025-02-10 Yuzhe Yang , Yifei Zhang , Yan Hu , Yilin Guo , Ruoli Gan , Yueru He , Mingcong Lei , Xiao Zhang , Haining Wang , Qianqian Xie , Jimin Huang , Honghai Yu , Benyou Wang

Financial large language models (FinLLMs) have been applied to various tasks in business, finance, accounting, and auditing. Complex financial regulations and standards are critical to financial services, which LLMs must comply with.…

Computational Engineering, Finance, and Science · Computer Science 2025-01-14 Keyi Wang , Jaisal Patel , Charlie Shen , Daniel Kim , Andy Zhu , Alex Lin , Luca Borella , Cailean Osborne , Matt White , Steve Yang , Kairong Xiao , Xiao-Yang Liu Yanglet

The emergence of Large Vision-Language Models (LVLMs) has substantially expanded model capabilities beyond text-only understanding, enabling unified inference across both visual and textual modalities and supporting a broader range of…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Qian Chen , Xianyin Zhang , Yanzhi Liu , Lifan Guo , Feng Chen , Chi Zhang

Large language models (LLMs) have demonstrated strong capabilities in language understanding, generation, and reasoning, yet their potential in finance remains underexplored due to the complexity and specialization of financial knowledge.…

Computation and Language · Computer Science 2025-01-03 Hanyu Zhang , Boyu Qiu , Yuhao Feng , Shuqi Li , Qian Ma , Xiyuan Zhang , Qiang Ju , Dong Yan , Jian Xie

As large language models (LLMs) advance, it becomes more challenging to reliably evaluate their output due to the high costs of human evaluation. To make progress towards better LLM autoraters, we introduce FLAMe, a family of Foundational…

Computation and Language · Computer Science 2024-07-16 Tu Vu , Kalpesh Krishna , Salaheddin Alzubi , Chris Tar , Manaal Faruqui , Yun-Hsuan Sung

The emergence of Large Language Models (LLMs), such as ChatGPT, has revolutionized general natural language preprocessing (NLP) tasks. However, their expertise in the financial domain lacks a comprehensive evaluation. To assess the ability…

Computation and Language · Computer Science 2023-10-20 Yue Guo , Zian Xu , Yi Yang

Jailbreaking poses a significant risk to the deployment of Large Language Models (LLMs) and Vision Language Models (VLMs). VLMs are particularly vulnerable because they process both text and images, creating broader attack surfaces.…

Computation and Language · Computer Science 2026-02-23 Mirae Kim , Seonghun Jeong , Youngjun Kwak

Language Models (LMs) struggle with complex, interdependent instructions, particularly in high-stakes domains like finance where precision is critical. We introduce FIFE, a novel, high-difficulty benchmark designed to assess LM…

Machine Learning · Computer Science 2025-12-11 Glenn Matlin , Siddharth , Anirudh JM , Aditya Shukla , Yahya Hassan , Sudheer Chava

FinanceBench is a first-of-its-kind test suite for evaluating the performance of LLMs on open book financial question answering (QA). It comprises 10,231 questions about publicly traded companies, with corresponding answers and evidence…

Computation and Language · Computer Science 2023-11-21 Pranab Islam , Anand Kannappan , Douwe Kiela , Rebecca Qian , Nino Scherrer , Bertie Vidgen

The financial industry's growing demand for advanced natural language processing (NLP) capabilities has highlighted the limitations of generalist large language models (LLMs) in handling domain-specific financial tasks. To address this gap,…

Statistical Finance · Quantitative Finance 2025-11-13 Gaëtan Caillaut , Raheel Qader , Jingshu Liu , Mariam Nakhlé , Arezki Sadoune , Massinissa Ahmim , Jean-Gabriel Barthelemy

We introduce FinanceReasoning, a novel benchmark designed to evaluate the reasoning capabilities of large reasoning models (LRMs) in financial numerical reasoning problems. Compared to existing benchmarks, our work provides three key…

Computation and Language · Computer Science 2025-08-07 Zichen Tang , Haihong E , Ziyan Ma , Haoyang He , Jiacheng Liu , Zhongjun Yang , Zihua Rong , Rongjin Li , Kun Ji , Qing Huang , Xinyang Hu , Yang Liu , Qianhe Zheng

In this paper, we introduce FAMMA, an open-source benchmark for \underline{f}in\underline{a}ncial \underline{m}ultilingual \underline{m}ultimodal question \underline{a}nswering (QA). Our benchmark aims to evaluate the abilities of large…

Computation and Language · Computer Science 2025-05-16 Siqiao Xue , Xiaojing Li , Fan Zhou , Qingyang Dai , Zhixuan Chu , Hongyuan Mei

Large Language Models (LLMs) are increasingly integrated into financial workflows, but evaluation practice has not kept up. Finance-specific biases can inflate performance, contaminate backtests, and make reported results useless for any…

Although large language models (LLMs) has shown great performance on natural language processing (NLP) in the financial domain, there are no publicly available financial tailtored LLMs, instruction tuning datasets, and evaluation…

Computation and Language · Computer Science 2023-06-12 Qianqian Xie , Weiguang Han , Xiao Zhang , Yanzhao Lai , Min Peng , Alejandro Lopez-Lira , Jimin Huang

Pre-trained language models have shown impressive performance on a variety of tasks and domains. Previous research on financial language models usually employs a generic training scheme to train standard model architectures, without…

Computation and Language · Computer Science 2022-11-02 Raj Sanjay Shah , Kunal Chawla , Dheeraj Eidnani , Agam Shah , Wendi Du , Sudheer Chava , Natraj Raman , Charese Smiley , Jiaao Chen , Diyi Yang