English
Related papers

Related papers: FinForge: Semi-Synthetic Financial Benchmark Gener…

200 papers

Large language models (LLMs) have become powerful tools for advancing natural language processing applications in the financial industry. However, existing financial LLMs often face challenges such as hallucinations or superficial parameter…

Computation and Language · Computer Science 2024-08-06 Shujuan Zhao , Lingfeng Qiao , Kangyang Luo , Qian-Wen Zhang , Junru Lu , Di Yin

Over the past three years, the financial services industry has witnessed Large Language Models (LLMs) and agents transitioning from the exploration stage to readiness and governance stages. Financial large language models (FinLLMs), such as…

Computational Engineering, Finance, and Science · Computer Science 2026-02-24 Shengyuan Lin , Kaiwen He , Jaisal Patel , Qinchuan Zhang , Chris Ding , James Tang , Keyi Wang , Yupeng Cao , Yan Wang , Kairong Xiao , Vincent Caldeira , Matt White , Xiao-Yang Liu Yanglet

Search has emerged as core infrastructure for LLM-based agents and is widely viewed as critical on the path toward more general intelligence. Finance is a particularly demanding proving ground: analysts routinely conduct complex, multi-step…

As large language models (LLMs) increasingly permeate the financial sector, there is a pressing need for a standardized method to comprehensively assess their performance. Existing financial benchmarks often suffer from limited language and…

Computation and Language · Computer Science 2025-12-09 Xiaojun Wu , Junxi Liu , Huanyi Su , Zhouchi Lin , Yiyan Qi , Chengjin Xu , Jiajun Su , Jiajie Zhong , Fuwei Wang , Saizhuo Wang , Fengrui Hua , Jia Li , Jian Guo

Financial reporting systems increasingly leverage Large Language Models (LLMs) to extract and summarize corporate disclosures. However, most existing approaches assume a single-market setting and overlook structural differences across…

Large language models are increasingly used for financial analysis and investment research, yet systematic evaluation of their financial reasoning capabilities remains limited. In this work, we introduce the AI Financial Intelligence…

Can language models (LMs) self-refine their own responses? This question is increasingly relevant as a wide range of real-world user interactions involve refinement requests. However, prior studies have largely tested LMs' refinement…

Computation and Language · Computer Science 2025-12-01 Young-Jun Lee , Seungone Kim , Byung-Kwan Lee , Minkyeong Moon , Yechan Hwang , Jong Myoung Kim , Graham Neubig , Sean Welleck , Ho-Jin Choi

The rapid advancement of large language models (LLMs) has created a diverse landscape of models, each excelling at different tasks. This diversity drives researchers to employ multiple LLMs in practice, leaving behind valuable multi-LLM log…

Machine Learning · Computer Science 2025-09-30 Tao Feng , Haozhen Zhang , Zijie Lei , Pengrui Han , Mostofa Patwary , Mohammad Shoeybi , Bryan Catanzaro , Jiaxuan You

Answering questions within business and finance requires reasoning, precision, and a wide-breadth of technical knowledge. Together, these requirements make this domain difficult for large language models (LLMs). We introduce BizBench, a…

Computation and Language · Computer Science 2024-03-13 Rik Koncel-Kedziorski , Michael Krumdick , Viet Lai , Varshini Reddy , Charles Lovering , Chris Tanner

In natural language processing (NLP), the focus has shifted from encoder-only tiny language models like BERT to decoder-only large language models(LLMs) such as GPT-3. However, LLMs' practical application in the financial sector has…

Information Retrieval · Computer Science 2025-07-08 Xuan Xu , Fufang Wen , Beilin Chu , Zhibing Fu , Qinhong Lin , Jiaqi Liu , Binjie Fei , Yu Li , Linna Zhou , Zhongliang Yang

Evaluating the factuality of long-form text generated by large language models (LMs) is non-trivial because (1) generations often contain a mixture of supported and unsupported pieces of information, making binary judgments of quality…

Computation and Language · Computer Science 2023-10-12 Sewon Min , Kalpesh Krishna , Xinxi Lyu , Mike Lewis , Wen-tau Yih , Pang Wei Koh , Mohit Iyyer , Luke Zettlemoyer , Hannaneh Hajishirzi

Formal mathematical reasoning remains a critical challenge for artificial intelligence, hindered by limitations of existing benchmarks in scope and scale. To address this, we present FormalMATH, a large-scale Lean4 benchmark comprising…

Existing benchmarks for fake news detection have significantly contributed to the advancement of models in assessing the authenticity of news content. However, these benchmarks typically focus solely on news pertaining to a single semantic…

Computation and Language · Computer Science 2024-10-16 Ziyi Zhou , Xiaoming Zhang , Litian Zhang , Jiacheng Liu , Senzhang Wang , Zheng Liu , Xi Zhang , Chaozhuo Li , Philip S. Yu

Large language models (LLMs) have shown the potential of revolutionizing natural language processing tasks in diverse domains, sparking great interest in finance. Accessing high-quality financial data is the first challenge for financial…

Statistical Finance · Quantitative Finance 2025-11-18 Hongyang Yang , Xiao-Yang Liu , Christina Dan Wang

Large language models (LLMs) have demonstrated remarkable capabilities in code generation across various domains. However, their effectiveness in generating simulation scripts for domain-specific environments like ns-3 remains…

Networking and Internet Architecture · Computer Science 2025-07-16 Tasnim Ahmed , Mirza Mohammad Azwad , Salimur Choudhury

Most reasoning benchmarks for LLMs emphasize factual accuracy or step-by-step logic. In finance, however, professionals must not only converge on optimal decisions but also generate creative, plausible futures under uncertainty. We…

Artificial Intelligence · Computer Science 2025-07-25 Zhuang Qiang Bok , Watson Wei Khong Chua

Modern work relies on an assortment of digital collaboration tools, yet routine processes continue to suffer from human error and delay. To address this gap, this dissertation extends TheAgentCompany with a finance-focused environment and…

Artificial Intelligence · Computer Science 2025-12-03 Rory Milsom

Large Language models (LLMs) have demonstrated significant potential in text-to-SQL reasoning tasks, yet a substantial performance gap persists between existing open-source models and their closed-source counterparts. In this paper, we…

Computation and Language · Computer Science 2025-09-23 Yu Guo , Dong Jin , Shenghao Ye , Shuangwu Chen , Jian Yang , Xiaobin Tan

Financial question answering (QA) over long corporate filings requires evidence to satisfy strict constraints on entities, financial metrics, fiscal periods, and numeric values. However, existing LLM-based rerankers primarily optimize…

Information Retrieval · Computer Science 2026-05-01 Yixi Zhou , Fan Zhang , Yu Chen , Haipeng Zhang , Preslav Nakov , Zhuohan Xie

As financial applications of large language models (LLMs) gain attention, accurate Information Retrieval (IR) remains crucial for reliable AI services. However, existing benchmarks fail to capture the complex and domain-specific information…

Information Retrieval · Computer Science 2025-11-10 Hyunkyu Kim , Yeeun Yoo , Youngjun Kwak