中文
相关论文

相关论文: Finance Agent Benchmark: Benchmarking LLMs on Real…

200 篇论文

Although large language models (LLMs) has shown great performance on natural language processing (NLP) in the financial domain, there are no publicly available financial tailtored LLMs, instruction tuning datasets, and evaluation…

计算与语言 · 计算机科学 2023-06-12 Qianqian Xie , Weiguang Han , Xiao Zhang , Yanzhao Lai , Min Peng , Alejandro Lopez-Lira , Jimin Huang

Environmental, social, and governance (ESG) criteria are essential for evaluating corporate sustainability and ethical performance. However, professional ESG analysis is hindered by data fragmentation across unstructured sources, and…

人工智能 · 计算机科学 2026-01-15 Yilei Zhao , Wentao Zhang , Lei Xiao , Yandan Zheng , Mengpu Liu , Wei Yang Bryan Lim

AI research agents accelerate ML research by automating hypothesis generation, experimentation, and empirical refinement. Existing agent strategies range from greedy hill-climbing to tree search and evolutionary optimization, yet which…

Large Language Models (LLMs) and Large Vision-Language Models (LVLMs) have demonstrated impressive language/vision reasoning abilities, igniting the recent trend of building agents for targeted applications such as shopping assistants or AI…

人工智能 · 计算机科学 2025-04-14 Liqiang Jing , Zhehui Huang , Xiaoyang Wang , Wenlin Yao , Wenhao Yu , Kaixin Ma , Hongming Zhang , Xinya Du , Dong Yu

Financial tasks are pivotal to global economic stability; however, their execution faces challenges including labor intensive processes, low error tolerance, data fragmentation, and tool limitations. Although large language models (LLMs)…

人工智能 · 计算机科学 2025-05-21 Junzhe Jiang , Chang Yang , Aixin Cui , Sihan Jin , Ruiyu Wang , Bo Li , Xiao Huang , Dongning Sun , Xinrun Wang

The emergence of agentic recommender systems powered by Large Language Models (LLMs) represents a paradigm shift in personalized recommendations, leveraging LLMs' advanced reasoning and role-playing capabilities to enable autonomous,…

信息检索 · 计算机科学 2025-05-29 Yu Shang , Peijie Liu , Yuwei Yan , Zijing Wu , Leheng Sheng , Yuanqing Yu , Chumeng Jiang , An Zhang , Fengli Xu , Yu Wang , Min Zhang , Yong Li

As benchmarks grow in complexity, many apparent agent failures are not failures of the agent at all - they are failures of the benchmark itself: broken specifications, implicit assumptions, and rigid evaluation scripts that penalize valid…

计算与语言 · 计算机科学 2026-04-29 Xinming Tu , Tianze Wang , Yingzhou , Lu , Kexin Huang , Yuanhao Qu , Sara Mostafavi

Evaluating AI agents in finance faces two key challenges: static benchmarks require costly expert annotation yet miss the dynamic decision-making central to real-world trading, while LLM-based judges introduce uncontrolled variance on…

人工智能 · 计算机科学 2026-03-03 Xiaochuang Yuan , Hui Xu , Silvia Xu , Cui Zou , Jing Xiong

As language models (LMs) evolve from chat assistants to long-horizon agents capable of multi-step reasoning and tool use, existing benchmarks remain largely confined to structured or exam-style tasks that fall short of real-world…

The financial domain poses substantial challenges for vision-language models (VLMs) due to specialized chart formats and knowledge-intensive reasoning requirements. However, existing financial benchmarks are largely single-turn and rely on…

计算机视觉与模式识别 · 计算机科学 2026-02-04 Chenxi Zhang , Ziliang Gan , Liyun Zhu , Youwei Pang , Qing Zhang , Rongjunchen Zhang

LLM agents are increasingly expected to carry out end-to-end workflows, producing complete artifacts from high-level user instructions. To meet enterprise needs, frontier AI labs have developed agents that can construct entire spreadsheets…

The advancement of large language model (LLM) based agents has shifted AI evaluation from single-turn response assessment to multi-step task completion in interactive environments. We present an empirical study evaluating frontier AI models…

人工智能 · 计算机科学 2026-01-15 Logan Ritchie , Sushant Mehta , Nick Heiner , Mason Yu , Edwin Chen

The current paper presents the development and validation of SelfScore, a novel benchmark designed to assess the performance of automated Large Language Model (LLM) agents on help desk and professional consultation tasks. Given the…

计算机与社会 · 计算机科学 2024-10-23 John Mavi , Nathan Summers , Sergio Coronado

The literature has witnessed an emerging interest in AI agents for automated assessment of scientific papers. Existing benchmarks focus primarily on the computational aspect of this task, testing agents' ability to reproduce or replicate…

Trust in AI agents has been extensively studied in the literature, resulting in significant advancements in our understanding of this field. However, the rapid advancements in Large Language Models (LLMs) and the emergence of LLM-based AI…

人工智能 · 计算机科学 2023-08-11 Sivan Schwartz , Avi Yaeli , Segev Shlomov

Large Language Models (LLMs) have demonstrated impressive capabilities across various domains, but their effectiveness in financial decision-making remains inadequately evaluated. Current benchmarks primarily assess LLMs' understanding on…

多智能体系统 · 计算机科学 2025-06-27 Changlun Li , Yao Shi , Yuyu Luo , Nan Tang

The advances made by Large Language Models (LLMs) have led to the pursuit of LLM agents that can solve intricate, multi-step reasoning tasks. As with any research pursuit, benchmarking and evaluation are key corner stones to efficient and…

The rapid advancements in Large Language Models (LLMs) have unlocked transformative possibilities in natural language processing, particularly within the financial sector. Financial data is often embedded in intricate relationships across…

统计金融 · 定量金融 2026-05-21 Alejandro Lopez-Lira , Jihoon Kwon , Sangwoon Yoon , Jy-yong Sohn , Chanyeol Choi

Artificial intelligence (AI) is the core technology of technological revolution and industrial transformation. As one of the new intelligent needs in the AI 2.0 era, financial intelligence has elicited much attention from the academia and…

人工智能 · 计算机科学 2018-08-28 Xiaolin Zheng , Mengying Zhu , Qibing Li , Chaochao Chen , Yanchao Tan

Current financial large language models (FinLLMs) struggle with two critical limitations: the absence of objective evaluation metrics to assess the quality of stock analysis reports and a lack of depth in stock analysis, which impedes their…

人工智能 · 计算机科学 2025-07-10 Shijie Han , Jingshu Zhang , Yiqing Shen , Kaiyuan Yan , Hongguang Li