中文
相关论文

相关论文: Herculean: An Agentic Benchmark for Financial Inte…

200 篇论文

Developing autonomous agents for web-based tasks is a core challenge in AI. While Large Language Model (LLM) agents can interpret complex user requests, they often operate as black boxes, making it difficult to diagnose why they fail or how…

人工智能 · 计算机科学 2026-03-16 Orit Shahnovsky , Rotem Dror

Structured finance, which involves restructuring diverse assets into securities like MBS, ABS, and CDOs, enhances capital market efficiency but presents significant due diligence challenges. This study explores the integration of artificial…

人工智能 · 计算机科学 2024-05-08 Xiangpeng Wan , Haicheng Deng , Kai Zou , Shiqi Xu

Large language and vision-language models increasingly power agents that act on a user's behalf through command-line interface (CLI) harnesses. However, most agent benchmarks still rely on synthetic sandboxes, short-horizon tasks,…

Globally, artificial intelligence (AI) implementation is growing, holding the capability to fundamentally alter organisational processes and decision making. Simultaneously, this brings a multitude of emergent risks to organisations,…

计算机与社会 · 计算机科学 2024-04-10 Finlay McGee

AI-assisted task delegation is increasingly common, yet human effort in such systems is costly and typically unobserved. Recent work by Bastani and Cachon (2025); Sambasivan et al. (2021) shows that accuracy-based payment schemes suffer…

机器学习 · 统计学 2026-03-31 Qichuan Yin , Ziwei Su , Shuangning Li

Deploying machine learning in regulated financial environments -- credit risk, fraud detection, and anti-money laundering -- exposes critical vulnerabilities in algorithmic reproducibility. While early financial ML addressed statistical…

人工智能 · 计算机科学 2026-05-28 Ruizhe Zhou , Xiaoyang Liu , Gaoyuan Du , Yi Zheng , Shouxi Ren , Deepayan Chakrabarti , Dengdu Jiang

Financial intelligence generation from vast data sources has typically relied on traditional methods of knowledge-graph construction or database engineering. Recently, fine-tuned financial domain-specific Large Language Models (LLMs), have…

人工智能 · 计算机科学 2024-10-28 Nicole Cho , Nishan Srishankar , Lucas Cecchi , William Watson

Integration of artificial intelligent (AI) agents in higher education is transforming teaching, learning and administrative processes. Although existing AI agents effectively support individual tasks, their implementation remains fragmented…

Current approaches to AI governance often fall short in anticipating a future where AI agents manage critical tasks, such as financial operations, administrative functions, and beyond. While cryptocurrencies could serve as the foundation…

多智能体系统 · 计算机科学 2025-04-28 Tomer Jordi Chaffer

Artificial intelligence (AI) can undermine financial stability because of malicious use, misinformation, misalignment, and the AI analytics market structure. The low frequency and uniqueness of financial crises, coupled with mutable and…

综合经济学 · 经济学 2024-06-07 Jon Danielsson , Andreas Uthemann

Training GUI agents with traditional centralized methods faces significant cost and scalability challenges. Federated learning (FL) offers a promising solution, yet its potential is hindered by the lack of benchmarks that capture…

多智能体系统 · 计算机科学 2026-04-17 Wenhao Wang , Haoting Shi , Mengying Yuan , Yiquan Lin , Panrong Tong , Hanzhang Zhou , Guangyi Liu , Pengxiang Zhao , Yue Wang , Siheng Chen

Driven by the rapid ascent of artificial intelligence (AI), organizations are at the epicenter of a seismic shift, facing a crucial question: How can AI be successfully integrated into existing operations? To help answer it, manage…

软件工程 · 计算机科学 2024-02-09 Tamen Jadad-Garcia , Alejandro R. Jadad

European financial institutions face mounting regulatory pressure while their security operations centres remain constrained not by data or staffing but by reasoning capacity: enterprise SIEMs cover only a fraction of MITRE ATT&CK…

The field of AI is undergoing a fundamental transition from generative models that can produce synthetic content to artificial agents that can plan and execute complex tasks with only limited human involvement. Companies that pioneered the…

人工智能 · 计算机科学 2025-02-12 Noam Kolt

We introduce the AI Productivity Index for Agents (APEX-Agents), a benchmark for assessing whether AI agents can execute long-horizon, cross-application tasks created by investment banking analysts, management consultants, and corporate…

As AI agents increasingly operate in open, real-world environments, they require a deep synergy of multimodal perception, tool invocation with multi-hop reasoning, and dynamic interaction with users. However, existing benchmarks fail to…

人工智能 · 计算机科学 2026-05-28 Yunqi Liu , Tong Niu , Zitong Wang , Zhenlong Dai , Yuqi Qing , Weiqiang Wang , Jian Liu

The rapid pace of recent research in AI has been driven in part by the presence of fast and challenging simulation environments. These environments often take the form of games; with tasks ranging from simple board games, to competitive…

This position paper argues that agentic AI systems should be designed and evaluated as \emph{marginal token allocation economies} rather than as text generators priced by the unit. We follow a single request -- a developer asking a coding…

人工智能 · 计算机科学 2026-05-05 Siqi Zhu

Although large language models (LLMs) has shown great performance on natural language processing (NLP) in the financial domain, there are no publicly available financial tailtored LLMs, instruction tuning datasets, and evaluation…

计算与语言 · 计算机科学 2023-06-12 Qianqian Xie , Weiguang Han , Xiao Zhang , Yanzhao Lai , Min Peng , Alejandro Lopez-Lira , Jimin Huang

As large language model (LLM) agents increasingly undertake digital work, reliable frameworks are needed to evaluate their real-world competence, adaptability, and capacity for human collaboration. Existing benchmarks remain largely static,…

人工智能 · 计算机科学 2025-12-15 Darvin Yi , Teng Liu , Mattie Terzolo , Lance Hasson , Ayan Sinha , Pablo Mendes , Andrew Rabinovich
‹ 上一页 1 8 9 10 下一页 ›