English
Related papers

Related papers: Bench to the Future: A Pastcasting Benchmark for F…

200 papers

Forecasting accuracy is reliant on the quality of available past data. Data disruptions can adversely affect the quality of the generated model (e.g. unexpected events such as out-of-stock products when forecasting demand). We address this…

Machine Learning · Computer Science 2021-06-29 André Baptista , Yassine Baghoussi , Carlos Soares , João Mendes-Moreira , Miguel Arantes

Benchmarking is crucial for testing and validating any system, even more so in real-time systems. Typical real-time applications adhere to well-understood abstractions: they exhibit a periodic behavior, operate on a well-defined working…

Software Engineering · Computer Science 2022-08-02 Mattia Nicolella , Shahin Roozkhosh , Denis Hoornaert , Andrea Bastoni , Renato Mancuso

Generative AI, particularly large language models (LLMs), is beginning to transform the financial industry by automating tasks and helping to make sense of complex financial information. One especially promising use case is the automatic…

Statistical Finance · Quantitative Finance 2025-11-11 Zonghan Wu , Congyuan Zou , Junlin Wang , Chenhan Wang , Hangjing Yang , Yilei Shao

While agent evaluation has shifted toward long-horizon tasks, most benchmarks still emphasize local, step-level reasoning rather than the global constrained optimization (e.g., time and financial budgets) that demands genuine planning…

Artificial Intelligence · Computer Science 2026-01-27 Yinger Zhang , Shutong Jiang , Renhao Li , Jianhong Tu , Yang Su , Lianghao Deng , Xudong Guo , Chenxu Lv , Junyang Lin

Scaling up data, parameters, and test-time computation has been the mainstream methods to improve LLM systems (LLMsys), but their upper bounds are almost reached due to the gradual depletion of high-quality data and marginal gains obtained…

Machine Learning · Computer Science 2026-05-12 Qingyao Ai , Yichen Tang , Changyue Wang , Jianming Long , Weihang Su , Yiqun Liu

Many existing evaluation benchmarks for Large Language Models (LLMs) quickly become outdated due to the emergence of new models and training data. These benchmarks also fall short in assessing how LLM performance changes over time, as they…

Computation and Language · Computer Science 2025-07-09 Hui Dai , Ryan Teehan , Mengye Ren

A tool that could suggest new personalized research directions and ideas by taking insights from the scientific literature could significantly accelerate the progress of science. A field that might benefit from such an approach is…

Modern AI progress has been driven by ML methods that are generalizable across settings and scalable to larger regimes. As large language models demonstrate advanced capabilities in reasoning, coding, and engineering tasks, it is…

Evaluating the pedagogical capabilities of AI-based tutoring models is critical for making guided progress in the field. Yet, we lack a reliable, easy-to-use, and simple-to-run evaluation that reflects the pedagogical abilities of models.…

Computation and Language · Computer Science 2025-10-14 Jakub Macina , Nico Daheim , Ido Hakimi , Manu Kapur , Iryna Gurevych , Mrinmaya Sachan

Forecasting has always been at the forefront of decision making and planning. The uncertainty that surrounds the future is both exciting and challenging, with individuals and organisations seeking to minimise risks and maximise utilities.…

Applications · Statistics 2022-02-09 Fotios Petropoulos , Daniele Apiletti , Vassilios Assimakopoulos , Mohamed Zied Babai , Devon K. Barrow , Souhaib Ben Taieb , Christoph Bergmeir , Ricardo J. Bessa , Jakub Bijak , John E. Boylan , Jethro Browell , Claudio Carnevale , Jennifer L. Castle , Pasquale Cirillo , Michael P. Clements , Clara Cordeiro , Fernando Luiz Cyrino Oliveira , Shari De Baets , Alexander Dokumentov , Joanne Ellison , Piotr Fiszeder , Philip Hans Franses , David T. Frazier , Michael Gilliland , M. Sinan Gönül , Paul Goodwin , Luigi Grossi , Yael Grushka-Cockayne , Mariangela Guidolin , Massimo Guidolin , Ulrich Gunter , Xiaojia Guo , Renato Guseo , Nigel Harvey , David F. Hendry , Ross Hollyman , Tim Januschowski , Jooyoung Jeon , Victor Richmond R. Jose , Yanfei Kang , Anne B. Koehler , Stephan Kolassa , Nikolaos Kourentzes , Sonia Leva , Feng Li , Konstantia Litsiou , Spyros Makridakis , Gael M. Martin , Andrew B. Martinez , Sheik Meeran , Theodore Modis , Konstantinos Nikolopoulos , Dilek Önkal , Alessia Paccagnini , Anastasios Panagiotelis , Ioannis Panapakidis , Jose M. Pavía , Manuela Pedio , Diego J. Pedregal , Pierre Pinson , Patrícia Ramos , David E. Rapach , J. James Reade , Bahman Rostami-Tabar , Michał Rubaszek , Georgios Sermpinis , Han Lin Shang , Evangelos Spiliotis , Aris A. Syntetos , Priyanga Dilini Talagala , Thiyanga S. Talagala , Len Tashman , Dimitrios Thomakos , Thordis Thorarinsdottir , Ezio Todini , Juan Ramón Trapero Arenas , Xiaoqian Wang , Robert L. Winkler , Alisa Yusupova , Florian Ziel

Recent works have shown that large language model (LLM) agents are able to improve themselves from experience, which is an important ability for continuous enhancement post-deployment. However, existing benchmarks primarily evaluate their…

Computation and Language · Computer Science 2024-11-01 Cheng-Kuang Wu , Zhi Rui Tam , Chieh-Yen Lin , Yun-Nung Chen , Hung-yi Lee

Data governance ensures data quality, security, and compliance through policies and standards, a critical foundation for scaling modern AI development. Recently, large language models (LLMs) have emerged as a promising solution for…

Artificial Intelligence · Computer Science 2025-12-09 Zhou Liu , Zhaoyang Han , Guochen Yan , Hao Liang , Bohan Zeng , Xing Chen , Yuanfeng Song , Wentao Zhang

Modern work relies on an assortment of digital collaboration tools, yet routine processes continue to suffer from human error and delay. To address this gap, this dissertation extends TheAgentCompany with a finance-focused environment and…

Artificial Intelligence · Computer Science 2025-12-03 Rory Milsom

With the advancement of web techniques, they have significantly revolutionized various aspects of people's lives. Despite the importance of the web, many tasks performed on it are repetitive and time-consuming, negatively impacting overall…

Artificial Intelligence · Computer Science 2025-08-06 Liangbo Ning , Ziran Liang , Zhuohang Jiang , Haohao Qu , Yujuan Ding , Wenqi Fan , Xiao-yong Wei , Shanru Lin , Hui Liu , Philip S. Yu , Qing Li

Accurate prediction of climate in the subseasonal-to-seasonal scale is crucial for disaster preparedness and robust decision making amidst climate change. Yet, forecasting beyond the weather timescale is challenging because it deals with…

Computer Vision and Pattern Recognition · Computer Science 2024-11-22 Juan Nathaniel , Yongquan Qu , Tung Nguyen , Sungduk Yu , Julius Busecke , Aditya Grover , Pierre Gentine

Risk assessments for a pediatric population are often conducted across multiple stages. For example, clinicians may evaluate risks prenatally, at birth, and during Well-Child visits. Although predictions made at later stages typically…

Machine Learning · Computer Science 2025-08-18 Minghui Sun , Matthew M. Engelhard , Benjamin A. Goldstein

Who is the US President? The answer changes depending on when the question is asked. While large language models (LLMs) are evaluated on various reasoning tasks, they often miss a crucial dimension: time. In real-world scenarios, the…

Computation and Language · Computer Science 2025-05-16 David Herel , Vojtech Bartek , Jiri Jirak , Tomas Mikolov

LLM-based agents are increasingly expected to handle real-world assistant tasks, yet existing benchmarks typically evaluate them under isolated sources of difficulty, such as a single environment or fully specified instructions. This leaves…

Computation and Language · Computer Science 2026-04-16 Xiang Long , Li Du , Yilong Xu , Fangcheng Liu , Haoqing Wang , Ning Ding , Ziheng Li , Jianyuan Guo , Yehui Tang

While machine learning has witnessed significant advancements, the emphasis has largely been on data acquisition and model creation. However, achieving a comprehensive assessment of machine learning solutions in real-world settings…