English
Related papers

Related papers: FOReCAst: The Future Outcome Reasoning and Confide…

200 papers

Forecasts of future events are essential inputs into informed decision-making. Machine learning (ML) systems have the potential to deliver forecasts at scale, but there is no framework for evaluating the accuracy of ML systems on a…

Machine Learning · Computer Science 2025-03-03 Ezra Karger , Houtan Bastani , Chen Yueh-Han , Zachary Jacobs , Danny Halawi , Fred Zhang , Philip E. Tetlock

Forecasting is a challenging task that offers a clearly measurable way to study AI systems. Forecasting requires a large amount of research on the internet, and evaluations require time for events to happen, making the development of…

Computation and Language · Computer Science 2025-06-30 FutureSearch , : , Jack Wildman , Nikos I. Bosse , Daniel Hnyk , Peter Mühlbacher , Finn Hambly , Jon Evans , Dan Schwarz , Lawrence Phillips

Without the ability to estimate and benchmark AI capability advancements, organizations are left to respond to each change reactively, impeding their ability to build viable mid and long-term strategies. This paper explores the recent…

Computers and Society · Computer Science 2023-04-03 Emily Dardaman , Abhishek Gupta

Forecasting has always been at the forefront of decision making and planning. The uncertainty that surrounds the future is both exciting and challenging, with individuals and organisations seeking to minimise risks and maximise utilities.…

Applications · Statistics 2022-02-09 Fotios Petropoulos , Daniele Apiletti , Vassilios Assimakopoulos , Mohamed Zied Babai , Devon K. Barrow , Souhaib Ben Taieb , Christoph Bergmeir , Ricardo J. Bessa , Jakub Bijak , John E. Boylan , Jethro Browell , Claudio Carnevale , Jennifer L. Castle , Pasquale Cirillo , Michael P. Clements , Clara Cordeiro , Fernando Luiz Cyrino Oliveira , Shari De Baets , Alexander Dokumentov , Joanne Ellison , Piotr Fiszeder , Philip Hans Franses , David T. Frazier , Michael Gilliland , M. Sinan Gönül , Paul Goodwin , Luigi Grossi , Yael Grushka-Cockayne , Mariangela Guidolin , Massimo Guidolin , Ulrich Gunter , Xiaojia Guo , Renato Guseo , Nigel Harvey , David F. Hendry , Ross Hollyman , Tim Januschowski , Jooyoung Jeon , Victor Richmond R. Jose , Yanfei Kang , Anne B. Koehler , Stephan Kolassa , Nikolaos Kourentzes , Sonia Leva , Feng Li , Konstantia Litsiou , Spyros Makridakis , Gael M. Martin , Andrew B. Martinez , Sheik Meeran , Theodore Modis , Konstantinos Nikolopoulos , Dilek Önkal , Alessia Paccagnini , Anastasios Panagiotelis , Ioannis Panapakidis , Jose M. Pavía , Manuela Pedio , Diego J. Pedregal , Pierre Pinson , Patrícia Ramos , David E. Rapach , J. James Reade , Bahman Rostami-Tabar , Michał Rubaszek , Georgios Sermpinis , Han Lin Shang , Evangelos Spiliotis , Aris A. Syntetos , Priyanga Dilini Talagala , Thiyanga S. Talagala , Len Tashman , Dimitrios Thomakos , Thordis Thorarinsdottir , Ezio Todini , Juan Ramón Trapero Arenas , Xiaoqian Wang , Robert L. Winkler , Alisa Yusupova , Florian Ziel

High-stakes decision making involves reasoning under uncertainty about the future. In this work, we train language models to make predictions on open-ended forecasting questions. To scale up training data, we synthesize novel forecasting…

Machine Learning · Computer Science 2026-01-06 Nikhil Chandak , Shashwat Goel , Ameya Prabhu , Moritz Hardt , Jonas Geiping

Event forecasting is a challenging, yet important task, as humans seek to constantly plan for the future. Existing automated forecasting studies rely mostly on structured data, such as time-series or event-based knowledge graphs, to help…

Machine Learning · Computer Science 2021-06-09 Woojeong Jin , Rahul Khanna , Suji Kim , Dong-Ho Lee , Fred Morstatter , Aram Galstyan , Xiang Ren

Time series forecasts are widely used to inform decisions. Human decision-makers interpret these forecasts, incorporate prior experience and uncertainty about future outcomes, and then make a decision. In this paper, we propose a new…

Machine Learning · Statistics 2026-05-01 Daniel Andrew Coulson , Martin T. Wells

Real-world settings where language models (LMs) are deployed -- in domains spanning healthcare, finance, and other forms of knowledge work -- require models to grapple with incomplete information and reason under uncertainty. Yet most LM…

Artificial Intelligence · Computer Science 2026-04-24 Alana Renda , Jillian Ross , Michael Cafarella , Jacob Andreas

While pre-trained language models achieve impressive performance on various NLP benchmarks, they still struggle with tasks that require numerical reasoning. Recent advances in improving numerical reasoning are mostly achieved using very…

Computation and Language · Computer Science 2023-05-30 Jasivan Alex Sivakumar , Nafise Sadat Moosavi

People vary in their ability to make accurate predictions about the future. Prior studies have shown that some individuals can predict the outcome of future events with consistently better accuracy. This leads to a natural question: what…

Computation and Language · Computer Science 2020-06-17 Shi Zong , Alan Ritter , Eduard Hovy

Forecasting has become a natural benchmark for reasoning under uncertainty. Yet existing evaluations of large language models remain limited to judgmental tasks in simple formats, such as binary or multiple-choice questions. In practice,…

Machine Learning · Computer Science 2026-04-20 Jeremy Qin , Maksym Andriushchenko

Predicting future events is an important activity with applications across multiple fields and domains. For example, the capacity to foresee stock market trends, natural disasters, business developments, or political events can facilitate…

Computation and Language · Computer Science 2025-01-13 Petraq Nako , Adam Jatowt

Prior work has largely treated forecasting as a static task, failing to consider how forecasts and the confidence in them should evolve as new evidence emerges. To address this gap, we introduce EvolveCast, a framework for evaluating…

Computation and Language · Computer Science 2026-02-10 Zhangdie Yuan , Zifeng Ding , Andreas Vlachos

We introduce TFRBench, the first benchmark designed to evaluate the reasoning capabilities of forecasting systems. Traditionally, time-series forecasting has been evaluated solely on numerical accuracy, treating foundation models as ``black…

Artificial Intelligence · Computer Science 2026-04-08 Md Atik Ahamed , Mihir Parmar , Palash Goyal , Yiwen Song , Long T. Le , Qiang Cheng , Chun-Liang Li , Hamid Palangi , Jinsung Yoon , Tomas Pfister

For users to trust model predictions, they need to understand model outputs, particularly their confidence - calibration aims to adjust (calibrate) models' confidence to match expected accuracy. We argue that the traditional calibration…

Computation and Language · Computer Science 2022-10-25 Chenglei Si , Chen Zhao , Sewon Min , Jordan Boyd-Graber

Physical reasoning requires forward prediction: the ability to forecast what will happen next given some initial world state. We study the performance of state-of-the-art forward-prediction models in the complex physical-reasoning tasks of…

Machine Learning · Computer Science 2021-03-31 Rohit Girdhar , Laura Gustafson , Aaron Adcock , Laurens van der Maaten

Forecasting is a critical task in decision-making across numerous domains. While historical numerical data provide a start, they fail to convey the complete context for reliable and accurate predictions. Human forecasters frequently rely on…

Advances in deep learning systems have allowed large models to match or surpass human accuracy on a number of skills such as image classification, basic programming, and standardized test taking. As the performance of the most capable…

Machine Learning · Computer Science 2024-06-10 Sarah Pratt , Seth Blumberg , Pietro Kreitlon Carolino , Meredith Ringel Morris

Language models (LMs) trained on web-scale datasets are largely successful due to their ability to memorize large amounts of training data, even if only present in a few examples. These capabilities are often desirable in evaluation on…

Machine Learning · Computer Science 2024-11-04 Elvis Hsieh , Preston Fu , Jonathan Chen

Confidence estimation (CE) indicates how reliable the answers of large language models are and impacts user trust and decision-making. Existing evaluations mainly concern the alignment between confidence and correctness, but ignore the…

Computation and Language · Computer Science 2026-05-29 Yuxi Xia , Dennis Ulmer , Terra Blevins , Yihong Liu , Hinrich Schütze , Benjamin Roth
‹ Prev 1 2 3 10 Next ›