English
Related papers

Related papers: Evaluating Strategic Reasoning in Forecasting Agen…

200 papers

Many studies have shown that humans are "predictably irrational": they do not act in a fully rational way, but their deviations from rational behavior are quite systematic. Our goal is to see the extent to which we can explain and justify…

Computer Science and Game Theory · Computer Science 2023-07-27 Xinming Liu , Joseph Y. Halpern

Time series forecasting (TSF) is a fundamental and widely studied task, spanning methods from classical statistical approaches to modern deep learning and multimodal language modeling. Despite their effectiveness, these methods often follow…

Machine Learning · Computer Science 2025-12-23 Mingyue Cheng , Jiahao Wang , Daoyu Wang , Xiaoyu Tao , Qi Liu , Enhong Chen

We introduce rStar2-Agent, a 14B math reasoning model trained with agentic reinforcement learning to achieve frontier-level performance. Beyond current long CoT, the model demonstrates advanced cognitive behaviors, such as thinking…

Recent years have seen significant advances in explainable AI as the need to understand deep learning models has gained importance with the increased emphasis on trust and ethics in AI. Comprehensible models for sequential decision tasks…

Artificial Intelligence · Computer Science 2022-08-19 Pedro Sequeira , Daniel Elenius , Jesse Hostetler , Melinda Gervasio

As autonomous AI agents increasingly mediate online platform markets, a fundamental question emerges: do these markets generate stable strategic outcomes? In repeated strategic environments, the Nash equilibrium provides a natural benchmark…

Artificial Intelligence · Computer Science 2026-04-28 Enoch Hyunwook Kang

Multi-agent systems have demonstrated exceptional performance in downstream tasks beyond diverse single agent baselines. A growing body of work has explored ways to improve their reasoning and collaboration, from vote, debate, to complex…

Artificial Intelligence · Computer Science 2026-02-13 Yu Yao , Jiayi Dong , Yang Yang , Ju Li , Yilun Du

How should an AI-based explanation system explain an agent's complex behavior to ordinary end users who have no background in AI? Answering this question is an active research area, for if an AI-based explanation system could effectively…

Human-Computer Interaction · Computer Science 2017-11-21 Jonathan Dodge , Sean Penney , Claudia Hilderbrand , Andrew Anderson , Margaret Burnett

This paper proposes a framework in which agents are constrained to use simple models to forecast economic variables and characterizes the resulting biases. It considers agents who can only entertain state-space models with no more than d…

Theoretical Economics · Economics 2024-10-10 Pooya Molavi

Multi-agent systems (MAS) built on large language models (LLMs) offer a promising path toward solving complex, real-world tasks that single-agent systems often struggle to manage. While recent advancements in test-time scaling (TTS) have…

Artificial Intelligence · Computer Science 2025-08-20 Can Jin , Hongwu Peng , Qixin Zhang , Yujin Tang , Dimitris N. Metaxas , Tong Che

Repository-level code agents have shown strong promise in real-world feature addition tasks, making reliable evaluation of their capabilities increasingly important. However, existing benchmarks primarily evaluate these agents as black…

Software Engineering · Computer Science 2026-03-30 Shuhan Liu , Zhiyi Zhao , Xing Hu , Kui Liu , Xiaohu Yang , Xin Xia

Long-horizon AI agents execute complex workflows spanning hundreds of sequential actions, yet a single wrong assumption early on can cascade into irreversible errors. When instructions are incomplete, the agent must decide not only whether…

Computation and Language · Computer Science 2026-05-11 Anmol Gulati , Hariom Gupta , Elias Lumer , Sahil Sen , Vamse Kumar Subbiah

Multi-agent systems built on large language models (LLMs) are expected to enhance decision-making by pooling distributed information, yet systematically evaluating this capability has remained challenging. We introduce HiddenBench, a…

Computation and Language · Computer Science 2026-05-14 Yuxuan Li , Aoi Naito , Hirokazu Shirado

Forecasting has become a natural benchmark for reasoning under uncertainty. Yet existing evaluations of large language models remain limited to judgmental tasks in simple formats, such as binary or multiple-choice questions. In practice,…

Machine Learning · Computer Science 2026-04-20 Jeremy Qin , Maksym Andriushchenko

Forecasting future events is highly valuable in decision-making and is a robust measure of general intelligence. As forecasting is probabilistic, developing and evaluating AI forecasters requires generating large numbers of diverse and…

Machine Learning · Computer Science 2026-03-11 Nikos I. Bosse , Peter Mühlbacher , Jack Wildman , Lawrence Phillips , Dan Schwarz

RLHF-aligned LMs have shown unprecedented ability on both benchmarks and long-form text generation, yet they struggle with one foundational task: next-token prediction. As RLHF models become agent models aimed at interacting with humans,…

Computation and Language · Computer Science 2024-07-03 Margaret Li , Weijia Shi , Artidoro Pagnoni , Peter West , Ari Holtzman

We introduce and study the problem of detecting whether an agent is updating their prior beliefs given new evidence in an optimal way that is Bayesian, or whether they are biased towards their own prior. In our model, biased agents form…

Computer Science and Game Theory · Computer Science 2024-10-31 Yiling Chen , Tao Lin , Ariel D. Procaccia , Aaditya Ramdas , Itai Shapira

We consider the problem of multi-task reasoning (MTR), where an agent can solve multiple tasks via (first-order) logic reasoning. This capability is essential for human-like intelligence due to its strong generalizability and simplicity for…

Artificial Intelligence · Computer Science 2022-02-15 Daoming Lyu , Bo Liu , Jianshu Chen

The advancement of large language model (LLM) based agents has shifted AI evaluation from single-turn response assessment to multi-step task completion in interactive environments. We present an empirical study evaluating frontier AI models…

Artificial Intelligence · Computer Science 2026-01-15 Logan Ritchie , Sushant Mehta , Nick Heiner , Mason Yu , Edwin Chen

Autonomous planning has been an ongoing pursuit since the inception of artificial intelligence. Based on curated problem solvers, early planning agents could deliver precise solutions for specific tasks but lacked generalization. The…

Artificial Intelligence · Computer Science 2024-10-17 Jian Xie , Kexun Zhang , Jiangjie Chen , Siyu Yuan , Kai Zhang , Yikai Zhang , Lei Li , Yanghua Xiao

Structured deliberation has been found to improve the performance of human forecasters. This study investigates whether a similar intervention, i.e. allowing LLMs to review each other's forecasts before updating, can improve accuracy in…

Artificial Intelligence · Computer Science 2025-12-30 Paul Schneider , Amalie Schramm
‹ Prev 1 4 5 6 7 8 10 Next ›