English
Related papers

Related papers: PuzzlePlex: Benchmarking Foundation Models on Reas…

200 papers

Algorithmic reasoning is a fundamental cognitive ability that plays a pivotal role in problem-solving and decision-making processes. Reinforcement Learning (RL) has demonstrated remarkable proficiency in tasks such as motor control,…

Machine Learning · Computer Science 2024-07-02 Benjamin Estermann , Luca A. Lanzendörfer , Yannick Niedermayr , Roger Wattenhofer

Puzzlehunts are a genre of complex, multi-step puzzles lacking well-defined problem definitions. In contrast to conventional reasoning benchmarks consisting of tasks with clear instructions and constrained environments, puzzlehunts requires…

As language models master existing reasoning benchmarks, we need new challenges to evaluate their cognitive frontiers. Puzzle-solving events are rich repositories of challenging multimodal problems that test a wide range of advanced…

Artificial Intelligence · Computer Science 2025-02-17 Clinton J. Wang , Dean Lee , Cristina Menghini , Johannes Mols , Jack Doughty , Adam Khoja , Jayson Lynch , Sean Hendryx , Summer Yue , Dan Hendrycks

We introduce PuzzleJAX, a GPU-accelerated puzzle game engine and description language designed to support rapid benchmarking of tree search, reinforcement learning, and LLM reasoning abilities. Unlike existing GPU-accelerated learning…

Artificial Intelligence · Computer Science 2025-08-26 Sam Earle , Graham Todd , Yuchen Li , Ahmed Khalifa , Muhammad Umair Nasir , Zehua Jiang , Andrzej Banburski-Fahey , Julian Togelius

We introduce Pencil Puzzle Bench, a framework for evaluating large language model reasoning through pencil puzzles, a family of constraint-satisfaction problems closely related to NP-complete problems, with deterministic, step-level…

Artificial Intelligence · Computer Science 2026-03-03 Justin Waugh

Solving topological grid puzzles requires reasoning over global spatial invariants such as connectivity, loop closure, and region symmetry and remains challenging for even the most powerful large language models (LLMs). To study these…

Artificial Intelligence · Computer Science 2026-03-13 Mayug Maniparambil , Nils Hoehing , Janak Kapuriya , Arjun Karuvally , Ellen Rushe , Anthony Ventresque , Noel O'Connor , Fergal Reid

We present PLUGH (https://www.urbandictionary.com/define.php?term=plugh), a modern benchmark that currently consists of 5 tasks, each with 125 input texts extracted from 48 different games and representing 61 different (non-isomorphic)…

Computation and Language · Computer Science 2024-08-12 Alexey Tikhonov

High-quality mathematical and logical datasets with verifiable answers are essential for strengthening the reasoning capabilities of large language models (LLMs). While recent data augmentation techniques have facilitated the creation of…

Artificial Intelligence · Computer Science 2026-05-29 Kai Xiong , Yanwei Huang , Rongjunchen Zhang , Kun Chen , Haipang Wu , Yingcai Wu

While recent advances in artificial intelligence have achieved human-level performance in environments like Starcraft and Go, many physical reasoning tasks remain challenging for modern algorithms. To date, few algorithms have been…

Artificial Intelligence · Computer Science 2023-02-02 Ken Kansky , Skanda Vaidyanath , Scott Swingle , Xinghua Lou , Miguel Lazaro-Gredilla , Dileep George

While advancements in NLP have significantly improved the performance of Large Language Models (LLMs) on tasks requiring vertical thinking, their lateral thinking capabilities remain under-explored and challenging to measure due to the…

Computation and Language · Computer Science 2024-10-10 Qi Chen , Bowen Zhang , Gang Wang , Qi Wu

Large language models (LLMs) excel at many supervised tasks but often struggle with structured reasoning in unfamiliar settings. This discrepancy suggests that standard fine-tuning pipelines may instill narrow, domain-specific heuristics…

Machine Learning · Computer Science 2025-06-06 Zhen Hao Wong , Jingwen Deng , Runming He , Zirong Chen , Qijie You , Hejun Dong , Hao Liang , Chengyu Shen , Bin Cui , Wentao Zhang

Existing reasoning benchmarks for large language models (LLMs) frequently fail to capture authentic creativity, often rewarding memorization of previously observed patterns. We address this shortcoming with Sudoku-Bench, a curated benchmark…

Artificial Intelligence · Computer Science 2025-05-23 Jeffrey Seely , Yuki Imajuku , Tianyu Zhao , Edoardo Cetin , Llion Jones

Reasoning is a fundamental capability of large language models (LLMs), enabling them to comprehend, analyze, and solve complex problems. In this paper, we introduce TextGames, an innovative benchmark specifically crafted to assess LLMs…

Computation and Language · Computer Science 2025-02-26 Frederikus Hudi , Genta Indra Winata , Ruochen Zhang , Alham Fikri Aji

Puzzles have long served as compact and revealing probes of human cognition, isolating abstraction, rule discovery, and systematic reasoning with minimal reliance on prior knowledge. Leveraging these properties, visual puzzles have recently…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Maria Lymperaiou , Vasileios Karampinis , Giorgos Filandrianos , Angelos Vlachos , Chrysoula Zerva , Athanasios Voulodimos

Evaluating the reasoning capabilities of Large Language Models is increasingly challenging as models improve. Human curation of hard questions is highly expensive, especially in recent benchmarks using PhD-level domain knowledge to…

Artificial Intelligence · Computer Science 2026-05-19 Simon Henniger , Gabriel Poesia

Large Reasoning Models (LRMs) have demonstrated impressive performance on complex tasks, including logical puzzle games that require deriving solutions satisfying all constraints. However, whether they can flexibly apply appropriate rules…

Artificial Intelligence · Computer Science 2026-03-03 Jingcong Liang , Shijun Wan , Xuehai Wu , Yitong Li , Qianglong Chen , Duyu Tang , Siyuan Wang , Zhongyu Wei

We introduce LLM CHESS, an evaluation framework designed to probe the generalization of reasoning and instruction-following abilities in large language models (LLMs) through extended agentic interaction in the domain of chess. We rank over…

Artificial Intelligence · Computer Science 2025-12-02 Sai Kolasani , Maxim Saplin , Nicholas Crispino , Kyle Montgomery , Jared Quincy Davis , Matei Zaharia , Chi Wang , Chenguang Wang

We introduce FrontierCS, a benchmark of 156 open-ended problems across diverse areas of computer science, designed and reviewed by experts, including CS PhDs and top-tier competitive programming participants and problem setters. Unlike…

In experimental applications of bounded-reasoning models, behavior is often summarized by distributions of "levels". We argue that such summaries conflate two conceptually distinct dimensions: a player's type, capturing beliefs about what…

Theoretical Economics · Economics 2026-04-15 Shuige Liu , Gabriel Ziegler

Existing benchmarks for frontier models often test specialized, "PhD-level" knowledge that is difficult for non-experts to grasp. In contrast, we present a benchmark with 613 problems based on the NPR Sunday Puzzle Challenge that requires…

‹ Prev 1 2 3 10 Next ›