English
Related papers

Related papers: Baba Is AI: Break the Rules to Beat the Benchmark

200 papers

Existing benchmarks do not test language agents on their interaction with human users or ability to follow domain-specific rules, both of which are vital for deploying them in real world applications. We propose $\tau$-bench, a benchmark…

Artificial Intelligence · Computer Science 2024-06-19 Shunyu Yao , Noah Shinn , Pedram Razavi , Karthik Narasimhan

In dynamic collaborative settings, for artificial intelligence (AI) agents to better align with humans, they must adapt to novel teammates who utilise unforeseen strategies. While adaptation is often simple for humans, it can be challenging…

Machine Learning · Computer Science 2025-04-22 Ravi Hammond , Dustin Craggs , Mingyu Guo , Jakob Foerster , Ian Reid

Recent large language models (LLMs) have demonstrated great potential toward intelligent agents and next-gen automation, but there currently lacks a systematic benchmark for evaluating LLMs' abilities as agents. We introduce SmartPlay: both…

Machine Learning · Computer Science 2024-03-19 Yue Wu , Xuan Tang , Tom M. Mitchell , Yuanzhi Li

The rapid evolution of large language models (LLMs) has expanded their capabilities from basic dialogue to advanced scientific reasoning. However, existing benchmarks in biology often fail to assess a critical skill required of researchers:…

Artificial Intelligence · Computer Science 2026-02-06 Junting Zhou , Jin Chen , Linfeng Hao , Denghui Cao , Zheyu Wang , Qiguang Chen , Chaoyou Fu , Jiaze Chen , Yuchen Wu , Ge Zhang , Mingxuan Wang , Wenhao Huang , Tong Yang

Artificial intelligence-based systems for player risk detection have become central to harm prevention efforts in the gambling industry. However, growing concerns around transparency and effectiveness have highlighted the absence of…

While Artificial Intelligence has successfully outperformed humans in complex combinatorial games (such as chess and checkers), humans have retained their supremacy in social interactions that require intuition and adaptation, such as…

Computers and Society · Computer Science 2014-04-22 Fatimah Ishowo-Oloko , Jacob Crandall , Manuel Cebrian , Sherief Abdallah , Iyad Rahwan

Recent advances in artificial intelligence have been strongly driven by the use of game environments for training and evaluating agents. Games are often accessible and versatile, with well-defined state-transitions and goals allowing for…

Machine Learning · Computer Science 2019-09-19 Benjamin Beyret , José Hernández-Orallo , Lucy Cheke , Marta Halina , Murray Shanahan , Matthew Crosby

The performance of AI models on safety benchmarks does not indicate their real-world performance after deployment. This opaqueness of AI models impedes existing regulatory frameworks constituted on benchmark performance, leaving them…

Machine Learning · Computer Science 2025-12-16 Gabriel Stanovsky , Renana Keydar , Gadi Perl , Eliya Habba

AI is becoming increasingly integrated into everyday life, both in professional work environments and in leisure and entertainment contexts. This integration requires AI to move beyond acting as an assistant for informational or…

Human-Computer Interaction · Computer Science 2026-02-26 Christian Poelitz , Finale Doshi-Velez , Siân Lindley

AI-controlled characters in fighting games are expected to possess reasonably high skills and behave in a believable, human-like manner, exhibiting a diversity of play styles and strategies. Thus, the development of fighting game AI…

Artificial Intelligence · Computer Science 2021-08-10 Kaori Yuda , Shota Kamei , Riku Tanji , Ryoya Ito , Ippo Wakana , Maxim Mozgovoy

The rise of agentic AI systems, where agents collaborate to perform diverse tasks, poses new challenges with observing, analyzing and optimizing their behavior. Traditional evaluation and benchmarking approaches struggle to handle the…

Artificial Intelligence · Computer Science 2025-03-11 Dany Moshkovich , Hadar Mulian , Sergey Zeltyn , Natti Eder , Inna Skarbovsky , Roy Abitbol

We propose a novel in-context learning algorithm for building autonomous decision-making language agents. The language agent continuously attempts to solve the same task by self-correcting each time the task fails. Our selected language…

Artificial Intelligence · Computer Science 2024-11-06 Abhishek Dutta , Yen-Che Hsiao

The rapid pace of recent research in AI has been driven in part by the presence of fast and challenging simulation environments. These environments often take the form of games; with tasks ranging from simple board games, to competitive…

For over a decade now, robotics and the use of artificial agents have become a common thing.Testing the performance of new path finding or search space optimization algorithms has also become a challenge as they require simulation or an…

Machine Learning · Computer Science 2022-07-29 Jerin Paul Selvan , Pravin S. Game

This paper examines the reasoning capabilities of Large Language Models (LLMs) from a novel perspective, focusing on their ability to operate within formally specified, rule-governed environments. We evaluate four LLMs (Gemini 2.5 Pro and…

Artificial Intelligence · Computer Science 2026-02-24 Maciej Świechowski , Adam Żychowski , Jacek Mańdziuk

The rapid advancement of General Purpose AI (GPAI) models necessitates robust evaluation frameworks, especially with emerging regulations like the EU AI Act and its associated Code of Practice (CoP). Current AI evaluation practices depend…

Artificial Intelligence · Computer Science 2025-08-11 Matteo Prandi , Vincenzo Suriani , Federico Pierucci , Marcello Galisai , Daniele Nardi , Piercosma Bisconti

Although recent developments in generative AI have greatly enhanced the capabilities of conversational agents such as Google's Gemini (formerly Bard) or OpenAI's ChatGPT, it's unclear whether the usage of these agents aids users across…

Human-Computer Interaction · Computer Science 2024-04-03 Crystal Qian , James Wexler

Benchmarking has long served as a foundational practice in machine learning and, increasingly, in modern AI systems such as large language models, where shared tasks, metrics, and leaderboards offer a common basis for measuring progress and…

Artificial Intelligence · Computer Science 2026-02-16 Philip Waggoner

This paper introduces an information-theoretic method for selecting a subset of problems which gives the most information about a group of problem-solving algorithms. This method was tested on the games in the General Video Game AI (GVGAI)…

Artificial Intelligence · Computer Science 2020-05-19 Matthew Stephenson , Damien Anderson , Ahmed Khalifa , John Levine , Jochen Renz , Julian Togelius , Christoph Salge

Artificial agents have traditionally been trained to maximize reward, which may incentivize power-seeking and deception, analogous to how next-token prediction in language models (LMs) may incentivize toxicity. So do agents naturally learn…