English
Related papers

Related papers: Evaluating Language Models' Evaluations of Games

200 papers

We introduce AI rationalization, an approach for generating explanations of autonomous system behavior as if a human had performed the behavior. We describe a rationalization technique that uses neural machine translation to translate…

Artificial Intelligence · Computer Science 2017-12-20 Upol Ehsan , Brent Harrison , Larry Chan , Mark O. Riedl

Defining and measuring decision-making styles, also known as playstyles, is crucial in gaming, where these styles reflect a broad spectrum of individuality and diversity. However, finding a universally applicable measure for these styles…

Artificial Intelligence · Computer Science 2024-09-02 Chiu-Chou Lin , Wei-Chen Chiu , I-Chen Wu

Do AI systems truly understand human concepts or merely mimic surface patterns? We investigate this through chess, where human creativity meets precise strategic concepts. Analyzing a 270M-parameter transformer that achieves…

Machine Learning · Computer Science 2025-11-05 Semyon Lomasov , Judah Goldfeder , Mehmet Hamza Erol , Matthew So , Yao Yan , Addison Howard , Nathan Kutz , Ravid Shwartz Ziv

We explore how an AI model's decision fairness affects people's engagement with and perceived fairness of the model if they are subject to its decisions, but could repeatedly and strategically respond to these decisions. Two types of…

Human-Computer Interaction · Computer Science 2024-10-07 Meric Altug Gemalmaz , Ming Yin

Artificial intelligence (AI)-based decision support systems can be highly accurate yet still fail to support users or improve decisions. Existing theories of AI-assisted decision-making focus on calibrating reliance on AI advice, leaving it…

Human-Computer Interaction · Computer Science 2026-02-03 Venkatesh Sivaraman , Eric P. Mason , Mengfan Ellen Li , Jessica Tong , Andrew J. King , Jeremy M. Kahn , Adam Perer

Deceptive agents are a challenge for the safety, trustworthiness, and cooperation of AI systems. We focus on the problem that agents might deceive in order to achieve their goals (for instance, in our experiments with language models, the…

Artificial Intelligence · Computer Science 2023-12-05 Francis Rhys Ward , Francesco Belardinelli , Francesca Toni , Tom Everitt

As technologies become more and more pervasive, there is a need for considering the affective dimension of interaction with computer systems to make them more human-like. Current demands for this matter include accurate emotion recognition,…

Human-Computer Interaction · Computer Science 2018-06-13 Barbara Giżycka , Grzegorz J. Nalepa , Paweł Jemioło

Adversarial board games, as a paradigmatic domain of strategic reasoning and intelligence, have long served as both a popular competitive activity and a benchmark for evaluating artificial intelligence (AI) systems. Building on this…

Artificial Intelligence · Computer Science 2025-08-08 Yingjie Zhou , Jiezhang Cao , Farong Wen , Li Xu , Yanwei Jiang , Jun Jia , Ronghui Li , Xiaohong Liu , Yu Zhou , Xiongkuo Min , Jie Guo , Zicheng Zhang , Guangtao Zhai

If you are an artificial intelligence researcher, you should look to video games as ideal testbeds for the work you do. If you are a video game developer, you should look to AI for the technology that makes completely new types of games…

Artificial Intelligence · Computer Science 2016-12-07 Julian Togelius

Playing games is inherently human, and a lot of games are created to challenge different human characteristics. However, these tasks are often left out when evaluating the human-like nature of artificial models. The objective of this work…

Computer Vision and Pattern Recognition · Computer Science 2025-09-04 Nuria Alabau-Bosque , Jorge Vila-Tomás , Paula Daudén-Oliver , Pablo Hernández-Cámara , Jose Manuel Jaén-Lorites , Valero Laparra , Jesús Malo

Language carries thought and coordination among humans but rarely reaches further along the spectrum of diverse intelligence. Yet non-neural systems -- from gene regulatory networks and microbial consortia to fungi -- are increasingly…

Machine Learning · Computer Science 2026-05-19 Yanbo Zhang , Michael Levin

It is a long-standing goal of artificial intelligence (AI) to be superior to human beings in decision making. Games are suitable for testing AI capabilities of making good decisions in non-numerical tasks. In this paper, we develop a new AI…

Artificial Intelligence · Computer Science 2021-02-16 Ran Tian , Nan Li , Ilya Kolmanovsky , Anouck Girard

Physical reasoning is a crucial aspect in the development of general AI systems, given that human learning starts with interacting with the physical world before progressing to more complex concepts. Although researchers have studied and…

Artificial Intelligence · Computer Science 2023-12-19 Andrew Melnik , Robin Schiewer , Moritz Lange , Andrei Muresanu , Mozhgan Saeidi , Animesh Garg , Helge Ritter

Social deduction games like Werewolf combine language, reasoning, and strategy, providing a testbed for studying natural language and social intelligence. However, most studies reduce the game to LLM-based self-play, yielding templated…

Computation and Language · Computer Science 2025-10-14 Zirui Song , Yuan Huang , Junchang Liu , Haozhe Luo , Chenxi Wang , Lang Gao , Zixiang Xu , Mingfei Han , Xiaojun Chang , Xiuying Chen

As evaluation designs of large language models may shape our trajectory toward artificial general intelligence, comprehensive and forward-looking assessment is essential. Existing benchmarks primarily assess static knowledge, while…

Computation and Language · Computer Science 2025-08-07 Jiayin Wang , Zhiquang Guo , Weizhi Ma , Min Zhang

Artificial Intelligence (AI) technologies have been developed rapidly, and AI-based systems have been widely used in various application domains with opportunities and challenges. However, little is known about the architecture decisions…

Software Engineering · Computer Science 2022-12-29 Beiqi Zhang , Tianyang Liu , Peng Liang , Chong Wang , Mojtaba Shahin , Jiaxin Yu

The rapid advancements in large Language models (LLMs) have significantly enhanced their reasoning capabilities, driven by various strategies such as multi-agent collaboration. However, unlike the well-established performance improvements…

Artificial Intelligence · Computer Science 2026-04-23 Zihan Chen , Song Wang , Zhen Tan , Xingbo Fu , Zhenyu Lei , Peng Wang , Huan Liu , Cong Shen , Jundong Li

In this paper, we argue that simulation platforms enable a novel type of embodied spatial reasoning, one facilitated by a formal model of object and event semantics that renders the continuous quantitative search space of an open-world,…

Artificial Intelligence · Computer Science 2019-02-07 James Pustejovsky , Nikhil Krishnaswamy

The rapid advancement of LLMs sparked significant interest in their potential to augment or automate managerial functions. One of the most recent trends in AI benchmarking is performance of Large Language Models (LLMs) over longer time…

Artificial Intelligence · Computer Science 2025-10-01 Berdymyrat Ovezmyradov

An essential element of human mathematical reasoning is our number sense -- an abstract understanding of numbers and their relationships -- which allows us to solve problems involving vast number spaces using limited computational…

Artificial Intelligence · Computer Science 2025-04-02 Roussel Rahman
‹ Prev 1 8 9 10 Next ›