English
Related papers

Related papers: Towards Game-Playing AI Benchmarks via Performance…

200 papers

The rapid uptake of generative artificial intelligence (AI) in higher education is reshaping assessment practices and intensifying concerns around academic integrity, fairness, and learning quality. While institutional responses…

Computers and Society · Computer Science 2026-05-28 Ndidi Bianca Ogbo , Zhao Song , Shatha Ghareeb , The Anh Han

Serious games are widely used for learning and training across domains such as healthcare, defense, and education. Persistent challenges remain, however, including static scenario design, authoring bottlenecks, limited learner modeling, and…

Artificial Intelligence · Computer Science 2026-05-22 Priyamvada Tripathi , Bill Kapralos

As AI becomes increasingly embedded in digital games, players' attitudes de-pend not only on whether AI is used, but also on where and how it intervenes in gameplay. This study examines players' evaluative patterns toward eight AI…

Human-Computer Interaction · Computer Science 2026-05-01 Ting-Chen Hsu , Jiangxu Lin , Wenran Chen , Fei Qin , Zheyuan Zhang

Humans solve problems by following existing rules and procedures, and also by leaps of creativity to redefine those rules and objectives. To probe these abilities, we developed a new benchmark based on the game Baba Is You where an agent…

Computation and Language · Computer Science 2025-09-11 Nathan Cloos , Meagan Jens , Michelangelo Naim , Yen-Ling Kuo , Ignacio Cases , Andrei Barbu , Christopher J. Cueva

General game playing artificial intelligence has recently seen important advances due to the various techniques known as 'deep learning'. However the advances conceal equally important limitations in their reliance on: massive data sets;…

Human-Computer Interaction · Computer Science 2016-06-22 Benjamin Ultan Cowley

I would like to share recommendations on how to do performance benchmarks for the purpose of computer science research evaluation. Research in my field (programming language research) often involves performance considerations, but it is…

Programming Languages · Computer Science 2026-05-05 Gabriel Scherer

In 2016, 2017, and 2018 at the IEEE Conference on Computational Intelligence in Games, the authors of this paper ran a competition for agents that can play classic text-based adventure games. This competition fills a gap in existing game AI…

Artificial Intelligence · Computer Science 2019-01-25 Timothy Atkinson , Hendrik Baier , Tara Copplestone , Sam Devlin , Jerry Swan

The creation of benchmarks to evaluate the safety of Large Language Models is one of the key activities within the trusted AI community. These benchmarks allow models to be compared for different aspects of safety such as toxicity, bias,…

Artificial Intelligence · Computer Science 2025-06-23 Lina Berrayana , Sean Rooney , Luis Garcés-Erice , Ioana Giurgiu

AI systems face a growing number of AI security threats that are increasingly exploited in the real world. Hence, shared AI incident reporting practices are emerging in industry as best practice and as mandated by regulatory requirements.…

Cryptography and Security · Computer Science 2026-05-06 Lukas Bieringer , Sean McGregor , Nicole Nichols , Kevin Paeth , Jochen Stängler , Andreas Wespi , Alexandre Alahi , Kathrin Grosse

The evaluation of clustering algorithms can involve running them on a variety of benchmark problems, and comparing their outputs to the reference, ground-truth groupings provided by experts. Unfortunately, many research papers and graduate…

Machine Learning · Computer Science 2023-10-27 Marek Gagolewski

Assessing fairness in artificial intelligence (AI) typically involves AI experts who select protected features, fairness metrics, and set fairness thresholds to assess outcome fairness. However, little is known about how stakeholders,…

Artificial Intelligence · Computer Science 2026-02-27 Lin Luo , Yuri Nakao , Mathieu Chollet , Hiroya Inakoshi , Simone Stumpf

Benchmarks shape scientific conclusions about model capabilities and steer model development. This creates a feedback loop: stronger benchmarks drive better models, and better models demand more discriminative benchmarks. Ensuring benchmark…

Computation and Language · Computer Science 2025-10-01 Arda Uzunoglu , Tianjian Li , Daniel Khashabi

Many works in the domain of artificial intelligence in games focus on board or video games due to the ease of reimplementing their mechanics. Decision-making problems in real-world sports share many similarities to such domains.…

Artificial Intelligence · Computer Science 2024-08-13 Carlo Nübel , Alexander Dockhorn , Sanaz Mostaghim

This paper argues that market governance mechanisms should be considered a key approach in the governance of artificial intelligence (AI), alongside traditional regulatory frameworks. While current governance approaches have predominantly…

General Economics · Economics 2025-03-06 Philip Moreira Tomei , Rupal Jain , Matija Franklin

People enjoy encounters with generative software, but rarely are they encouraged to interact with, understand or engage with it. In this paper we define the term 'PCG-based game', and explain how this concept follows on from the idea of an…

Artificial Intelligence · Computer Science 2016-10-12 Michael Cook , Mirjam Eladhari , Andy Nealen , Mike Treanor , Eddy Boxerman , Alex Jaffe , Paul Sottosanti , Steve Swink

The meteoric rise of AI, with its rapidly expanding market capitalization, presents both transformative opportunities and critical challenges. Chief among these is the urgent need for a new, unified paradigm for trustworthy evaluation, as…

Participants in recent discussions of AI-related issues ranging from intelligence explosion to technological unemployment have made diverse claims about the nature, pace, and drivers of progress in AI. However, these theories are rarely…

Artificial Intelligence · Computer Science 2015-12-21 Miles Brundage

In this position paper, we argue that human baselines in foundation model evaluations must be more rigorous and more transparent to enable meaningful comparisons of human vs. AI performance, and we provide recommendations and a reporting…

The rapid advancement of LLMs sparked significant interest in their potential to augment or automate managerial functions. One of the most recent trends in AI benchmarking is performance of Large Language Models (LLMs) over longer time…

Artificial Intelligence · Computer Science 2025-10-01 Berdymyrat Ovezmyradov