English
Related papers

Related papers: Meterstick: Benchmarking Performance Variability i…

200 papers

Modern language models (LMs) pose a new challenge in capability assessment. Static benchmarks inevitably saturate without providing confidence in the deployment tolerances of LM-based systems, but developers nonetheless claim that their…

Software Engineering · Computer Science 2024-07-31 Michael Saxon , Ari Holtzman , Peter West , William Yang Wang , Naomi Saphra

Achieving optimal balance in games is essential to their success, yet reliant on extensive manual work and playtesting. To facilitate this process, the Procedural Content Generation via Reinforcement Learning (PCGRL) framework has recently…

Human-Computer Interaction · Computer Science 2024-09-10 Florian Rupp , Alessandro Puddu , Christian Becker-Asano , Kai Eckert

Quantification of human group-behavior has so far defied an empirical, falsifiable approach. This is due to tremendous difficulties in data acquisition of social systems. Massive multiplayer online games (MMOG) provide a fascinating new way…

Physics and Society · Physics 2013-07-10 Michael Szell , Stefan Thurner

Progress in multiagent intelligence research is fundamentally limited by the number and quality of environments available for study. In recent years, simulated games have become a dominant research platform within reinforcement learning, in…

Machine Learning · Computer Science 2020-04-20 Joseph Suarez , Yilun Du , Igor Mordatch , Phillip Isola

Competitive Self-Play (CSP) based Multi-Agent Reinforcement Learning (MARL) has shown phenomenal breakthroughs recently. Strong AIs are achieved for several benchmarks, including Dota 2, Glory of Kings, Quake III, StarCraft II, to name a…

Machine Learning · Computer Science 2020-12-01 Peng Sun , Jiechao Xiong , Lei Han , Xinghai Sun , Shuxing Li , Jiawei Xu , Meng Fang , Zhengyou Zhang

As large language models (LLMs) advance across diverse tasks, the need for comprehensive evaluation beyond single metrics becomes increasingly important. To fully assess LLM intelligence, it is crucial to examine their interactive dynamics…

Computation and Language · Computer Science 2025-09-23 Junhao Chen , Jingbo Sun , Xiang Li , Haidong Xin , Yuhao Xue , Yibin Xu , Hao Zhao

Vision-language models (VLMs) have achieved strong results on coding and math benchmarks that are challenging for humans, yet their ability to perform tasks that come naturally to humans--such as perception, spatial navigation, and memory…

Artificial Intelligence · Computer Science 2026-05-18 Alex L. Zhang , Thomas L. Griffiths , Karthik R. Narasimhan , Ofir Press

The concept of the Metaverse has garnered growing interest from both academic and industry circles. The decentralization of both the integrity and security of digital items has spurred the popularity of play-to-earn (P2E) games, where…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-12-12 Chang Liu , Terence Jie Chua , Jun Zhao

Many online services running in datacenters are implemented using a microservice software architecture characterized by strict latency requirements. Consequently, this popular software paradigm is increasingly used for the performance…

Hardware Architecture · Computer Science 2024-10-16 Georgia Antoniou , Haris Volos , Yiannakis Sazeides

Data ecosystems are becoming larger and more complex due to online tracking, wearable computing, and the Internet of Things. But privacy concerns are threatening to erode the potential benefits of these systems. Recently, users have…

Cryptography and Security · Computer Science 2017-10-17 Jeffrey Pawlick , Quanyan Zhu

Recent advances in generative models have significantly impacted game generation. However, despite producing high-quality graphics and adequately receiving player input, existing models often fail to maintain fundamental game properties…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Jingye Chen , Yuzhong Zhao , Yupan Huang , Lei Cui , Li Dong , Tengchao Lv , Qifeng Chen , Furu Wei

Large Language Models (LLMs) have shown great success as high-level planners for zero-shot game-playing agents. However, these agents are primarily evaluated on Minecraft, where long-term planning is relatively straightforward. In contrast,…

Artificial Intelligence · Computer Science 2024-03-04 Dominik Jeurissen , Diego Perez-Liebana , Jeremy Gow , Duygu Cakmak , James Kwan

The pursuit of leaderboard rankings in Large Language Models (LLMs) has created a fundamental paradox: models excel at standardized tests while failing to demonstrate genuine language understanding and adaptability. Our systematic analysis…

Computation and Language · Computer Science 2024-12-06 Sourav Banerjee , Ayushi Agarwal , Eishkaran Singh

Adversarial board games, as a paradigmatic domain of strategic reasoning and intelligence, have long served as both a popular competitive activity and a benchmark for evaluating artificial intelligence (AI) systems. Building on this…

Artificial Intelligence · Computer Science 2025-08-08 Yingjie Zhou , Jiezhang Cao , Farong Wen , Li Xu , Yanwei Jiang , Jun Jia , Ronghui Li , Xiaohong Liu , Yu Zhou , Xiongkuo Min , Jie Guo , Zicheng Zhang , Guangtao Zhai

We propose GuessBench, a novel benchmark that evaluates Vision Language Models (VLMs) on modeling the pervasive, noisy, and pluralistic human creativity. GuessBench sources data from "Guess the Build", an online multiplayer Minecraft…

Computation and Language · Computer Science 2025-06-09 Zifeng Zhu , Shangbin Feng , Herun Wan , Ningnan Wang , Minnan Luo , Yulia Tsvetkov

Multi-agent reinforcement learning (MARL), as a thriving field, explores how multiple agents independently make decisions in a shared dynamic environment. Due to environmental uncertainties, policies in MARL must remain robust to tackle the…

Machine Learning · Computer Science 2025-12-02 Na Li , Zewu Zheng , Wei Ni , Hangguan Shan , Wenjie Zhang , Xinyu Li

Developing AI agents capable of interacting with open-world environments to solve diverse tasks is a compelling challenge. However, evaluating such open-ended agents remains difficult, with current benchmarks facing scalability limitations.…

Artificial Intelligence · Computer Science 2025-06-04 Xinyue Zheng , Haowei Lin , Kaichen He , Zihao Wang , Zilong Zheng , Yitao Liang

Mobile games are extremely popular and engage millions of people every day. Even if they are often quite simple, their development features a high degree of difficulty and requires close attention to both achieve a high satisfaction from…

Computers and Society · Computer Science 2017-03-14 Daniele Grassi , Giacomo Barigazzi , Giacomo Cabri

Hyperscalars run services across a large fleet of servers, serving billions of users worldwide. These services, however, behave differently than commonly available benchmark suites, resulting in server architectures that are not optimized…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-05-03 Suyash Mahar , Hao Wang , Wei Shu , Abhishek Dhanotia

Open world games present players with more freedom than games with linear progression structures. However, without clearly-defined objectives, they often leave players without a sense of purpose. Most of the time, quests and objectives are…

Multimedia · Computer Science 2017-05-02 Ryan Alexander , Chris Martens