MAD Chairs: A new tool to evaluate AI
Abstract
This paper contributes a new way to evaluate AI. Much as one might evaluate a machine in terms of its performance at chess, this approach involves evaluating a machine in terms of its performance at a game called "MAD Chairs". At the time of writing, evaluation with this game exposed opportunities to improve Claude, Gemini, ChatGPT, Qwen and DeepSeek. Furthermore, this paper sets a stage for future innovation in game theory and AI safety by providing an example of success with non-standard approaches to each: studying a game beyond the scope of previous game theoretic tools and mitigating a serious AI safety risk in a way that requires neither determination of values nor their enforcement.
Keywords
Cite
@article{arxiv.2503.20986,
title = {MAD Chairs: A new tool to evaluate AI},
author = {Chris Santos-Lang},
journal= {arXiv preprint arXiv:2503.20986},
year = {2025}
}
Comments
17 pages, 1 figure, reproduced with permission from Springer Nature from Coordination, Organizations, Institutions, Norms, and Ethics for Governance of Multi-Agent Systems XVIII (COINE 2025)