Overcoming the Machine Penalty with Imperfectly Fair AI Agents
Abstract
Despite rapid technological progress, effective human-machine cooperation remains a significant challenge. Humans tend to cooperate less with machines than with fellow humans, a phenomenon known as the machine penalty. Here, we show that artificial intelligence (AI) agents powered by large language models can overcome this penalty in social dilemma games with communication. In a pre-registered experiment with 1,152 participants, we deploy AI agents exhibiting three distinct personas: selfish, cooperative, and fair. However, only fair agents elicit human cooperation at rates comparable to human-human interactions. Analysis reveals that fair agents, similar to human participants, occasionally break pre-game cooperation promises, but nonetheless effectively establish cooperation as a social norm. These results challenge the conventional wisdom of machines as altruistic assistants or rational actors. Instead, our study highlights the importance of AI agents reflecting the nuanced complexity of human social behaviors -- imperfect yet driven by deeper social cognitive processes.
Cite
@article{arxiv.2410.03724,
title = {Overcoming the Machine Penalty with Imperfectly Fair AI Agents},
author = {Zhen Wang and Ruiqi Song and Chen Shen and Shiya Yin and Zhao Song and Balaraju Battu and Lei Shi and Danyang Jia and Talal Rahwan and Shuyue Hu},
journal= {arXiv preprint arXiv:2410.03724},
year = {2025}
}