English

How Different AI Chatbots Behave? Benchmarking Large Language Models in Behavioral Economics Games

Artificial Intelligence 2024-12-18 v1 Computation and Language

Abstract

The deployment of large language models (LLMs) in diverse applications requires a thorough understanding of their decision-making strategies and behavioral patterns. As a supplement to a recent study on the behavioral Turing test, this paper presents a comprehensive analysis of five leading LLM-based chatbot families as they navigate a series of behavioral economics games. By benchmarking these AI chatbots, we aim to uncover and document both common and distinct behavioral patterns across a range of scenarios. The findings provide valuable insights into the strategic preferences of each LLM, highlighting potential implications for their deployment in critical decision-making roles.

Keywords

Cite

@article{arxiv.2412.12362,
  title  = {How Different AI Chatbots Behave? Benchmarking Large Language Models in Behavioral Economics Games},
  author = {Yutong Xie and Yiyao Liu and Zhuang Ma and Lin Shi and Xiyuan Wang and Walter Yuan and Matthew O. Jackson and Qiaozhu Mei},
  journal= {arXiv preprint arXiv:2412.12362},
  year   = {2024}
}

Comments

Presented at The First Workshop on AI Behavioral Science (AIBS 2024)