大语言模型中认知偏见的全面评估
计算与语言
2025-11-04 v2 人工智能
摘要
我们对 30 种认知偏见在 20 个领先的大语言模型(LLMs)中进行了大规模评估,涉及各种决策情景。我们的贡献包括:一个用于可靠且大规模生成 LLM 测试的新型通用测试框架、一个包含 30,000 项用于检测 LLMs 认知偏见的基准数据集,以及对 20 个评估模型中发现的偏见的全面评估。我们的工作确认并拓宽了之前的研究发现,即 LLMs 存在认知偏见,报告了在至少部分 20 个 LLMs 中检测到所有 30 种偏见的证据。我们发布了框架代码以鼓励未来研究: https://github.com/simonmalberg/cognitive-biases-in-llms
引用
@article{arxiv.2410.15413,
title = {A Comprehensive Evaluation of Cognitive Biases in LLMs},
author = {Simon Malberg and Roman Poletukhin and Carolin M. Schuster and Georg Groh},
journal= {arXiv preprint arXiv:2410.15413},
year = {2025}
}
备注
Published in "Proceedings of the 5th International Conference on Natural Language Processing for Digital Humanities"