大型语言模型诚实性综述
计算与语言
2024-09-30 v1 人工智能
摘要
诚实性是使大型语言模型与人类价值观对齐的一项基本原则,要求这些模型识别它们所知与不知,并能够忠实地表达其知识。尽管前景广阔,当前的 LLM 仍表现出显著的 dishonest 行为,例如自信地给出错误答案或未能表达其所知。此外,关于 LLM 诚实性的研究也面临挑战,包括对诚实性的定义不一、难以区分已知与未知知识,以及缺乏对相关研究的全面理解。为了解决这些问题,我们提供了一篇关于 LLM 诚实性的综述,涵盖其概念澄清、评估方法以及改进策略。此外,我们提供了对未来研究的见解,旨在激发对这一重要领域的进一步探索。
引用
@article{arxiv.2409.18786,
title = {A Survey on the Honesty of Large Language Models},
author = {Siheng Li and Cheng Yang and Taiqiang Wu and Chufan Shi and Yuji Zhang and Xinyu Zhu and Zesen Cheng and Deng Cai and Mo Yu and Lemao Liu and Jie Zhou and Yujiu Yang and Ngai Wong and Xixin Wu and Wai Lam},
journal= {arXiv preprint arXiv:2409.18786},
year = {2024}
}
备注
Project Page: https://github.com/SihengLi99/LLM-Honesty-Survey