On the Reliability of Large Language Models to Misinformed and Demographically-Informed Prompts
计算与语言
2024-10-18 v2 人工智能
计算机与社会
人机交互
摘要
我们调查并观察了基于大型语言模型(LLM)的聊天机器人在应对误信息提示和包含人口统计学信息的提问时的行为和性能,这些提问涉及气候变化和精神健康领域。通过定量和定性方法,我们评估了聊天机器人辨别陈述真伪的能力、其对事实的遵循程度,以及其响应中偏见或误information 的存在。我们的定量分析使用真/假问题表明,这些聊天机器人能够可靠地给出正确答案。然而,来自领域专家的定性见解显示,仍然存在关于隐私、伦理问题以及聊天机器人需要将用户引导至专业服务的必要性的顾虑。我们得出结论,尽管这些聊天机器人具有巨大的潜力,但在敏感领域的部署仍需谨慎考虑、伦理监督和严格的改进,以确保它们作为人类专长的有益补充而非自主解决方案。
引用
@article{arxiv.2410.10850,
title = {On the Reliability of Large Language Models to Misinformed and Demographically-Informed Prompts},
author = {Toluwani Aremu and Oluwakemi Akinwehinmi and Chukwuemeka Nwagu and Syed Ishtiaque Ahmed and Rita Orji and Pedro Arnau Del Amo and Abdulmotaleb El Saddik},
journal= {arXiv preprint arXiv:2410.10850},
year = {2024}
}
备注
Study conducted between August and December 2023. Under review at AAAI-AI Magazine. Submitted for archival purposes only