语言模型预测并不可靠地区分不可能与不太可能的事件
计算与语言
2025-06-10 v1 人工智能
摘要
语言模型能否可靠地预测可能的事件比仅不太可能的事件更可能?通过分离可能性、典型性和情境关联性,我们显示尽管已有研究结果表明,但语言模型的能力远非稳健。在事实上,在某些条件下,所有测试的模型——包括 Llama 3、Gemma 2 和 Mistral NeMo ——都以劣于随机水平的表现,给予不可能的句子如 'the car was given a parking ticket by the brake' 更高的概率,而给予仅不太可能的句子如 'the car was given a parking ticket by the explorer' 更低的概率。
引用
@article{arxiv.2506.06808,
title = {Not quite Sherlock Holmes: Language model predictions do not reliably differentiate impossible from improbable events},
author = {James A. Michaelov and Reeka Estacio and Zhien Zhang and Benjamin K. Bergen},
journal= {arXiv preprint arXiv:2506.06808},
year = {2025}
}
备注
Accepted to Findings of ACL 2025