规模的误区:探究大语言模型中的重新定义逆任务
计算与语言
2025-06-03 v2
摘要
逆向任务可揭示大规模语言模型(LLM)潜在的推理缺陷。在本工作中,我们探讨了重新定义任务,即为众所周知的物理常数和计量单位指定替代值,要求模型作出相应响应。我们的发现表明,模型性能随规模而下降,同时其虚假自信心也随之上升。此外,尽管提示策略和响应格式等因素具有影响,但它们并不能防止LLM锚定于记忆中的值。
引用
@article{arxiv.2502.12821,
title = {Pitfalls of Scale: Investigating the Inverse Task of Redefinition in Large Language Models},
author = {Elena Stringli and Maria Lymperaiou and Giorgos Filandrianos and Athanasios Voulodimos and Giorgos Stamou},
journal= {arXiv preprint arXiv:2502.12821},
year = {2025}
}
备注
Accepted at Findings of ACL 2025