中文

规模的误区:探究大语言模型中的重新定义逆任务

计算与语言 2025-06-03 v2

摘要

逆向任务可揭示大规模语言模型(LLM)潜在的推理缺陷。在本工作中,我们探讨了重新定义任务,即为众所周知的物理常数和计量单位指定替代值,要求模型作出相应响应。我们的发现表明,模型性能随规模而下降,同时其虚假自信心也随之上升。此外,尽管提示策略和响应格式等因素具有影响,但它们并不能防止LLM锚定于记忆中的值。

关键词

引用

@article{arxiv.2502.12821,
  title  = {Pitfalls of Scale: Investigating the Inverse Task of Redefinition in Large Language Models},
  author = {Elena Stringli and Maria Lymperaiou and Giorgos Filandrianos and Athanasios Voulodimos and Giorgos Stamou},
  journal= {arXiv preprint arXiv:2502.12821},
  year   = {2025}
}

备注

Accepted at Findings of ACL 2025