知识编辑方法能多好地处理难以理解的知识吗?
摘要
大型语言模型(LLM)展现出惊人的能力,但在训练后更新其知识仍是一个关键挑战。虽然最近的模型编辑技术,如单秩模型编辑(Rank-One Model Editing, ROME)显示出前景,但其效果可能因所编辑知识的性质而有所不同。我们提出“难以理解性”(perplexingness)的概念:即新知识与 LLM 所学习概念层次结构和范畴关系冲突的程度。例如,将“British Shorthair is a kind of cat”编辑为“British Shorthair is a kind of dog”属于低难以理解性编辑,因其位于同一分类层级内;而将“A cat is a kind of animal”编辑为“A cat is a kind of plant”则属于高难以理解性编辑,因其违反了基本的范畴边界。为系统地研究这一现象,我们构建了 HierarchyData 数据集,包含 99 个跨越多样类别的同义词-原义词配对。通过在三个模型和四种编辑方法下的控制实验,我们证明了新知识的难以理解性与知识编辑效果之间存在强烈的负相关。我们的分析表明,涉及更抽象概念(上位词)的编辑通常难度更大,且比具体对象(下位词)更难以修改。这些发现凸显了 LLM 知识编辑中的一个根本挑战:当新事实与 LLM 所学概念层次结构相矛盾时,便更难可靠地编码该知识。
引用
@article{arxiv.2406.17253,
title = {How Well Can Knowledge Edit Methods Edit Perplexing Knowledge?},
author = {Huaizhi Ge and Frank Rudzicz and Zining Zhu},
journal= {arXiv preprint arXiv:2406.17253},
year = {2025}
}
备注
A previous version of this document contained a hidden prompt entered by Z Zhu without knowledge of -- or consent by -- his co-authors. This version does not contain the prompt