检测大语言模型与知识图谱间元语言分歧的基准测试
计算与语言
2025-02-06 v1 人工智能
摘要
评估大语言模型(LLMs)在支持知识图谱构建的事实提取等任务中的表现,通常涉及使用基于知识图谱(KG)的真实基准计算准确率指标。这些评估假设错误代表事实性分歧。然而,人类话语中经常出现元语言分歧,即主体之间的分歧不在于事实,而在于用于表达事实的语言的含义。鉴于使用LLMs进行自然语言处理与生成的复杂性,我们提出疑问:LLMs与KGs之间是否会发生元语言分歧?基于使用T-REx知识对齐数据集的调查,我们假设LLMs与KGs之间确实会发生元语言分歧,这可能与知识图谱工程的实践相关。我们提出了一个用于评估检测LLMs与KGs之间事实性分歧和元语言分歧的基准测试。该基准测试的初步概念验证已在Github上提供。
引用
@article{arxiv.2502.02896,
title = {A Benchmark for the Detection of Metalinguistic Disagreements between LLMs and Knowledge Graphs},
author = {Bradley P. Allen and Paul T. Groth},
journal= {arXiv preprint arXiv:2502.02896},
year = {2025}
}
备注
6 pages, 2 tables, to appear in Reham Alharbi, Jacopo de Berardinis, Paul Groth, Albert Mero\~no-Pe\~nuela, Elena Simperl, Valentina Tamma (eds.), ISWC 2024 Special Session on Harmonising Generative AI and Semantic Web Technologies. CEUR-WS.org (forthcoming), for associated code and data see https://github.com/bradleypallen/trex-metalinguistic-disagreement