中文

OpenFactCheck:构建与基准测试定制化事实核查系统及评估声明与 LLM 的事实性

计算与语言 2025-10-30 v3

摘要

大型语言模型(LLMs)在各种实际应用中的日益广泛使用,需要相应的机制来验证其输出的事实准确性。评估开放域中自由格式回复的事实性存在困难。此外,不同论文使用各异的评估基准和测量方法,导致难以进行比较并阻碍了未来的进展。为了缓解这些问题,我们提出了 OpenFactCheck,这是一个用于构建定制化自动事实核查系统、对系统准确性进行基准测试、评估 LLM 事实性以及验证文档中声明的统一框架。OpenFactCheck 由三个模块组成: CUSTCHECKER 允许用户轻松定制自动事实核查器,并验证文档和声明的事实正确性; LLMEVAL 是一个统一的评估框架,从多个角度公平地评估 LLM 的事实性能力; CHECKEREVAL 是一个可扩展的解决方案,利用人工标注数据集来衡量自动事实核查器验证结果的可靠性。数据和代码已在 https://github.com/yuxiaw/openfactcheck 公开。

关键词

引用

@article{arxiv.2405.05583,
  title  = {OpenFactCheck: Building, Benchmarking Customized Fact-Checking Systems and Evaluating the Factuality of Claims and LLMs},
  author = {Yuxia Wang and Minghan Wang and Hasan Iqbal and Georgi Georgiev and Jiahui Geng and Preslav Nakov},
  journal= {arXiv preprint arXiv:2405.05583},
  year   = {2025}
}

备注

23 pages, 8 tables, 11 figures, Published In Proceedings of the 31st International Conference on Computational Linguistics 2025