中文

LLM 作为丹麦难民决定中的可信度评估标注器:评估分类性能及超越聚合指标的错误

计算与语言 2026-05-14 v1 人工智能

摘要

面向不可见语言和专业领域的大语言模型 (LLM) 用于自动化文本标注,其有效性仍未得到充分探索,其中类定义需要细微的 expert 理解。我们调查 LLM 用于 novel legal NLP 任务的标注:识别 asylum decision 文本中可信度评估的存在及情感。我们引入 RAB-Cred,一个丹麦文本分类数据集,包含高质量的 expert 标注以及有价值的元数据,如标注者信心和难民案件结果。我们对 21 个 open-weight 模型和 30 套 system-user prompt 组合进行基准测试,系统性评估 model 和 prompt 选择对 zero-shot 和 few-shot 分类的影响。我们聚焦于顶级模型和 prompt 所犯的错误,探讨 error 在 LLM 之间的一致性、类间混淆、与人类信心和样本难度及 LLM 错误的严重程度之间的相关性。我们的结果确认 LLM 用于成本效益高的难民决定标注的潜力,但突显 LLM 标注器的不完美与不一致性,以及超越单个任意所选模型预测的必要性。RAB-Cred 数据集和代码已公开于 https://github.com/glhr/RAB-Cred。

关键词

引用

@article{arxiv.2605.13412,
  title  = {LLMs as annotators of credibility assessment in Danish asylum decisions: evaluating classification performance and errors beyond aggregated metrics},
  author = {Galadrielle Humblot-Renaux and Mohammad N. S. Jahromi and Rohat Bakuri-Jørgensen and Marieke Anne Heyl and Asta S. Stage Jarlner and Maria Vlachou and Anna Murphy Høgenhaug and Desmond Elliott and Thomas Gammeltoft-Hansen and Thomas B. Moeslund},
  journal= {arXiv preprint arXiv:2605.13412},
  year   = {2026}
}

备注

Accepted at the 20th Linguistic Annotation Workshop (LAW XX), co-located with ACL 2026 (https://sigann.github.io/LAW-XX-2026/)