中文

量化与预测graded human ratings中的 disagreement

计算与语言 2026-05-05 v1

摘要

人类标注员之间并非总是达成一致,这种不一致在许多标注任务中是内在的。然而,给定任务中并非所有实例都引发相同的意见分歧。本文调查了对不当语言(包括冒犯语言、仇恨言论和有害语言感知)中graded human ratings的annotation variation patterns。我们检验了annotation disagreement的程度是否可由文本特征预测。我们进一步提出Opposition Index这一指标,用于量化标注员对特定项目的观点对立程度,并调查具有潜在对立人类意见的实例的可预测性。我们的结果显示,估计值与观察到的annotation variance之间存在中等正相关。我们发现两种方法在variance prediction方面的性能相当:直接预测方差值以及从预测的annotation分布中估计它。我们的结果表明,具有高Opposition Index值的元素更难预测,常被模型低估。

关键词

引用

@article{arxiv.2605.01168,
  title  = {Quantifying and Predicting Disagreement in Graded Human Ratings},
  author = {Leixin Zhang and Çağrı Çöltekin},
  journal= {arXiv preprint arXiv:2605.01168},
  year   = {2026}
}

备注

Accepted by the 5th Workshop on Perspectivist Approaches to NLP at LREC