English

HaRiM$^+$: Evaluating Summary Quality with Hallucination Risk

Computation and Language 2022-11-28 v2

Abstract

One of the challenges of developing a summarization model arises from the difficulty in measuring the factual inconsistency of the generated text. In this study, we reinterpret the decoder overconfidence-regularizing objective suggested in (Miao et al., 2021) as a hallucination risk measurement to better estimate the quality of generated summaries. We propose a reference-free metric, HaRiM+, which only requires an off-the-shelf summarization model to compute the hallucination risk based on token likelihoods. Deploying it requires no additional training of models or ad-hoc modules, which usually need alignment to human judgments. For summary-quality estimation, HaRiM+ records state-of-the-art correlation to human judgment on three summary-quality annotation sets: FRANK, QAGS, and SummEval. We hope that our work, which merits the use of summarization models, facilitates the progress of both automated evaluation and generation of summary.

Keywords

Cite

@article{arxiv.2211.12118,
  title  = {HaRiM$^+$: Evaluating Summary Quality with Hallucination Risk},
  author = {Seonil Son and Junsoo Park and Jeong-in Hwang and Junghwa Lee and Hyungjong Noh and Yeonsoo Lee},
  journal= {arXiv preprint arXiv:2211.12118},
  year   = {2022}
}

Comments

9 pages (+ 21 pages of Appendix), AACL 2022

R2 v1 2026-06-28T06:34:21.122Z