English

Multilingual Fine-Grained News Headline Hallucination Detection

Computation and Language 2024-07-24 v1

Abstract

The popularity of automated news headline generation has surged with advancements in pre-trained language models. However, these models often suffer from the ``hallucination'' problem, where the generated headline is not fully supported by its source article. Efforts to address this issue have predominantly focused on English, using over-simplistic classification schemes that overlook nuanced hallucination types. In this study, we introduce the first multilingual, fine-grained news headline hallucination detection dataset that contains over 11 thousand pairs in 5 languages, each annotated with detailed hallucination types by experts. We conduct extensive experiments on this dataset under two settings. First, we implement several supervised fine-tuning approaches as preparatory solutions and demonstrate this dataset's challenges and utilities. Second, we test various large language models' in-context learning abilities and propose two novel techniques, language-dependent demonstration selection and coarse-to-fine prompting, to boost the few-shot hallucination detection performance in terms of the example-F1 metric. We release this dataset to foster further research in multilingual, fine-grained headline hallucination detection.

Keywords

Cite

@article{arxiv.2407.15975,
  title  = {Multilingual Fine-Grained News Headline Hallucination Detection},
  author = {Jiaming Shen and Tianqi Liu and Jialu Liu and Zhen Qin and Jay Pavagadhi and Simon Baumgartner and Michael Bendersky},
  journal= {arXiv preprint arXiv:2407.15975},
  year   = {2024}
}
R2 v1 2026-06-28T17:50:04.818Z