English

Evaluating the Efficacy of Summarization Evaluation across Languages

Computation and Language 2021-06-04 v1

Abstract

While automatic summarization evaluation methods developed for English are routinely applied to other languages, this is the first attempt to systematically quantify their panlinguistic efficacy. We take a summarization corpus for eight different languages, and manually annotate generated summaries for focus (precision) and coverage (recall). Based on this, we evaluate 19 summarization evaluation metrics, and find that using multilingual BERT within BERTScore performs well across all languages, at a level above that for English.

Keywords

Cite

@article{arxiv.2106.01478,
  title  = {Evaluating the Efficacy of Summarization Evaluation across Languages},
  author = {Fajri Koto and Jey Han Lau and Timothy Baldwin},
  journal= {arXiv preprint arXiv:2106.01478},
  year   = {2021}
}

Comments

Findings of ACL 2021

R2 v1 2026-06-24T02:46:24.708Z