English

Why We Need New Evaluation Metrics for NLG

Computation and Language 2017-09-18 v1

Abstract

The majority of NLG evaluation relies on automatic metrics, such as BLEU . In this paper, we motivate the need for novel, system- and data-independent automatic evaluation methods: We investigate a wide range of metrics, including state-of-the-art word-based and novel grammar-based ones, and demonstrate that they only weakly reflect human judgements of system outputs as generated by data-driven, end-to-end NLG. We also show that metric performance is data- and system-specific. Nevertheless, our results also suggest that automatic metrics perform reliably at system-level and can support system development by finding cases where a system performs poorly.

Keywords

Cite

@article{arxiv.1707.06875,
  title  = {Why We Need New Evaluation Metrics for NLG},
  author = {Jekaterina Novikova and Ondřej Dušek and Amanda Cercas Curry and Verena Rieser},
  journal= {arXiv preprint arXiv:1707.06875},
  year   = {2017}
}

Comments

accepted to EMNLP 2017

R2 v1 2026-06-22T20:53:54.902Z