English

USR: An Unsupervised and Reference Free Evaluation Metric for Dialog Generation

Computation and Language 2020-05-04 v1 Machine Learning

Abstract

The lack of meaningful automatic evaluation metrics for dialog has impeded open-domain dialog research. Standard language generation metrics have been shown to be ineffective for evaluating dialog models. To this end, this paper presents USR, an UnSupervised and Reference-free evaluation metric for dialog. USR is a reference-free metric that trains unsupervised models to measure several desirable qualities of dialog. USR is shown to strongly correlate with human judgment on both Topical-Chat (turn-level: 0.42, system-level: 1.0) and PersonaChat (turn-level: 0.48 and system-level: 1.0). USR additionally produces interpretable measures for several desirable properties of dialog.

Keywords

Cite

@article{arxiv.2005.00456,
  title  = {USR: An Unsupervised and Reference Free Evaluation Metric for Dialog Generation},
  author = {Shikib Mehri and Maxine Eskenazi},
  journal= {arXiv preprint arXiv:2005.00456},
  year   = {2020}
}

Comments

Accepted to ACL 2020 as long paper

R2 v1 2026-06-23T15:14:39.635Z