English

Unsupervised Evaluation of Interactive Dialog with DialoGPT

Computation and Language 2020-06-25 v1 Artificial Intelligence Human-Computer Interaction

Abstract

It is important to define meaningful and interpretable automatic evaluation metrics for open-domain dialog research. Standard language generation metrics have been shown to be ineffective for dialog. This paper introduces the FED metric (fine-grained evaluation of dialog), an automatic evaluation metric which uses DialoGPT, without any fine-tuning or supervision. It also introduces the FED dataset which is constructed by annotating a set of human-system and human-human conversations with eighteen fine-grained dialog qualities. The FED metric (1) does not rely on a ground-truth response, (2) does not require training data and (3) measures fine-grained dialog qualities at both the turn and whole dialog levels. FED attains moderate to strong correlation with human judgement at both levels.

Keywords

Cite

@article{arxiv.2006.12719,
  title  = {Unsupervised Evaluation of Interactive Dialog with DialoGPT},
  author = {Shikib Mehri and Maxine Eskenazi},
  journal= {arXiv preprint arXiv:2006.12719},
  year   = {2020}
}

Comments

Published at to SIGdial 2020

R2 v1 2026-06-23T16:32:33.884Z