English

Towards Unified Dialogue System Evaluation: A Comprehensive Analysis of Current Evaluation Protocols

Computation and Language 2020-06-12 v1

Abstract

As conversational AI-based dialogue management has increasingly become a trending topic, the need for a standardized and reliable evaluation procedure grows even more pressing. The current state of affairs suggests various evaluation protocols to assess chat-oriented dialogue management systems, rendering it difficult to conduct fair comparative studies across different approaches and gain an insightful understanding of their values. To foster this research, a more robust evaluation protocol must be set in place. This paper presents a comprehensive synthesis of both automated and human evaluation methods on dialogue systems, identifying their shortcomings while accumulating evidence towards the most effective evaluation dimensions. A total of 20 papers from the last two years are surveyed to analyze three types of evaluation protocols: automated, static, and interactive. Finally, the evaluation dimensions used in these papers are compared against our expert evaluation on the system-user dialogue data collected from the Alexa Prize 2020.

Keywords

Cite

@article{arxiv.2006.06110,
  title  = {Towards Unified Dialogue System Evaluation: A Comprehensive Analysis of Current Evaluation Protocols},
  author = {Sarah E. Finch and Jinho D. Choi},
  journal= {arXiv preprint arXiv:2006.06110},
  year   = {2020}
}

Comments

Accepted by SIGDIAL 2020

R2 v1 2026-06-23T16:13:18.804Z