Towards Unified Dialogue System Evaluation: A Comprehensive Analysis of Current Evaluation Protocols
As conversational AI-based dialogue management has increasingly become a\ntrending topic, the need for a standardized and reliable evaluation procedure\ngrows even more pressing. The current state of affairs suggests various\nevaluation protocols to assess chat-oriented dialogue management systems,\nrendering it difficult to conduct fair comparative studies across different\napproaches and gain an insightful understanding of their values. To foster this\nresearch, a more robust evaluation protocol must be set in place. This paper\npresents a comprehensive synthesis of both automated and human evaluation\nmethods on dialogue systems, identifying their shortcomings while accumulating\nevidence towards the most effective evaluation dimensions. A total of 20 papers\nfrom the last two years are surveyed to analyze three types of evaluation\nprotocols: automated, static, and interactive. Finally, the evaluation\ndimensions used in these papers are compared against our expert evaluation on\nthe system-user dialogue data collected from the Alexa Prize 2020.\n
Paper
References (37)
Scroll for more · 25 remaining