As transparency becomes key for robotics and AI, it will be necessary to\nevaluate the methods through which transparency is provided, including\nautomatically generated natural language (NL) explanations. Here, we explore\nparallels between the generation of such explanations and the much-studied\nfield of evaluation of Natural Language Generation (NLG). Specifically, we\ninvestigate which of the NLG evaluation measures map well to explanations. We\npresent the ExBAN corpus: a crowd-sourced corpus of NL explanations for\nBayesian Networks. We run correlations comparing human subjective ratings with\nNLG automatic measures. We find that embedding-based automatic NLG evaluation\nmethods, such as BERTScore and BLEURT, have a higher correlation with human\nratings, compared to word-overlap metrics, such as BLEU and ROUGE. This work\nhas implications for Explainable AI and transparent robotic and autonomous\nsystems.\n
Paper
References (85)
Scroll for more · 38 remaining