On the Challenges of Evaluating Compositional Explanations in Multi-Hop Inference: Relevance, Completeness, and Expert Ratings
Building compositional explanations requires models to combine two or more\nfacts that, together, describe why the answer to a question is correct.\nTypically, these "multi-hop" explanations are evaluated relative to one (or a\nsmall number of) gold explanations. In this work, we show these evaluations\nsubstantially underestimate model performance, both in terms of the relevance\nof included facts, as well as the completeness of model-generated explanations,\nbecause models regularly discover and produce valid explanations that are\ndifferent than gold explanations. To address this, we construct a large corpus\nof 126k domain-expert (science teacher) relevance ratings that augment a corpus\nof explanations to standardized science exam questions, discovering 80k\nadditional relevant facts not rated as gold. We build three strong models based\non different methodologies (generation, ranking, and schemas), and empirically\nshow that while expert-augmented ratings provide better estimates of\nexplanation quality, both original (gold) and expert-augmented automatic\nevaluations still substantially underestimate performance by up to 36% when\ncompared with full manual expert judgements, with different models being\ndisproportionately affected. This poses a significant methodological challenge\nto accurately evaluating explanations produced by compositional reasoning\nmodels.\n
Paper
References (41)
Scroll for more · 29 remaining