Learning to Explain: Datasets and Models for Identifying Valid Reasoning Chains in Multihop Question-Answering

Despite the rapid progress in multihop question-answering (QA), models still\nhave trouble explaining why an answer is correct, with limited explanation\ntraining data available to learn from. To address this, we introduce three\nexplanation datasets in which explanations formed from corpus facts are\nannotated. Our first dataset, eQASC, contains over 98K explanation annotations\nfor the multihop question answering dataset QASC, and is the first that\nannotates multiple candidate explanations for each answer. The second dataset\neQASC-perturbed is constructed by crowd-sourcing perturbations (while\npreserving their validity) of a subset of explanations in QASC, to test\nconsistency and generalization of explanation prediction models. The third\ndataset eOBQA is constructed by adding explanation annotations to the OBQA\ndataset to test generalization of models trained on eQASC. We show that this\ndata can be used to significantly improve explanation quality (+14% absolute F1\nover a strong retrieval baseline) using a BERT-based classifier, but still\nbehind the upper bound, offering a new challenge for future research. We also\nexplore a delexicalized chain representation in which repeated noun phrases are\nreplaced by variables, thus turning them into generalized reasoning chains (for\nexample: "X is a Y" AND "Y has Z" IMPLIES "X has Z"). We find that generalized\nchains maintain performance while also being more robust to certain\nperturbations.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC