A Framework for Evaluation of Machine Reading Comprehension Gold Standards

Machine Reading Comprehension (MRC) is the task of answering a question over\na paragraph of text. While neural MRC systems gain popularity and achieve\nnoticeable performance, issues are being raised with the methodology used to\nestablish their performance, particularly concerning the data design of gold\nstandards that are used to evaluate them. There is but a limited understanding\nof the challenges present in this data, which makes it hard to draw comparisons\nand formulate reliable hypotheses. As a first step towards alleviating the\nproblem, this paper proposes a unifying framework to systematically investigate\nthe present linguistic features, required reasoning and background knowledge\nand factual correctness on one hand, and the presence of lexical cues as a\nlower bound for the requirement of understanding on the other hand. We propose\na qualitative annotation schema for the first and a set of approximative\nmetrics for the latter. In a first application of the framework, we analyse\nmodern MRC gold standards and present our findings: the absence of features\nthat contribute towards lexical ambiguity, the varying factual correctness of\nthe expected answers and the presence of lexical cues, all of which potentially\nlower the reading comprehension complexity and quality of the evaluation data.\n

Paper

References (50)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC