ASSET: A Dataset for Tuning and Evaluation of Sentence Simplification Models with Multiple Rewriting Transformations
In order to simplify a sentence, human editors perform multiple rewriting\ntransformations: they split it into several shorter sentences, paraphrase words\n(i.e. replacing complex words or phrases by simpler synonyms), reorder\ncomponents, and/or delete information deemed unnecessary. Despite these varied\nrange of possible text alterations, current models for automatic sentence\nsimplification are evaluated using datasets that are focused on a single\ntransformation, such as lexical paraphrasing or splitting. This makes it\nimpossible to understand the ability of simplification models in more realistic\nsettings. To alleviate this limitation, this paper introduces ASSET, a new\ndataset for assessing sentence simplification in English. ASSET is a\ncrowdsourced multi-reference corpus where each simplification was produced by\nexecuting several rewriting transformations. Through quantitative and\nqualitative experiments, we show that simplifications in ASSET are better at\ncapturing characteristics of simplicity when compared to other standard\nevaluation datasets for the task. Furthermore, we motivate the need for\ndeveloping better methods for automatic evaluation using ASSET, since we show\nthat current popular metrics may not be suitable when multiple simplification\ntransformations are performed.\n