LCP-RIT at SemEval-2021 Task 1: Exploring Linguistic Features for Lexical Complexity Prediction
This paper describes team LCP-RIT's submission to the SemEval-2021 Task 1:\nLexical Complexity Prediction (LCP). The task organizers provided participants\nwith an augmented version of CompLex (Shardlow et al., 2020), an English\nmulti-domain dataset in which words in context were annotated with respect to\ntheir complexity using a five point Likert scale. Our system uses logistic\nregression and a wide range of linguistic features (e.g. psycholinguistic\nfeatures, n-grams, word frequency, POS tags) to predict the complexity of\nsingle words in this dataset. We analyze the impact of different linguistic\nfeatures in the classification performance and we evaluate the results in terms\nof mean absolute error, mean squared error, Pearson correlation, and Spearman\ncorrelation.\n