UPB at SemEval-2021 Task 1: Combining Deep Learning and Hand-Crafted Features for Lexical Complexity Prediction
Reading is a complex process which requires proper understanding of texts in\norder to create coherent mental representations. However, comprehension\nproblems may arise due to hard-to-understand sections, which can prove\ntroublesome for readers, while accounting for their specific language skills.\nAs such, steps towards simplifying these sections can be performed, by\naccurately identifying and evaluating difficult structures. In this paper, we\ndescribe our approach for the SemEval-2021 Task 1: Lexical Complexity\nPrediction competition that consists of a mixture of advanced NLP techniques,\nnamely Transformer-based language models, pre-trained word embeddings, Graph\nConvolutional Networks, Capsule Networks, as well as a series of hand-crafted\ntextual complexity features. Our models are applicable on both subtasks and\nachieve good performance results, with a MAE below 0.07 and a Person\ncorrelation of .73 for single word identification, as well as a MAE below 0.08\nand a Person correlation of .79 for multiple word targets. Our results are just\n5.46% and 6.5% lower than the top scores obtained in the competition on the\nfirst and the second subtasks, respectively.\n