Are pre-trained text representations useful for multilingual and multi-dimensional language proficiency modeling?
Development of language proficiency models for non-native learners has been\nan active area of interest in NLP research for the past few years. Although\nlanguage proficiency is multidimensional in nature, existing research typically\nconsiders a single "overall proficiency" while building models. Further,\nexisting approaches also considers only one language at a time. This paper\ndescribes our experiments and observations about the role of pre-trained and\nfine-tuned multilingual embeddings in performing multi-dimensional,\nmultilingual language proficiency classification. We report experiments with\nthree languages -- German, Italian, and Czech -- and model seven dimensions of\nproficiency ranging from vocabulary control to sociolinguistic appropriateness.\nOur results indicate that while fine-tuned embeddings are useful for\nmultilingual proficiency modeling, none of the features achieve consistently\nbest performance for all dimensions of language proficiency. All code, data and\nrelated supplementary material can be found at:\nhttps://github.com/nishkalavallabhi/MultidimCEFRScoring.\n