Understanding how language learners develop linguistic complexity remains a central concern in second language research. This study introduces a transformer-based natural language processing framework designed to analyse learner texts, measure linguistic sophistication, and estimate proficiency in a transparent and interpretable manner. The approach integrates lexical and syntactic indicators with contextual embeddings from multilingual BERT, creating a unified representation that captures both surface and deep language patterns. Experiments on learner essays and spoken transcripts demonstrate that combining handcrafted linguistic features with transformer embeddings enhances both predictive accuracy and interpretability. The findings highlight the potential of NLP-driven models to support data-informed, pedagogically meaningful assessment of language proficiency.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex