Do Large Language Models Encode Second-Language Writing Proficiency? A CALF-Based Perspective
Abstract This study investigates how Large Language Models (LLMs) encode second language (L2) writing proficiency distinctions compared to human learners, focusing on the structural alignment between synthetic outputs and human developmental patterns. We analyzed CEFR-graded Write & Improve 2024 learner essays and matched LLM generations (A2–C1) using statistical models controlling for prompt and text length. With a CALF feature set (Complexity, Accuracy, Lexical Complexity, Fluency), level explained 2.9 % of variance in human texts ( R 2 = 0.029) but 10.6 % in pooled LLM outputs ( R 2 = 0.106), indicating clearer A2–C1 features in LLM-generated texts overall. LLM level effects were robust ( R 2 range = 0.038–0.454), with sharper differences between neighboring CEFR bands than in learner essays (e.g., A2–B1 distance, D 2 = 4.81 versus 1.07). Prompt-wise analyses found effects for all 10 prompts in the LLM set but only 2/10 in the learner set, with much greater separation than in the learner cohort. These results support CALF and CEFR construct validity: with prompt and length controlled, A2–C1 texts show ordered CALF gradients, especially for LLM texts across prompts. The sharper LLM separation likely reflects idealized, low-variance production; therefore, CALF captures core structural proficiency (syntax, lexis, accuracy, fluency) but not discourse-pragmatic qualities or human variability. The analyses provide level-calibrated evidence that LLM texts show clearer distinctions than learners’, under identical prompts, whose development is gradual and overlapping. This positions LLMs as both tool (exemplars, rubric calibration) and challenge (assessment validity and fairness), while offering SLA researchers insight into how proficiency constructs are encoded in human versus model-based writing.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex