Automated English speaking assessment with large language models: A framework for multi-dimensional scoring and feedback generation
Traditional automated English speaking assessment systems are limited in their ability to provide meaningful improvement guidance to learners, as they typically focus solely on generating overall scores. To address this limitation, this study proposes an integrated LLM-based assessment model that simultaneously performs quantitative score prediction and qualitative feedback generation for comprehensive English speaking evaluation. The model combines Whisper, BEATs, and QFormer for multidimensional audio feature extraction, utilizes ChatGPT-generated training data for Llama-based instruction tuning, and employs large language models to predict scores across four domains (task completion, delivery, accuracy, appropriateness) while generating specific feedback and corrections. Experimental results demonstrate reasonable correlations with human evaluators (Pearson correlation coefficients ranging from 0.730 to 0.789) and feedback quality with average scores above 4.0 points in all evaluation categories as validated by LLM-as-a-Judge methodology.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex