Automatic Assessment Of Spoken English Proficiency Based On Multimodal & Multitask Transformers
This paper describes technology developed to automatically grade students on their English spontaneous spoken language proficiency with common european framework of reference for languages (CEFR) level.Our automated assessment system contains two tasks: elicited imitation and spontaneous speech assessment.Spontaneous speech assessment is a challenging task that requires evaluating various aspects of speech quality, content, and coherence.In this paper, we propose a multimodal and multitask transformer model that leverages both audio and text features to perform three tasks: scoring, coherence modeling, and prompt relevancy scoring.Our model uses a fusion of multiple features and multiple modality attention to capture the interactions between audio and text modalities and learn from different sources of information.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex