Improving Automated Assessment of English Spoken Discourse Using Transfer Learning Models-Wav2Vec 2.0
Automated Assessment of English Spoken Discourse is invaluable tool used in the process of learning foreign languages, teaching and communication as it is capable of delivering a large number of assessments of spoken language without significant human interferences. However, common approaches may yield some issues in terms of accuracy, methods' applicability to a variety of accents and languages, and the scalability of the techniques. Conventional techniques including HMM-GMM and contemporary DL techniques like the CNN-RNN triangulate unsatisfactory results due to the inability to capture more intricate spoken discourse since they are unable to pick tune from the din. In response to these difficulties, this research proposal presents the following framework using Wav2Vec 2.0, a self-supervised learning model involved in speech signals. The proposed method makes use of Wav2Vec 2.0 to capture rich representations directly from the raw input data to minimize dependence on labelled data and improve robustness to different patterns of speech. Applicable in Python, the model is fine-tuned in spoken discourse datasets and is assessed based on primary metrics of performance; test performance: 99.1%. These results are much better than traditional and baseline methods and clearly demonstrate the high stability and potential of the model. This way, the study that addresses prior limitations and offers a stable solution contributes to the development of the automated speech assessment. Further work that could be suggested is the selection of multimodal integration approaches and real time as a way to increase more the scores of the different models and the flexibility of the system that has been proposes in education and work environments.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex