Enhanced Machine Learning Framework for English Speaking Assessment: Addressing Accent and Speech Variability Through BiLSTM-Transformer Architecture

Automated speech assessment still faces the challenge of accurate recognition of diverse English accents. However, benefitting from temporal and contextual variations in speech data, traditional models that we use such as CNNs and HMMs are not sufficiently able to do so. This study overcomes the above challenges with a hybrid BiLSTM-Transformer model, where BiLSTM handles sequential dependencies in the speech signal, and the Transformer is for the long-range contextual relationship. Together, these generate a more complex view of accent and speech variability. It uses Mel Frequency Cepstral Coefficients as features and is trained on the CSTR VCTK Corpus of English that includes English accents. The proposed model carries an accuracy of 95.6%, precision of 93.4%, recall of 92.5%, and 92.9% as F1 score, which corresponds to a recognition error of 4.4%. This quantitatively shows a considerable improvement compared to traditional models. The model is in fact built with such high responsiveness that it is well suited for live situations, i.e. automated interviews or language learning tools, with an average latency in the order of 180 ms per utterance. Also, the system demonstrates state of the art performance between speech and vegetative users, demonstrating its robustness and generalizability. Based on this research, a reliable and accurate solution for English accent classification and speech emotion recognition is established through the use of BiLSTM-Transformer architecture.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC