Speech Recognition and Machine Learning: Revolutionizing English Language Oral Assessments

The assessment of English oral skills is an important part of language acquisition process; historically implemented through human assessment, which has drawbacks including subjectivity and inability to scale. Current approaches are based on rule definition or on classical machine learning techniques including Hidden Markov Models and Recurrent Neural Networks that cannot adequately account for variations in accents, contextual effects, or long distance connections in spoken language. To overcome these drawbacks, this research proposes the Attention-based Transformer Model to assign the score to the oral assessment without human intervention, using the self-attention mechanism to improve the speech recognition rate. The research approach includes using a fine-tuned transformer model that transcribes speech, evaluates pronunciation, fluency and syntactic complexity. Learner oral performance is assessed relative to the reference text by comparing transcribed speech with text, delivering real-time detailed feedback. For the experiments the approach was used the dataset that involves the various speakers of the English language and it demonstrated a well increase of the recognition accuracy on 25% more than traditional model in terms of WER. The feedback loop of the system was another feature that improved learning results as it provided individual recommendations. In this paper, the capabilities of transformer models are exemplified in an attempt to redefine language tests as comprehensive, efficient, and affordable for learners globally.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC