Combining Speech Recognition and Machine Learning for Context-Aware English Pronunciation Assessment

The use of pronunciation assessment is vital for the improvement of the aspect of English language for learners learning it like foreigners. The existing approaches are based upon the conceptual representations and the scoring functions which necessarily focus on simple comparison, and typically fail to offer adaptive feedback in response to contextual changes. Current methods for building up LSTM models fail to address the complex relationship between phonemes and the patterns in which they appear, thus producing inaccurate evaluations. This research work therefore recommended the usage of Bidirectional Long Short-Term Memory (Bi-Lstm) networks and an Attention Mechanism, for enhancing the first approach’s precision and readability of the pronunciation evaluation. In the Bi-Lstm network, considered forward phonetic sequences and backward phonetic sequences so that temporal dependencies are also learned, and the AM forced the model to pay more attention to any mispronunciation occurrence. The proposed system takes advantage of speech recognition to get the phonetic features, and uses the Bi-Lstm-Attention model to give the quality of pronunciation scores. Native vs nonnative English speakers experimentations show positive results the proposed model achieves better accuracy coupled with better error detecting capability than the LSTM Based models. This context-aware framework gives detailed, accurate, real-time commentary, and it is a very useful tool which language learners as well as educators can use to pretty precisely determine error-ridden or better: error-free pronunciation patterns.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC