Enhancing Automatic Speech Recognitionwith Contextual Understanding using Natural Language Processing
Automatic speech recognition (ASR) has advanced from responding to limited sound to fluently understanding natural language. Used in voice search, virtual assistants, and speech-to-text systems to enhance user experience and productivity. Began with basic sound recognition and evolved into comprehensive language comprehension. Despite significant advancements in automatic speech recognition (ASR) technology, existing systems often struggle to accurately transcribe spoken language in context where semantic nuances and contextual cues play a crucial role. The problem arises from the inherent limitations of conventional ASR approaches to comprehensively understand and intercept the contextual information efficiently, resulting in inaccuracies misinterpretations and errors in transcriptions, especially in scenarios involving ambiguous or context dependent speech. Incorporating Natural Language Processing (NLP) techniques into ASR systems presents a promising avenue to address this challenge.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex