Self-Training and Error Correction using Large Language Models for Medical Speech Recognition
In the healthcare sector, medical professionals must dedicate substantial time and effort to documentation, which directly impacts patient care and clinical decision-making. Automatic speech recognition (ASR) systems offer a potential solution to reduce the burden of these documentation tasks. However, conventional ASR systems often perform poorly in the medical domain due to the use of specialized terminology, as well as variations in accent and pronunciation. Fine-tuning ASR models with large amounts of medical speech data is challenging due to confidentiality concerns and limitations in recording conditions. This study explores two approaches to enhance the performance of ASR for medical speech. The first approach utilizes a self-training mechanism, where transcriptions generated by a baseline ASR model are used to train the model further. The pseudo-labels produced by the baseline model introduce greater data diversity, particularly when the model transcribes with high confidence. The second approach employs open-source large language models (LLMs) to select the most probable transcription from the fivebest hypotheses generated by the ASR model. We conducted experiments using the Whisper ASR model as the baseline and investigated the effectiveness of these two approaches. Our findings indicate that self-training reduced the word error rate (WER) by an absolute $1-2 \%$. However, the use of five-best hypothesis selection resulted in an increased WER, which we attribute to the limited medical knowledge of the open-source LLMs.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex