Accurate localization of the vocal folds is critical for diagnostic and therapeutic applications in medical imaging. This paper developed a deep learning-based object detection system to address different tasks related to vocal fold detection. A two-step transfer learning approach was proposed for model training. The YOLOv8 was pre-trained with a large scale of high-speed video recordings available in a public dataset. Then, we fine-tuned the model using a 1:1 combination of the data from the public dataset and the low-frame rate (30 frames/sec) dataset collected from the hospital CGMH in Taiwan to fine-tune the final optimized glottis detection model. The results show that the recall and precision of ROI multiple bounding box predictions are 97.6% and 98.3%, respectively. In comparison, the recall and precision of single bounding box predictions are 97.5% and 100%, respectively. These results demonstrate the successful use of deep learning technology for vocal fold localization with superior performance. Our study provides valuable information for selecting appropriate object detection modalities in medical imaging applications for diagnosis and treatment planning in laryngology and otorhinolaryngology.
Paper
Full text
Deep Learning-Based Object Detection System for Vocal Cords in Laryngoscopy Images
OpenAlex · Speech Recognition and Synthesis · 2025
Abstract
Accurate localization of the vocal folds is critical for diagnostic and therapeutic applications in medical imaging. This paper developed a deep learning-based object detection system to address different tasks related to vocal fold detection. A two-step transfer learning approach was proposed for model training. The YOLOv8 was pre-trained with a large scale of high-speed video recordings available in a public dataset. Then, we fine-tuned the model using a 1:1 combination of the data from the public dataset and the low-frame rate (30 frames/sec) dataset collected from the hospital CGMH in Taiwan to fine-tune the final optimized glottis detection model. The results show that the recall and precision of ROI multiple bounding box predictions are 97.6% and 98.3%, respectively. In comparison, the recall and precision of single bounding box predictions are 97.5% and 100%, respectively. These results demonstrate the successful use of deep learning technology for vocal fold localization with superior performance. Our study provides valuable information for selecting appropriate object detection modalities in medical imaging applications for diagnosis and treatment planning in laryngology and otorhinolaryngology.