Deep Learning-Based Object Detection System for Vocal Cords in Laryngoscopy Images

Accurate localization of the vocal folds is critical for diagnostic and therapeutic applications in medical imaging. This paper developed a deep learning-based object detection system to address different tasks related to vocal fold detection. A two-step transfer learning approach was proposed for model training. The YOLOv8 was pre-trained with a large scale of high-speed video recordings available in a public dataset. Then, we fine-tuned the model using a 1:1 combination of the data from the public dataset and the low-frame rate (30 frames/sec) dataset collected from the hospital CGMH in Taiwan to fine-tune the final optimized glottis detection model. The results show that the recall and precision of ROI multiple bounding box predictions are 97.6% and 98.3%, respectively. In comparison, the recall and precision of single bounding box predictions are 97.5% and 100%, respectively. These results demonstrate the successful use of deep learning technology for vocal fold localization with superior performance. Our study provides valuable information for selecting appropriate object detection modalities in medical imaging applications for diagnosis and treatment planning in laryngology and otorhinolaryngology.

Paper

Full text

PDF

Deep Learning-Based Object Detection System for Vocal Cords in Laryngoscopy Images

OpenAlex · Speech Recognition and Synthesis · 2025

Abstract

Accurate localization of the vocal folds is critical for diagnostic and therapeutic applications in medical imaging. This paper developed a deep learning-based object detection system to address different tasks related to vocal fold detection. A two-step transfer learning approach was proposed for model training. The YOLOv8 was pre-trained with a large scale of high-speed video recordings available in a public dataset. Then, we fine-tuned the model using a 1:1 combination of the data from the public dataset and the low-frame rate (30 frames/sec) dataset collected from the hospital CGMH in Taiwan to fine-tune the final optimized glottis detection model. The results show that the recall and precision of ROI multiple bounding box predictions are 97.6% and 98.3%, respectively. In comparison, the recall and precision of single bounding box predictions are 97.5% and 100%, respectively. These results demonstrate the successful use of deep learning technology for vocal fold localization with superior performance. Our study provides valuable information for selecting appropriate object detection modalities in medical imaging applications for diagnosis and treatment planning in laryngology and otorhinolaryngology.

Similar papers

© 2026 NYSGPT2525 LLC