Assessing the Performance of the DINOv2 Self-supervised Learning Vision Transformer Model for the Segmentation of the Left Atrium from MRI Images
Accurate segmentation of the left atrium from pre-operative scans is required for diagnosing atrial fibrillation, treatment planning, intraoperative guidance, and supporting computer-assisted surgical interventions. While deep learning models are pivotal in medical image segmentation, they often require extensive manually annotated datasets. However, the emergence of foundation models trained on larger datasets has helped reduce this dependency, enhancing generalizability and robustness through transfer learning capabilities. In this work, we explore the out-of-the-box potential of DINOv2, a self-supervised learning vision transformer-based foundation model trained on natural images, by evaluating its performance when used for the left atrium (LA) segmentation task using MRI images. The challenges include the complex anatomical structures, thin myocardial walls of the left atrium, and limited annotated data, making it difficult to accurately segment the desired LA structures both prior to or during the image-guided intervention. We aim to demonstrate DINOv2’s ability to provide accurate and consistent segmentation in this specific context. We comprehensively evaluated the performance of DINOv2 for left atrial segmentation utilizing end-to-end fine-tuning and achieved, utilizing end-to-end finetuning, and achieved a mean Dice score of 87.1% and an Intersection over Union (IoU) of 79.2%. Our study included data-level few-shot learning across different dataset sizes and patient counts, consistently finding that DINOv2 outperforms all baseline models. Furthermore, these comparisons suggest that DINOv2 can perform well out-of-the-box to match the above instances in the medical domain and effectively adapt and generalize to MRI data, even with minimal fine-tuning and limited data. These findings highlight DINOv2’s potential as a competitive tool for cardiac segmentation, providing accurate results essential for pre-procedural planning and pre-operative applications. Our study aims to inform medical researchers about DINOv2’s potential for broader implementation in other medical imaging modalities.