The skin, the largest organ of the human body, is vulnerable to numerous pathological conditions collectively referred to as skin lesions, encompassing a wide spectrum of dermatoses. Diagnosing these lesions remains challenging for medical practitioners due to their subtle visual differences, many of which are imperceptible to the naked eye. While not all lesions are malignant, some serve as early indicators of serious diseases such as skin cancer, emphasizing the urgent need for accurate and timely diagnostic tools. This study advances dermatological diagnostics by curating a comprehensive and balanced dataset containing 9360 dermoscopic and clinical images across 39 lesion categories, synthesized from five publicly available datasets. Five state‐of‐the‐art deep learning architectures—MobileNetV2, Xception, InceptionV3, EfficientNetB1, and Vision Transformer (ViT)—were systematically evaluated on this dataset. To enhance model precision and robustness, Efficient Channel Attention (ECA) and Convolutional Block Attention Module (CBAM) mechanisms were integrated into these architectures. Extensive evaluation across multiple performance metrics demonstrated that the Vision Transformer with CBAM achieved the best results, with 93.46% accuracy, 94% precision, 93% recall, 93% F1‐score, and 93.67% specificity. These findings highlight the effectiveness of attention‐guided Vision Transformers in addressing complex, large‐scale, multi‐class skin lesion classification. By combining dataset diversity with advanced attention mechanisms, the proposed framework provides a reliable and interpretable tool to assist medical professionals in accurate and efficient lesion diagnosis, thereby contributing to improved clinical decision‐making and patient outcomes.
Paper
References (59)
Scroll for more · 38 remaining