Accurate skin-lesion segmentation remains a key technical challenge for computer-aided diagnosis of skin cancer. Convolutional neural networks, while effective, are constrained by limited receptive fields and thus struggle to model long-range dependencies. Vision Transformers capture global context, yet their quadratic complexity and large parameter budgets hinder use on the small-sample medical datasets common in dermatology. We introduce the MedLiteNet, a lightweight CNN–Transformer hybrid tailored for dermoscopic segmentation that achieves high precision through hierarchical feature extraction and multiscale context aggregation. The encoder stacks depthwise mobile inverted bottleneck blocks to curb computation, inserts a bottleneck-level cross-scale token-mixing unit to exchange information between resolutions, and embeds a boundary-aware self-attention module to sharpen lesion contours. On the ISIC 2018 benchmark, a single MedLiteNet model attains a Dice score of <inline-formula><tex-math notation="LaTeX">$0.897 \pm 0.010$</tex-math></inline-formula> and an IoU of <inline-formula><tex-math notation="LaTeX">$0.821 \pm 0.015$</tex-math></inline-formula> with fewer than <inline-formula><tex-math notation="LaTeX">$3.3\,\mathrm{M}$</tex-math></inline-formula> parameters. A performance-weighted ensemble of three complementary variants raises accuracy to <inline-formula><tex-math notation="LaTeX">$0.904 \pm 0.012$</tex-math></inline-formula> Dice and <inline-formula><tex-math notation="LaTeX">$0.830 \pm 0.018$</tex-math></inline-formula> IoU while keeping the total parameter count below <inline-formula><tex-math notation="LaTeX">$10\,\mathrm{M}$</tex-math></inline-formula>—over 90% smaller than Vision-Transformer backbones. Qualitative results confirm superiority on irregular borders, low-contrast regions and multiscale lesions, indicating MedLiteNet’s suitability for real-time, resource-aware computer-aided dermatology.
Paper
References (30)
Scroll for more · 18 remaining