The shape of defects on steel surfaces is highly variable and training samples are limited, making it a significant challenge to transfer a high-performance pretrained vision language model to steel surface defect detection. Therefore, a Multi-level Supervised Vision Language Model based Steel Surface Defect Detection method MLS-VLM is proposed in this paper. MLS-VLM delves deeply into the extraction of profound features from limited samples with three levels of training: supervised contrast training from labeled areas and the entire image, as well as self-supervised contrast learning from Region Proposals. MLS-VLM can be rapidly transferred to two-stage object detector. Experimental results demonstrate that, compared to traditional object detection methods, MLS-VLM achieves 5.68~8.37 mAP improvement on three benchmark object detectors.
Paper
Full text
Multi-level supervised vision language model based steel surface defect detection
Semantic Scholar · Engineering · 2024
Abstract
The shape of defects on steel surfaces is highly variable and training samples are limited, making it a significant challenge to transfer a high-performance pretrained vision language model to steel surface defect detection. Therefore, a Multi-level Supervised Vision Language Model based Steel Surface Defect Detection method MLS-VLM is proposed in this paper. MLS-VLM delves deeply into the extraction of profound features from limited samples with three levels of training: supervised contrast training from labeled areas and the entire image, as well as self-supervised contrast learning from Region Proposals. MLS-VLM can be rapidly transferred to two-stage object detector. Experimental results demonstrate that, compared to traditional object detection methods, MLS-VLM achieves 5.68~8.37 mAP improvement on three benchmark object detectors.