Automated Cleanup of the ImageNet Dataset by Model Consensus, Explainability and Confident Learning

The convolutional neural networks (CNNs) trained on ILSVRC12 ImageNet were\nthe backbone of various applications as a generic classifier, a feature\nextractor or a base model for transfer learning. This paper describes automated\nheuristics based on model consensus, explainability and confident learning to\ncorrect labeling mistakes and remove ambiguous images from this dataset. After\nmaking these changes on the training and validation sets, the ImageNet-Clean\nimproves the model performance by 2-2.4 % for SqueezeNet and EfficientNet-B0\nmodels. The results support the importance of larger image corpora and\nsemi-supervised learning, but the original datasets must be fixed to avoid\ntransmitting their mistakes and biases to the student learner. Further\ncontributions describe the training impacts of widescreen input resolutions in\nportrait and landscape orientations. The trained models and scripts are\npublished on Github (https://github.com/kecsap/imagenet-clean) to clean up\nImageNet and ImageNetV2 datasets for reproducible research.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC