Dataset Cleaning -- A Cross Validation Methodology for Large Facial Datasets using Face Recognition
In recent years, large "in the wild" face datasets have been released in an\nattempt to facilitate progress in tasks such as face detection, face\nrecognition, and other tasks. Most of these datasets are acquired from webpages\nwith automatic procedures. As a consequence, noisy data are often found.\nFurthermore, in these large face datasets, the annotation of identities is\nimportant as they are used for training face recognition algorithms. But due to\nthe automatic way of gathering these datasets and due to their large size, many\nidentities folder contain mislabeled samples which deteriorates the quality of\nthe datasets. In this work, it is presented a semi-automatic method for\ncleaning the noisy large face datasets with the use of face recognition. This\nmethodology is applied to clean the CelebA dataset and show its effectiveness.\nFurthermore, the list with the mislabelled samples in the CelebA dataset is\nmade available.\n