Deep Active Learning for Biased Datasets via Fisher Kernel Self-Supervision

Active learning (AL) aims to minimize labeling efforts for data-demanding\ndeep neural networks (DNNs) by selecting the most representative data points\nfor annotation. However, currently used methods are ill-equipped to deal with\nbiased data. The main motivation of this paper is to consider a realistic\nsetting for pool-based semi-supervised AL, where the unlabeled collection of\ntrain data is biased. We theoretically derive an optimal acquisition function\nfor AL in this setting. It can be formulated as distribution shift minimization\nbetween unlabeled train data and weakly-labeled validation dataset. To\nimplement such acquisition function, we propose a low-complexity method for\nfeature density matching using self-supervised Fisher kernel (FK) as well as\nseveral novel pseudo-label estimators. Our FK-based method outperforms\nstate-of-the-art methods on MNIST, SVHN, and ImageNet classification while\nrequiring only 1/10th of processing. The conducted experiments show at least\n40% drop in labeling efforts for the biased class-imbalanced data compared to\nexisting methods.\n

Paper

References (29)

Scroll for more · 17 remaining

Similar papers

© 2026 NYSGPT2525 LLC