Effective Evaluation of Deep Active Learning on Image Classification Tasks

With the goal of making deep learning more label-efficient, a growing number\nof papers have been studying active learning (AL) for deep models. However,\nthere are a number of issues in the prevalent experimental settings, mainly\nstemming from a lack of unified implementation and benchmarking. Issues in the\ncurrent literature include sometimes contradictory observations on the\nperformance of different AL algorithms, unintended exclusion of important\ngeneralization approaches such as data augmentation and SGD for optimization, a\nlack of study of evaluation facets like the labeling efficiency of AL, and\nlittle or no clarity on the scenarios in which AL outperforms random sampling\n(RS). In this work, we present a unified re-implementation of state-of-the-art\nAL algorithms in the context of image classification via our new open-source AL\ntoolkit DISTIL, and we carefully study these issues as facets of effective\nevaluation. On the positive side, we show that AL techniques are $2\\times$ to\n$4\\times$ more label-efficient compared to RS with the use of data\naugmentation. Surprisingly, when data augmentation is included, there is no\nlonger a consistent gain in using BADGE, a state-of-the-art approach, over\nsimple uncertainty sampling. We then do a careful analysis of how existing\napproaches perform with varying amounts of redundancy and number of examples\nper class. Finally, we provide several insights for AL practitioners to\nconsider in future work, such as the effect of the AL batch size, the effect of\ninitialization, the importance of retraining the model at every round, and\nother insights.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC