Active learning is the process of training a model with limited labeled data\nby selecting a core subset of an unlabeled data pool to label. The large scale\nof data sets used in deep learning forces most sample selection strategies to\nemploy efficient heuristics. This paper introduces an integer optimization\nproblem for selecting a core set that minimizes the discrete Wasserstein\ndistance from the unlabeled pool. We demonstrate that this problem can be\ntractably solved with a Generalized Benders Decomposition algorithm. Our\nstrategy uses high-quality latent features that can be obtained by unsupervised\nlearning on the unlabeled pool. Numerical results on several data sets show\nthat our optimization approach is competitive with baselines and particularly\noutperforms them in the low budget regime where less than one percent of the\ndata set is labeled.\n
Paper
References (52)
Scroll for more · 38 remaining