Low Budget Active Learning via Wasserstein Distance: An Integer Programming Approach

Active learning is the process of training a model with limited labeled data\nby selecting a core subset of an unlabeled data pool to label. The large scale\nof data sets used in deep learning forces most sample selection strategies to\nemploy efficient heuristics. This paper introduces an integer optimization\nproblem for selecting a core set that minimizes the discrete Wasserstein\ndistance from the unlabeled pool. We demonstrate that this problem can be\ntractably solved with a Generalized Benders Decomposition algorithm. Our\nstrategy uses high-quality latent features that can be obtained by unsupervised\nlearning on the unlabeled pool. Numerical results on several data sets show\nthat our optimization approach is competitive with baselines and particularly\noutperforms them in the low budget regime where less than one percent of the\ndata set is labeled.\n

Paper

References (52)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC