AutoDO: Robust AutoAugment for Biased Data with Label Noise via Scalable Probabilistic Implicit Differentiation

AutoAugment has sparked an interest in automated augmentation methods for\ndeep learning models. These methods estimate image transformation policies for\ntrain data that improve generalization to test data. While recent papers\nevolved in the direction of decreasing policy search complexity, we show that\nthose methods are not robust when applied to biased and noisy data. To overcome\nthese limitations, we reformulate AutoAugment as a generalized automated\ndataset optimization (AutoDO) task that minimizes the distribution shift\nbetween test data and distorted train dataset. In our AutoDO model, we\nexplicitly estimate a set of per-point hyperparameters to flexibly change\ndistribution of train data. In particular, we include hyperparameters for\naugmentation, loss weights, and soft-labels that are jointly estimated using\nimplicit differentiation. We develop a theoretical probabilistic interpretation\nof this framework using Fisher information and show that its complexity scales\nlinearly with the dataset size. Our experiments on SVHN, CIFAR-10/100, and\nImageNet classification show up to 9.3% improvement for biased datasets with\nlabel noise compared to prior methods and, importantly, up to 36.6% gain for\nunderrepresented SVHN classes.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC