Rethinking Importance Weighting for Deep Learning under Distribution Shift

Under distribution shift (DS) where the training data distribution differs\nfrom the test one, a powerful technique is importance weighting (IW) which\nhandles DS in two separate steps: weight estimation (WE) estimates the\ntest-over-training density ratio and weighted classification (WC) trains the\nclassifier from weighted training data. However, IW cannot work well on complex\ndata, since WE is incompatible with deep learning. In this paper, we rethink IW\nand theoretically show it suffers from a circular dependency: we need not only\nWE for WC, but also WC for WE where a trained deep classifier is used as the\nfeature extractor (FE). To cut off the dependency, we try to pretrain FE from\nunweighted training data, which leads to biased FE. To overcome the bias, we\npropose an end-to-end solution dynamic IW that iterates between WE and WC and\ncombines them in a seamless manner, and hence our WE can also enjoy deep\nnetworks and stochastic optimizers indirectly. Experiments with two\nrepresentative types of DS on three popular datasets show that our dynamic IW\ncompares favorably with state-of-the-art methods.\n

Paper

References (71)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC