Stability and Generalization of Bilevel Programming in Hyperparameter Optimization

The (gradient-based) bilevel programming framework is widely used in\nhyperparameter optimization and has achieved excellent performance empirically.\nPrevious theoretical work mainly focuses on its optimization properties, while\nleaving the analysis on generalization largely open. This paper attempts to\naddress the issue by presenting an expectation bound w.r.t. the validation set\nbased on uniform stability. Our results can explain some mysterious behaviours\nof the bilevel programming in practice, for instance, overfitting to the\nvalidation set. We also present an expectation bound for the classical\ncross-validation algorithm. Our results suggest that gradient-based algorithms\ncan be better than cross-validation under certain conditions in a theoretical\nperspective. Furthermore, we prove that regularization terms in both the outer\nand inner levels can relieve the overfitting problem in gradient-based\nalgorithms. In experiments on feature learning and data reweighting for noisy\nlabels, we corroborate our theoretical findings.\n

Paper

References (54)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC