Automatic Tuning of Stochastic Gradient Descent with Bayesian Optimisation

Many machine learning models require a training procedure based on running\nstochastic gradient descent. A key element for the efficiency of those\nalgorithms is the choice of the learning rate schedule. While finding good\nlearning rates schedules using Bayesian optimisation has been tackled by\nseveral authors, adapting it dynamically in a data-driven way is an open\nquestion. This is of high practical importance to users that need to train a\nsingle, expensive model. To tackle this problem, we introduce an original\nprobabilistic model for traces of optimisers, based on latent Gaussian\nprocesses and an auto-/regressive formulation, that flexibly adjusts to abrupt\nchanges of behaviours induced by new learning rate values. As illustrated, this\nmodel is well-suited to tackle a set of problems: first, for the on-line\nadaptation of the learning rate for a cold-started run; then, for tuning the\nschedule for a set of similar tasks (in a classical BO setup), as well as\nwarm-starting it for a new task.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC