Optimizing Loss Functions Through Multivariate Taylor Polynomial Parameterization

Metalearning of deep neural network (DNN) architectures and hyperparameters\nhas become an increasingly important area of research. Loss functions are a\ntype of metaknowledge that is crucial to effective training of DNNs, however,\ntheir potential role in metalearning has not yet been fully explored. Whereas\nearly work focused on genetic programming (GP) on tree representations, this\npaper proposes continuous CMA-ES optimization of multivariate Taylor polynomial\nparameterizations. This approach, TaylorGLO, makes it possible to represent and\nsearch useful loss functions more effectively. In MNIST, CIFAR-10, and SVHN\nbenchmark tasks, TaylorGLO finds new loss functions that outperform functions\npreviously discovered through GP, as well as the standard cross-entropy loss,\nin fewer generations. These functions serve to regularize the learning task by\ndiscouraging overfitting to the labels, which is particularly useful in tasks\nwhere limited training data is available. The results thus demonstrate that\nloss function optimization is a productive new avenue for metalearning.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC