Marginal-likelihood based model-selection, even though promising, is rarely\nused in deep learning due to estimation difficulties. Instead, most approaches\nrely on validation data, which may not be readily available. In this work, we\npresent a scalable marginal-likelihood estimation method to select both\nhyperparameters and network architectures, based on the training data alone.\nSome hyperparameters can be estimated online during training, simplifying the\nprocedure. Our marginal-likelihood estimate is based on Laplace's method and\nGauss-Newton approximations to the Hessian, and it outperforms cross-validation\nand manual-tuning on standard regression and image classification datasets,\nespecially in terms of calibration and out-of-distribution detection. Our work\nshows that marginal likelihoods can improve generalization and be useful when\nvalidation data is unavailable (e.g., in nonstationary settings).\n
Paper
References (53)
Scroll for more · 38 remaining