Benchmarking Bayesian neural networks and evaluation metrics for regression tasks

Due to the growing adoption of deep neural networks in many fields of science\nand engineering, modeling and estimating their uncertainties has become of\nprimary importance. Despite the growing literature about uncertainty\nquantification in deep learning, the quality of the uncertainty estimates\nremains an open question. In this work, we assess for the first time the\nperformance of several approximation methods for Bayesian neural networks on\nregression tasks by evaluating the quality of the confidence regions with\nseveral coverage metrics. The selected algorithms are also compared in terms of\npredictivity, kernelized Stein discrepancy and maximum mean discrepancy with\nrespect to a reference posterior in both weight and function space. Our\nfindings show that (i) some algorithms have excellent predictive performance\nbut tend to largely over or underestimate uncertainties (ii) it is possible to\nachieve good accuracy and a given target coverage with finely tuned\nhyperparameters and (iii) the promising kernel Stein discrepancy cannot be\nexclusively relied on to assess the posterior approximation. As a by-product of\nthis benchmark, we also compute and visualize the similarity of all algorithms\nand corresponding hyperparameters: interestingly we identify a few clusters of\nalgorithms with similar behavior in weight space, giving new insights on how\nthey explore the posterior distribution.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC