URSABench: Comprehensive Benchmarking of Approximate Bayesian Inference Methods for Deep Neural Networks
While deep learning methods continue to improve in predictive accuracy on a\nwide range of application domains, significant issues remain with other aspects\nof their performance including their ability to quantify uncertainty and their\nrobustness. Recent advances in approximate Bayesian inference hold significant\npromise for addressing these concerns, but the computational scalability of\nthese methods can be problematic when applied to large-scale models. In this\npaper, we describe initial work on the development ofURSABench(the Uncertainty,\nRobustness, Scalability, and Accu-racy Benchmark), an open-source suite of\nbench-marking tools for comprehensive assessment of approximate Bayesian\ninference methods with a focus on deep learning-based classification tasks\n