The ability to accurately track what happens during a conversation is\nessential for the performance of a dialogue system. Current state-of-the-art\nmulti-domain dialogue state trackers achieve just over 55% accuracy on the\ncurrent go-to benchmark, which means that in almost every second dialogue turn\nthey place full confidence in an incorrect dialogue state. Belief trackers, on\nthe other hand, maintain a distribution over possible dialogue states. However,\nthey lack in performance compared to dialogue state trackers, and do not\nproduce well calibrated distributions. In this work we present state-of-the-art\nperformance in calibration for multi-domain dialogue belief trackers using a\ncalibrated ensemble of models. Our resulting dialogue belief tracker also\noutperforms previous dialogue belief tracking models in terms of accuracy.\n