We propose an evaluation framework for class probability estimates (CPEs) in\nthe presence of label uncertainty, which is commonly observed as diagnosis\ndisagreement between experts in the medical domain. We also formalize\nevaluation metrics for higher-order statistics, including inter-rater\ndisagreement, to assess predictions on label uncertainty. Moreover, we propose\na novel post-hoc method called $alpha$-calibration, that equips neural network\nclassifiers with calibrated distributions over CPEs. Using synthetic\nexperiments and a large-scale medical imaging application, we show that our\napproach significantly enhances the reliability of uncertainty estimates:\ndisagreement probabilities and posterior CPEs.\n