The interpretation of medical images is a challenging task, often complicated\nby the presence of artifacts, occlusions, limited contrast and more. Most\nnotable is the case of chest radiography, where there is a high inter-rater\nvariability in the detection and classification of abnormalities. This is\nlargely due to inconclusive evidence in the data or subjective definitions of\ndisease appearance. An additional example is the classification of anatomical\nviews based on 2D Ultrasound images. Often, the anatomical context captured in\na frame is not sufficient to recognize the underlying anatomy. Current machine\nlearning solutions for these problems are typically limited to providing\nprobabilistic predictions, relying on the capacity of underlying models to\nadapt to limited information and the high degree of label noise. In practice,\nhowever, this leads to overconfident systems with poor generalization on unseen\ndata. To account for this, we propose a system that learns not only the\nprobabilistic estimate for classification, but also an explicit uncertainty\nmeasure which captures the confidence of the system in the predicted output. We\nargue that this approach is essential to account for the inherent ambiguity\ncharacteristic of medical images from different radiologic exams including\ncomputed radiography, ultrasonography and magnetic resonance imaging. In our\nexperiments we demonstrate that sample rejection based on the predicted\nuncertainty can significantly improve the ROC-AUC for various tasks, e.g., by\n8% to 0.91 with an expected rejection rate of under 25% for the classification\nof different abnormalities in chest radiographs. In addition, we show that\nusing uncertainty-driven bootstrapping to filter the training data, one can\nachieve a significant increase in robustness and accuracy.\n