This work presents a probabilistic deep neural network that combines LiDAR\npoint clouds and RGB camera images for robust, accurate 3D object detection. We\nexplicitly model uncertainties in the classification and regression tasks, and\nleverage uncertainties to train the fusion network via a sampling mechanism. We\nvalidate our method on three datasets with challenging real-world driving\nscenarios. Experimental results show that the predicted uncertainties reflect\ncomplex environmental uncertainty like difficulties of a human expert to label\nobjects. The results also show that our method consistently improves the\nAverage Precision by up to 7% compared to the baseline method. When sensors are\ntemporally misaligned, the sampling method improves the Average Precision by up\nto 20%, showing its high robustness against noisy sensor inputs.\n