In this paper, we present TANDEM a real-time monocular tracking and dense\nmapping framework. For pose estimation, TANDEM performs photometric bundle\nadjustment based on a sliding window of keyframes. To increase the robustness,\nwe propose a novel tracking front-end that performs dense direct image\nalignment using depth maps rendered from a global model that is built\nincrementally from dense depth predictions. To predict the dense depth maps, we\npropose Cascade View-Aggregation MVSNet (CVA-MVSNet) that utilizes the entire\nactive keyframe window by hierarchically constructing 3D cost volumes with\nadaptive view aggregation to balance the different stereo baselines between the\nkeyframes. Finally, the predicted depth maps are fused into a consistent global\nmap represented as a truncated signed distance function (TSDF) voxel grid. Our\nexperimental results show that TANDEM outperforms other state-of-the-art\ntraditional and learning-based monocular visual odometry (VO) methods in terms\nof camera tracking. Moreover, TANDEM shows state-of-the-art real-time 3D\nreconstruction performance.\n