Real-time 3D reconstruction enables fast dense mapping of the environment\nwhich benefits numerous applications, such as navigation or live evaluation of\nan emergency. In contrast to most real-time capable approaches, our approach\ndoes not need an explicit depth sensor. Instead, we only rely on a video stream\nfrom a camera and its intrinsic calibration. By exploiting the self-motion of\nthe unmanned aerial vehicle (UAV) flying with oblique view around buildings, we\nestimate both camera trajectory and depth for selected images with enough novel\ncontent. To create a 3D model of the scene, we rely on a three-stage processing\nchain. First, we estimate the rough camera trajectory using a simultaneous\nlocalization and mapping (SLAM) algorithm. Once a suitable constellation is\nfound, we estimate depth for local bundles of images using a Multi-View Stereo\n(MVS) approach and then fuse this depth into a global surfel-based model. For\nour evaluation, we use 55 video sequences with diverse settings, consisting of\nboth synthetic and real scenes. We evaluate not only the generated\nreconstruction but also the intermediate products and achieve competitive\nresults both qualitatively and quantitatively. At the same time, our method can\nkeep up with a 30 fps video for a resolution of 768x448 pixels.\n