We present a novel algorithm for self-supervised monocular depth completion.\nOur approach is based on training a neural network that requires only sparse\ndepth measurements and corresponding monocular video sequences without dense\ndepth labels. Our self-supervised algorithm is designed for challenging indoor\nenvironments with textureless regions, glossy and transparent surface,\nnon-Lambertian surfaces, moving people, longer and diverse depth ranges and\nscenes captured by complex ego-motions. Our novel architecture leverages both\ndeep stacks of sparse convolution blocks to extract sparse depth features and\npixel-adaptive convolutions to fuse image and depth features. We compare with\nexisting approaches in NYUv2, KITTI, and NAVERLABS indoor datasets, and observe\n5-34 % improvements in root-means-square error (RMSE) reduction.\n