Boosting Monocular Depth Estimation Models to High-Resolution via Content-Adaptive Multi-Resolution Merging

Neural networks have shown great abilities in estimating depth from a single\nimage. However, the inferred depth maps are well below one-megapixel resolution\nand often lack fine-grained details, which limits their practicality. Our\nmethod builds on our analysis on how the input resolution and the scene\nstructure affects depth estimation performance. We demonstrate that there is a\ntrade-off between a consistent scene structure and the high-frequency details,\nand merge low- and high-resolution estimations to take advantage of this\nduality using a simple depth merging network. We present a double estimation\nmethod that improves the whole-image depth estimation and a patch selection\nmethod that adds local details to the final result. We demonstrate that by\nmerging estimations at different resolutions with changing context, we can\ngenerate multi-megapixel depth maps with a high level of detail using a\npre-trained model.\n

Paper

References (58)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC