Depth Anywhere: Enhancing 360 Monocular Depth Estimation via Perspective Distillation and Unlabeled Data Augmentation

Accurately estimating depth in 360-degree imagery is crucial for virtual reality, autonomous navigation, and immersive media applications. Existing depth estimation methods designed for perspective-view imagery fail when applied to 360-degree images due to different camera projections and distortions, whereas 360-degree methods perform inferior due to the lack of labeled data pairs. We propose a new depth estimation framework that utilizes unlabeled 360-degree data effectively. Our approach uses state-of-the-art perspective depth estimation models as teacher models to generate pseudo labels through a six-face cube projection technique, enabling efficient labeling of depth in 360-degree images. This method leverages the increasing availability of large datasets. Our approach includes two main stages: offline mask generation for invalid regions and an online semi-supervised joint training regime. We tested our approach on benchmark datasets such as Matterport3D and Stanford2D3D, showing significant improvements in depth estimation accuracy, particularly in zero-shot scenarios. Our proposed training pipeline can enhance any 360 monocular depth estimator and demonstrates effective knowledge transfer across different camera projections and data types. See our project page for results: https://albert100121.github.io/Depth-Anywhere/

Paper

References (69)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer Kp4U5/10 · confidence 3/52024-07-10

Summary

The authors propose a technique that utilizes unlabeled 360-degree data to improve previous methods, which includes two main stages: offline mask generation for invalid regions and an online semi-supervised joint training regime. Experimental results indicate that the proposed method outperforms previous methods on the Matterport3D and Stanford2D3D datasets.

Strengths

The paper is well written and clearly structured; The performance is promising; The experiments are thorough and the proposed method is validated on multiple datasets.

Weaknesses

My major concern is that the technical contribution seems limited. The core idea of utilizing unlabeled data has been proposed by DepthAnything. The proposed method is more like an application of this idea to the 360 data; As shown in Tables 2 and 3, previous 360 depth estimation methods typically use the metric loss for training. While the affine-invariant approach employed by the authors enables training on multiple datasets, it also impedes real-world applications due to the lose of metric scale.

Questions

Please see the weakness.

Rating

5

Confidence

3

Soundness

3

Presentation

3

Contribution

2

Limitations

N/A

Reviewer ewNL5/10 · confidence 5/52024-07-10

Summary

This paper proposes a method to improve 360 monocular depth estimation using perspective distillation and augmentation with unlabeled data. It introduces the concept of "perspective distillation," which leverages the available 360 monocular depth maps and their corresponding equirectangular images to generate pixel-wise depth supervision signals. This technique helps to address the lack of ground truth depth data for training in the 360 domain. Additionally, the paper presents an unlabeled data augmentation approach that utilizes the geometric properties of 360-degree images. By exploiting the spherical geometry, the authors generate synthetic stereo pairs to augment the training dataset without requiring paired depth information. The proposed method is evaluated on benchmark datasets and achieves significant improvements in terms of depth estimation accuracy compared to existing approaches. The results demonstrate the effectiveness of perspective distillation and unlabeled data augmentation in enhancing the performance of 360 monocular depth estimation.

Strengths

1. Innovative techniques: The paper introduces novel approaches, such as perspective distillation and unlabeled data augmentation, to address the challenges of 360 monocular depth estimation. 2. Improved depth estimation: The proposed method achieves significant improvements in depth estimation accuracy compared to existing approaches, as demonstrated through rigorous evaluations on benchmark datasets. 3. Use of unlabeled data: By leveraging unlabeled data and synthetic stereo pairs, the method reduces the reliance on paired depth information, which is often difficult to obtain in the 360-degree domain.

Weaknesses

1. Complexity: The proposed method introduces additional complexity, such as perspective distillation and synthetic stereo pair generation, which may require more computational resources and training time. 2. Dataset dependency: The effectiveness of the proposed method heavily relies on the availability and quality of the benchmark datasets used for evaluation, which may affect its generalizability to real-world scenarios. 3. Limited scope: The paper focuses specifically on 360 monocular depth estimation, which may limit its applicability to other depth estimation tasks or domains, such as 360 monocular depth completion. 4. Insufficient related work: Adding the latest panoramic depth estimation and panoramic depth completion methods is preferred.

Questions

Overall, the idea of this paper is interesting, allowing existing depth estimation models to benefit from unlabeled data in a semi-supervised manner, but there are concerns: 1. Based on the experimental results, the author only validated the performance of this training technique on models that employ dual-projection fusion, such as Unifuse and Bifuse. There was no effective validation or analysis provided for other non-dual-projection fusion models, such as methods based on horizontal compression (e.g., HorizonNet, HohoNet) or transformer-based methods (e.g., EGFormer). However, it is essential to explicitly state the scope of applicability for this training strategy. Different training strategies, data processing methods, and even device variations can lead to unfair comparisons. The author should conduct fair experiments and comparisons under the same experimental conditions instead of directly copying the data results from the original paper. This is crucial for validating the effectiveness of this training strategy. Therefore, the author's mention of "as many of the aforementioned methods did not release pre-trained models or provide training code and implementation details. It’s worth noting that PanoFormer [29] is not included due to incorrect evaluation code and results, and EGFormer [29] is not included since its experiments are mainly conducted on other datasets and benchmarks" is not convincing. 2. The issue of cross-domain experiments is indeed important. Given the challenges of obtaining real-world data, it would be meaningful if this framework could benefit models trained on synthetic data (e.g., training on synthetic data and testing on real data). However, the significance of this training strategy in that regard is not yet clear. Further research and experimentation are necessary to determine whether training on synthetic data using this framework can indeed yield benefits when applied to real-world data. It would be valuable to investigate the effectiveness and generalization capabilities of the trained models in real-world scenarios. 3. Median alignment is currently not widely utilized in existing depth estimation methods. Based on existing findings, aligning the ground truth (GT) depth with the predicted depth can indeed lead to improved results. However, when comparing the proposed method with baseline methods like Unifuse, which explicitly state that median alignment is not used, directly comparing the results to those from the original paper leads to unfair comparisons. This raises concerns about the advantages claimed for this training framework. To ensure fair comparisons, it is important to apply the same alignment techniques consistently across all methods being compared.

Rating

5

Confidence

5

Soundness

3

Presentation

3

Contribution

3

Limitations

This paper proposes an interesting solution for panoramic depth estimation task, ie.e, perspective distillation and unlabeled data augmentation. It could contribute a lot for the community. The reviewer suggests introducing more related works, including the latest panoramic depth estimation and panoramic depth completion approaches.

Authorsrebuttal2024-08-09

Thank you for your suggestion. In addition to those already cited in the original paper, we will add the following papers to our related work discussion. **Panoramic Depth Estimation** - $ 𝑆^2 $ Net: Accurate Panorama Depth Estimation on Spherical Surface - High-Resolution Depth Estimation for 360-degree Panoramas through Perspective and Panoramic Depth Images Registration - Adversarial Mixture Density Network and Uncertainty-based Joint Learning for 360 Monocular Depth Estimation - Learning high-quality depth map from 360 multi-exposure imagery - Distortion-Aware Self-Supervised 360 Depth Estimation from A Single Equirectangular Projection Image - SphereDepth: Panorama Depth Estimation from Spherical Domain - 360 Depth Estimation in the Wild -- the Depth360 Dataset and the SegFuse Network - GLPanoDepth: Global-to-Local Panoramic Depth Estimation - Neural Contourlet Network for Monocular 360 Depth Estimation - HiMODE: A Hybrid Monocular Omnidirectional Depth Estimation Model - Geometric Structure Based and Regularized Depth Estimation From 360 Indoor Imagery - Deep Depth Estimation on 360 Images with a Double Quaternion Loss **Panoramic Depth Completion** - Cross-Modal 360° Depth Completion and Reconstruction for Large-Scale Indoor Environment - 360 ORB-SLAM: A Visual SLAM System for Panoramic Images with Depth Completion Network - Deep panoramic depth prediction and completion for indoor scenes - Multi-Modal Masked Pre-Training for Monocular Panoramic Depth Completion - Distortion and Uncertainty Aware Loss for Panoramic Depth Completion - Indoor Depth Completion with Boundary Consistency and Self-Attention Please let us know if any specific paper is missing.

Authorsrebuttal2024-08-13

Dear Reviewer ewNL, We have listed the latest **panoramic depth estimation** and **completion** methods and will discuss them in the related work section of the final version. Could you please confirm that we did not miss any essential references and addressed all your concerns? Thank you!

Reviewer ewNL2024-08-14

Thanks for the response. I tend to maintain the initial rating, since the current version of this paper needs to solve many weaknesses. But I am still willing to give a **borderline accept** score.

Reviewer AagT5/10 · confidence 5/52024-07-12

Summary

This paper effectively utilizes unlabeled data by employing the SAM and DepthAnything models to generate masks and pseudo-labels respectively. When projecting data onto a cube, the authors use random rotation techniques to minimize cube artifacts, thereby enhancing the accuracy of 360-degree monocular depth estimation. Additionally, the method was tested in zero-shot scenarios, demonstrating its effective knowledge transfer.

Strengths

This paper introduces a training technique for 360-degree imagery that enhances depth estimation performance by generating pseudo-labels with the DepthAnything model to leverage information from unlabeled data. Additionally, it employs the SAM model to segment irrelevant areas such as the sky in outdoor panoramic images. Furthermore, the method uses random rotation preprocessing to eliminate cube artifacts.

Weaknesses

- The paper stated that "As depicted in Figure 2, the rotation is applied to the equirectangular projection RGB images using a random rotation matrix, followed by cube projection. This results in a more diverse set of cube faces, effectively capturing the relative distances between ceilings, walls, windows, and other objects." However, from observing Figure 2, the unlabeled data seems to be entirely outdoor panoramas, which makes the mention of indoor elements such as ceilings, walls, and windows confusing. If there are indeed indoor panoramic images in the unlabeled data, what elements might the SAM model need to segment in such indoor panoramas? The description in the paper appears somewhat unclear. - The paper mentioned, "We chose UniFuse and BiFuse++ as our baseline models for experiments, as many of the aforementioned methods did not release pre-trained models or provide training code and implementation details." However, methods such as HRDfuse [1], EGFormer [50], BiFuse and BiFuse++ [35, 36], UniFuse [11], and PanoFormer [29] have all made their source codes available, making this reason in the paper seem insufficient. Additionally, "EGFormer is not included since its experiments are mainly conducted on other datasets and benchmarks" appears to be an inadequate reason for not including it in the experiments. - In Table 3, despite introducing more unlabeled data on the Unifuse training set, i.e., SP-all (p), the performance of the method is not significantly improved. This phenomenon seems to be only effective on the BiFuse model, but not on other methods. - In Table 3, there is a typographical error in the recording of the Abs Rel value; it should not be 0.858. Additionally, in Table 2, "UniFuse" is incorrectly written as "UniFise."

Questions

- The paper mentioned, "Subsequently, in the online stage, we adopt a semi-supervised learning strategy, loading half of the batch with labeled data and the other half with pseudo-labeled data." However, common semi-supervised strategies typically set the ratio of labeled to total data at 1/2, 1/4, 1/8, 1/16, etc. The experiments in the paper were only conducted at a 1:1 ratio (labeled: unlabeled), and thus the performance at other ratios remains unknown. - The third contribution mentioned in the paper refers to "interchangeability" and "This enables better results even as new SOTA techniques emerge in the future." This suggests that the strategy might be adaptable to a variety of models. However, the experiments were conducted only on models like UniFuse and BiFuse++ which use a dual projection fusion of Cube and ERP. Whether this approach would perform well with other transformer-based models remains an unresolved question.

Rating

5

Confidence

5

Soundness

3

Presentation

3

Contribution

2

Limitations

- The paper stated, "Cube projection and tangent projection are the most common techniques. We selected cube projection to ensure a larger field of view for each patch." This is merely a theoretical assertion, with no experimental evidence to prove which projection method is superior. Additionally, using cube projection directly can lead to cube artifacts. Following the suggestion in [1], setting each panoramic image to have 10 or 18 tangent images using more polyhedral faces, rather than the standard six faces, might reduce the artifacts caused by direct cube projection. [1] Cokelek, M., Imamoglu, N., Ozcinar, C., Erdem, E. and Erdem, A., 2023. Spherical Vision Transformer for 360-degree Video Saliency Prediction. BMVC 2023.

Reviewer F7bZ5/10 · confidence 5/52024-07-12

Summary

This paper introduces a novel depth estimation framework specifically designed for 360-degree data using an innovative two-stage process: offline mask generation and online semi-supervised joint training. Initially, invalid regions such as sky and watermarks are masked using detection and segmentation models. The method then employs a semi-supervised learning approach, blending labeled and pseudo-labeled data derived from state-of-the-art perspective depth models using a cube projection technique for effective training. This framework demonstrates adaptability across different state-of-the-art models and datasets, effectively tackling the challenges of depth estimation in 360-degree imagery.

Strengths

1. The proposed method employs models trained for traditional pinhole cameras to enhance 360-degree depth estimation, a first in the field. Moreover, the main motivation of this paper is reasonable. 2. It outperforms conventional methods by incorporating pseudo labels from foundational models into the loss function. 3. The paper demonstrates the model's generalizability to real-world scenarios, indicating its practical utility.

Weaknesses

1. The proposed method is straightforward and the performance gains provided by the proposed method are described as marginal. 2. The method's effectiveness is demonstrated only with specific models, UniFuse and BiFuse++, limiting evidence of its broader applicability. 3. Inconsistencies in the decimal points used in quantitative results tables make direct performance comparisons challenging.

Questions

1. Is there any reason the author only applied the proposed method to the UniFuse and BiFuse++? 2. There seems to be a discrepancy in the reported Absolute Relative (Abs Rel) error for BiFuse++ (Affine-Inv, M-all, SP-all(p)) at 0.858, which is ten times higher than results from competitive methods, while other metrics (\delta_1, \delta_2, \delta_3) align closely. Could this be an error?

Rating

5

Confidence

5

Soundness

2

Presentation

2

Contribution

2

Limitations

See the weakness and the question parts.

Reviewer 3NfA6/10 · confidence 4/52024-07-14

Summary

The authors present a training strategy for single-image depth estimation on 360-degree equirectangular images. The strategy centers around leveraging strong pre-trained models for perspective images as teacher networks. It does not depend on any particular network architecture and therefore can benefit any 360 depth estimation networks. Experimental results validate that the strategy improves models otherwise trained only with limited ground-truth annotations.

Strengths

**Originality and significance**: 360-degree imagery is becoming increasingly critical for many computer vision applications. However, there is still a significant gap in GT depth annotations compared to their perspective counterparts, and any solution to this issue can have a significant impact. The authors demonstrate that leveraging strong perspective depth models, e.g., Depth Anything, is a simple yet effective solution. **Quality**: The main idea and the several supporting procedures (e.g., random rotation processing, valid pixel masking, mixed labeled and unlabelled training) are all well-motivated and reasonably designed. Experimental results show consistent improvement by incorporating the proposed pseudo GT. Additional results including zero-shot and qualitative evaluations further help with understanding and make the approach overall more convincing. **Clarity**: Paper is well-written with good structure, clear expressions, and adequate details.

Weaknesses

1. A substantive assessment of the weaknesses of the paper. Focus on constructive and actionable insights on how the work could improve towards its stated goals. Be specific, and avoid generic remarks.  Despite focusing on a slightly different task (stereo depth), FoVA-Depth (Lichy et al. 3DV 2024) presents a few very similar concepts: - leveraging abundance of perspective depth GT - cube map as a intermediate representation to gap between 360 and perspective images - random rotation augmentation It is worth some discussion regarding similarities and differences. 2. A very simple yet critical baseline is missing: directly project pseudo GT on cubemap to equirectangular images. A good stitching strategy may be challenging, but with something simple or even without any additional scaling, it should help clarify how much the pre-trained depth anything model contribute to the overall performance. 3. The paper addresses only relative depth estimation. Since several baselines (e.g. upper section of Tab.2) already have metric counterparts, and depth anything has metric variants, I feel some experiments and analysis in that regards should be straightforward and nicely complement the relative depth results. 4. As the premise of the work is the usefulness of the abundant pseudo GT compard to limited 360 depth GT, it is necessary to understand how does the benefit from pseudo GT scale (how many pseudo GT can be as useful as a real 360 GT?), and how does the student network compare to the teacher network in terms of estimation quality (is the quality of teacher network already a bottleneck?). The paper offers little insight in these questions.

Questions

I am looking forward to answer to the questions raised above, namely: - How does training scale in terms of the amount of pseudo GT vs real GT? - Is the approach applicable to metric depth estimation? - Additional related work and baseline as described above.

Rating

6

Confidence

4

Soundness

3

Presentation

3

Contribution

3

Limitations

The only limitation the authors bring up (quality of unlabelled data) already has a solution in the paper, so not really a limitation? I do think there are a few other things worth mentioning: - 360 data, even without requiring GT, is still scarce compared to perspective data. The fact the authors can only evaluate with two such datasets is an evidance. This is a limitation since it prevents further scaling up training. - Only equirectangular images are supported (though it seems that, in principle, the approach should work in more general camera models).

Reviewer ewNL2024-08-09

Thanks for the responses. As mentioned in weakness 4, more related works should be involved in the Related Work section, including the latest **panoramic depth estimation** and **panoramic depth completion** methods.

Area Chair XoPq2024-08-09

Discussion about Submission 2414

Dear reviewers, please share your opinion about the rebuttal; it provides a detailed response to the questions raised in your reviews, including issues regarding novelty and experimental evaluation. Thank you AC

Reviewer F7bZ2024-08-09

Thank you for providing the additional experiments. The experiments for the other baseline are quite convincing, and I have increased my initial rating accordingly.

Authorsrebuttal2024-08-10

Thank you for your constructive review and valuable feedback. Your insights have been instrumental in enhancing the quality of our paper.

Authorsrebuttal2024-08-12

Please let us know if you have additional questions after reading our response

Dear Reviewers, We appreciate your reviews and comments. We hope our responses address your concerns. Please let us know if you have further questions after reading our rebuttal. We aim to address all the potential issues during the discussion period. Thank you! Best, Authors

Area Chair XoPq2024-08-12

Feedback

Dear reviewer Kp4U, You raised concerns about the paper's technical novelty and contribution, and the authors responded in detail. Please share your feedback with us. Thank you

Authorsrebuttal2024-08-13

Please let us know if you have additional questions after reading our response

Dear Reviewer, As we approach the end of the discussion period, we want to confirm whether we have successfully addressed your concerns. Should any lingering issues require further attention, please let us know as early as possible so we can answer them soon. We appreciate your time and effort in enhancing the quality of our manuscript. Thank you!

Authorsrebuttal2024-08-14

Dear Reviewer, As the discussion period ends soon, have we addressed all your concerns? If any issues remain, please inform us promptly. We appreciate your help in improving our manuscript. Thank you!

Area Chair XoPq2024-08-12

Feedback to authors

Dear Reviewers AagT and Reviewer 3NfA, The authors replied to the questions raised in your initial evaluation report. Did the author address your concerns? Please post your feedback about the rebuttal. Thank you

Authorsrebuttal2024-08-13

Please let us know if you have additional questions after reading our response

Dear Reviewer AagT, We appreciate your reviews and comments. We hope our responses address your concerns. Please let us know if you have further questions after reading our rebuttal. We aim to address all the potential issues during the discussion period. Thank you! Best, Authors

Reviewer 3NfA2024-08-12

Thank you for the responses! I think these are all valuable details and hopefully much of them will be integrated into the paper. I don't have further questions.

Authorsrebuttal2024-08-13

Thank you for your constructive review, which has significantly contributed to the improvement of our paper. We will ensure that the materials from the rebuttal are incorporated into the final version.

Reviewer AagT2024-08-14

Thanks for the rebuttal, which addressed some of my concerns. I would like to increase the rating. The updated clarification and experiments are expected to be presented in the final version.

Authorsrebuttal2024-08-14

Thank you for your valuable review and feedback, which have significantly improved the completeness of our paper. We will ensure that the updated clarifications and experiments are incorporated into the final version.

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC