Towards Label-free Scene Understanding by Vision Foundation Models

Vision foundation models such as Contrastive Vision-Language Pre-training (CLIP) and Segment Anything (SAM) have demonstrated impressive zero-shot performance on image classification and segmentation tasks. However, the incorporation of CLIP and SAM for label-free scene understanding has yet to be explored. In this paper, we investigate the potential of vision foundation models in enabling networks to comprehend 2D and 3D worlds without labelled data. The primary challenge lies in effectively supervising networks under extremely noisy pseudo labels, which are generated by CLIP and further exacerbated during the propagation from the 2D to the 3D domain. To tackle these challenges, we propose a novel Cross-modality Noisy Supervision (CNS) method that leverages the strengths of CLIP and SAM to supervise 2D and 3D networks simultaneously. In particular, we introduce a prediction consistency regularization to co-train 2D and 3D networks, then further impose the networks' latent space consistency using the SAM's robust feature representation. Experiments conducted on diverse indoor and outdoor datasets demonstrate the superior performance of our method in understanding 2D and 3D open environments. Our 2D and 3D network achieves label-free semantic segmentation with 28.4\% and 33.5\% mIoU on ScanNet, improving 4.7\% and 7.9\%, respectively. For nuImages and nuScenes datasets, the performance is 22.1\% and 26.8\% with improvements of 3.5\% and 6.0\%, respectively. Code is available. (https://github.com/runnanchen/Label-Free-Scene-Understanding).

Paper

Similar papers

Peer review

Reviewer inkR7/10 · confidence 4/52023-07-02

Summary

The paper proposes a method that utilizes the SAM and CLIP (Contrastive Language-Image Pretraining) models for understanding the 2D and 3D world in a label-free manner. The authors suggest a two-step process to improve the results obtained from CLIP. In the first step, the authors generate a noisy output from the CLIP model, which is primarily designed for semantic tasks. This output may contain imperfections or inaccuracies due to the inherent limitations of CLIP.To refine the output and obtain more visually appealing semantic results, the authors employ a SAM network in the second step. The SAM network is designed to enhance the quality of the semantic results generated by CLIP. Furthermore, the paper introduces a Cross-modality Noise Supervision module, which aims to optimize the training of both the 2D and 3D models simultaneously. The authors demonstrate promising results on datasets such as ScanNet and nuScenes. These datasets are commonly used in the field of computer vision and provide challenging scenarios for understanding the 2D and 3D world. The encouraging results obtained from these datasets suggest the effectiveness of the proposed approach in understanding and representing the visual and semantic aspects of the real world.

Strengths

Strength: 1. Paper pushes the current boundary of training models with label free training data, using pre-existing models which have shown good performance on zero-shot tasks. 2. Proposes to seamlessly refine labels predicted by CLIP model to achieve good results in Indoor Dataset. 3. Paper for most of the part is well written and has a definite flow to it. 4. Beats existing method such MaskCLIP, CLIP2scene models comprehensively to get state-of-arts results.

Weaknesses

1. Missing Ablation: Prediction Consistency Regularization, the second stage where there is random switching of pseudo labels between 2D, 3D and CLIP based intermediate outputs, is not well explained. 2. Missing Ablation: Ablation study which shows what is the impact of this random switching with and without. We also need to see what if we only employ one of the three pseudo labels and not all three, what impact does that have. Seeing the impact of each individual is necessary to comment about its efficacy. 3. Effect of Latent Space Consistency Regularization: There is only marginal improvement seen using SAM features to guide the feature of intermediate 2D/3D encoder features. Do the authors have intermediate results(qualitative), which justifies how SAM features actually help the 2D and 3D features ? Typos: 1. Line no 169: is the doc product operation. I presume this is supposed to be dot product. 2. In Fig:2, How do we achieve the Refined CLIP Pseudo Label and how is this different from Refined 2D Pseudo labels. Please add further information for this.

Questions

1. Could the authors provide intermediate feature space visualization which shows comparative results with and without latent space regularization ? Since from the ablation experiments provided there seems to be marginal improvement in mIoU numbers both in the ScanNet and nuScenes dataset.

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

3 good

Contribution

2 fair

Limitations

1. The method proposed assumes SAM to be perfect in sense of providing meaning masks and does not train the SAM model. But fails to give examples or quote scenarios where this might fail, considering SAM outputs a mask which is not meaningful or helping to noisy CLIP based labels.

Reviewer rM414/10 · confidence 4/52023-07-03

Summary

This work employs 2D segmentation foundation models to perform semantic segmentation for both 2D and 3D indoor scenes. The clip model provides a general understanding of the semantic content in each image, while the SAM model generates precise segmentation masks. By integrating these two models, this study demonstrates improved quantitative results on two datasets.

Strengths

1. The significant enhancement in segmentation results is achieved by leveraging two distinctively trained 2D segmentation models, SAM and Clip, surpassing the previous approach that solely utilized the Clip model. 2. It is reasonable to anticipate improved segmentation outcomes by combining two powerful models that provide both semantic understanding and precise mask contours. The merging of diverse masks through pooling is straightforward and appears to be the most effective element contributing to the performance gain.

Weaknesses

1. Writing: The authors are encouraged to provide further clarification on the "calibration matrix," as it serves as the initial step in connecting 2D pixels and 3D points. How was this connection established without knowledge of the depth of each pixel? Additionally, could you please clarify the meaning of "Image Number" mentioned in line 305 and indicate the corresponding table or figure for reference? A similar question arises regarding the ablation study mentioned in line 290. 2. Clarity: In the abstract, the authors claim that the proposed method outperforms the state-of-the-art by a significant margin, but the method's name is not mentioned. When the term "calibration matrix" is first introduced in line 51, could you please provide a clear explanation? Does the proposed method predict both 2D and 3D semantic label maps? In Figure 2, which model requires training/finetuning and which one is frozen? How is the SAM prompted to predict the segmentation maps? 3. As a straightforward baseline, the authors could utilize Clip to generate coarse semantic masks and merge these masks with SAM segmentation results. It is recommended to present the performance in Table 2 and the corresponding figure. The suggested experiments should be evaluated on the large-scale Scannet dataset. Furthermore, if the 2D foundation models already perform well, the motivation behind the "Prediction Consistency Regularization" is not clear. Why is there a need to fine-tune your own model? 4. Although the demonstrated segmentation results are close to the ground truth, the mIoU scores for both 2D and 3D remain below 35. 5. The improvement achieved through "Latent Space Consistency Regularization" appears to be minimal for both 2D and 3D segmentation, despite being listed as an important contribution. 6. Please include citations for NeRF-based and mesh-based scene segmentation: [1] Atlas: End-to-end 3d scene reconstruction from posed images (supervised and unsupervised) [2] Decomposing nerf for editing via feature field distillation [3] Nerf-sos: Any-view self-supervised object segmentation on complex scenes [4] Unsupervised Multi-View Object Segmentation Using Radiance Field Propagation

Questions

As the reviewer's question mainly lies in the clarity and writing, please refer to the "Weakness" section.

Rating

4: Borderline reject: Technically solid paper where reasons to reject, e.g., limited evaluation, outweigh reasons to accept, e.g., good evaluation. Please use sparingly.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

2 fair

Contribution

2 fair

Limitations

Please refer to the "Weakness" section.

Authorsrebuttal2023-08-18

Anticipating the response from Reviewer rM41

Thank you once again for your invaluable contribution to our paper. We hope that the response provided above adequately addresses your concerns. If you find any aspects of the paper unclear or have further questions, please don't hesitate to share them with us. We are committed to resolving any uncertainties and ensuring the clarity and quality of our work. Your response is highly appreciated. Thank you!

Reviewer Mxk96/10 · confidence 5/52023-07-03

Summary

This paper proposes using SAM to improve the quality of segmentation masks generated by CLIPs. The refined masks are then used as pseudo-labels to train networks that excel at segmentation tasks, following a similar training process to MaskCLIP+. Based on this idea, a two-stage self-training process is proposed to simultaneously train 2D and 3D networks with SAM-refined masks and self-generated masks iteratively. To mitigate mask noise, the authors further propose latent space consistency regularization to transfer knowledge from the SAM feature space to segmentation networks. The final trained segmentation networks demonstrate impressive performance on both 2D and 3D segmentation datasets.

Strengths

(1) The overall method is both simple and effective. Although the authors use SAM trained with labeled data, which means it is not strictly label-free, I believe that it is acceptable given the performance improvement and contribution to the community. (2) The performance is impressively prefect. (3) The ablation study proves the gains of the main contributions of the paper. (4) The paper is well-written and easy to read and understand except for some minor issues.

Weaknesses

(1) In this paper, the authors chose DeeplabV3 as the 2D backbone and MinkowskiNet34 as the 3D backbone. However, previous works such as MaskCLIP used DeeplabV2 as the 2D backbone and CLIP2Scene used MinkowskiNet14 as the 3D backbone. It is obviously that the backbones used in this paper have better performance, so it may not be entirely fair to compare them with previous works. (2) I'm a bit confused about the details of the Label Refinement by SAM. SAM often outputs very fine-grained masks and produces multi-level masks, while CLIP also often outputs masks with a lot of discrete noise points. How to handle these situations? Can the author provide a simple pseudo-code to describe this part of the processing? (3) I think the baseline selected by this paper is not appropriate. While the main contributions of this paper are the improvement of pseudo-labeling quality and more effective training strategies, the baseline should be a basic method that incorporates both of these points. For example, I think that for 2D scenes, MaskCLIP+ is a better baseline than CLIP. (4) I think the presentation of Table 2 could be improved. Except for the experimental settings in the first row, the configurations in the other rows are abridged versions of the full configuration in a specific structure. To better present the impact of each structure on the final configuration, I suggest comparing the results of each ablation experiment with the performance drop of the full configuration results. (5) MaskCLIP++ -> MaskCLIP+.

Questions

Please refer to Weaknesses.

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.

Soundness

4 excellent

Presentation

3 good

Contribution

4 excellent

Limitations

Yes, the authors adequately addressed the limitations.

Reviewer 9W6n6/10 · confidence 4/52023-07-07

Summary

The paper proposes a new method for jointly learning 2D and 3D segmentation models that can be queried with open-world text queries. The paper's main contribution is a new Cross-modality Noisy Supervision to reduce the influence of noise on the predictions. The proposed method uses a SAM model to predict masks for every pixel and then uses those noises to regularize CLIP's noisy dense pixel features. The results shown in the paper show that the proposed method outperforms baselines in both 2D and 3D segmentation. I have a couple of questions and suggestions to improve the paper and I am willing to increase my score if these are sufficiently addressed in the rebutal.

Strengths

The paper is overall well-written and mostly easy to follow (apart from the mathematical notation, see comments below). The paper provides a pretty comprehensive overview of the literature up until CVPR 2023, and I also liked the qualitative results and Figure 1, which was informative.

Weaknesses

I have some suggestions (in no particular order) to improve the paper: 1. I take some issue with the claim that the method is the first to combine SAM and CLIP for label-free scene understanding. Multiple different methods have combined 2D segmentation models with CLIP, and as these systems are built modularity, they have already been able to switch out their segmenters for SAM. Two papers that come to mind for this are OvSeg (2D segmentation) and ClipFusion (3D segmentation). I suggest that this claim be softened. 2. Figure 2 is very non-intuitive in my opinion and I did not fully understand the flow of information even after multiple reads. Maybe it would be good to improve this figure. 3. Section 3.2 seems to be overly mathy and the notation is extremely convoluted. Consider cleaning up the super/sub scripts, and I also have a few additional comments: * Equations 2 and 3 take a lot of space do not add anything to my understanding of the method. I suggest removing them. * It is not necessary to define the range of each variable. * Equations 5/6 are extremely convoluted. Consider simpler notation. * In equation 1, I think there is an error, and the dot product should be between c_i^T, t_l 4. Regarding the results in Table 1, the OpenScene results seem to be much lower compared to the ones reported in the original paper. How can this be explained? Was the baseline retrained for this paper?

Questions

I have a few clarification questions as well: 1. Line 159: What is meant by pixel-point calibration? Are the camera poses not known? Is there an additional step required apart from simple projection? If not, this section can be shortened in my opinion. 2. Line 160: What is meant by the domain gap here? And why does this matter in the first place? As I see it, the method tries to align the features from the 3D network to the 2D CLIP/SAM space, and the domain gap should not matter here, right? 3. Line 209: In the current method, you use cosine similarity to force 2D/3D network outputs to be similar to the SAM feature output. Is this objective not in contrast to enforcing similarity between CLIP features and the 3D/3D network outputs? Especially when looking at the results in Table 2, this specific regularization does not seem to add a lot in terms of performance (Latent Space Consistency Regularization)

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

The paper does not discuss limitations and broader impact in the paper.

Reviewer inkR2023-08-16

Reply to Authors

Thank authors for providing supporting answers and results. I think the authors have almost answered and provided results for most of my queries and I hope they would add these results in the final version of the paper. I had one query though the t-SNE visualization and statement given "By our latent space regularization, for those pixels/points with the same semantics (the same color), their feature space tends to be more compact. Meanwhile, for those pixels/points with different semantics (the different colors), their feature space tends to be more distinguishable." does not quite hold good for all the visualization as presented in the result-pdf file. Effect on latent feature space still needs some further investigation to evaluate its impact and identify probable weakness in the current approach. Never the less I will change my rating to Accept from Weak Accept after author's rebuttal.

Authorsrebuttal2023-08-17

Authors' Response to Reviewer inkR

Thanks so much for donating your time to our paper. As suggested, we will add these results in the final version. Besides, we have thought about multiple potential ways to better verify the effectiveness of the latent space feature regularization: - We could apply linear probing on the latent space feature, i.e., using a few labelled data to fine-tune the network and compare the improvements to that without latent space feature regularization. - We could unsupervised cluster the feature and qualitatively evaluate the performance using ground-truth labels. - We could set different anchor points and visualize the similarity map based on the similarity between the anchor point feature and all other point features. We will investigate and discuss the probable weakness in the final version. Thanks again for your valuable comments and positive feedback!

Reviewer Mxk92023-08-17

The rebuttal solves most of my concerns. I keep my rating as weak accept. I wish that the authors correct the minor errors and further improve quality of the paper in the revised version.

Authorsrebuttal2023-08-17

Authors' Response to Reviewer Mxk9

We thank Reviewer Mxk9 for participating in the Author-Reviewer discussion session and providing positive feedback. As suggested, we will revise the manuscript rigorously based on the reviewers' comments. --- Last but not least, we thank Reviewer Mxk9 again for the time and effort devoted and the valuable comments drawn during this review.

Reviewer 9W6n2023-08-17

I want to thank the reviewers for their in-depth explanations and clarifications. They have effectively addressed most of my concerns, and I will raise my score to a weak accept. I think that this paper is timely and addresses an interesting issue.

Authorsrebuttal2023-08-18

Authors' Response to Reviewer 9W6n

We thank Reviewer 9W6n for acknowledging that our rebuttal is helpful and the topic of this paper is interesting. We will revise the manuscript accordingly based on your comments and further improve the quality of this work. --- Last but not least, we thank Reviewer 9W6n again for the time and effort devoted and the valuable comments drawn during this review.

Reviewer rM412023-08-21

Thanks for the authors' response

After reviewing the authors' response to the posted inquiries, additional crucial details regarding the methodology have been incorporated, along with an explanation concerning the effectiveness of the "Latent Space Consistency Regularization." My anticipated score improvement is grounded in the following reasons: 1. The proposed approach is intriguing, employing two foundational models for scene comprehension, resulting in commendable outcomes. 2. The seemingly straightforward "Label Refinement" + Post-Refine technique exhibits substantial power in enhancing overall accuracy. 3. It's advisable to incorporate more representative outcomes within the main paper, ensuring quantitative figures align with qualitative findings: the overall mIoU is still far from the supervised conterpart, but the visualizations in the main draft seems perfect. 4. While the integration of diverse model insights can undoubtedly heighten accuracy, it's essential for the authors to perform runtime and memory consumption analyses as outlined in Table 1, ensuring equitable comparisons.

Authorsrebuttal2023-08-21

Authors' Response to Reviewer rM41

We appreciate your acknowledgement of the interest in our method and the effectiveness of the Label Refinement module. Additionally, we are committed to enhancing our paper based on your suggestions: 1. A major contribution of the paper is that we study how vision foundation models enable networks to comprehend 2D and 3D environments without relying on labelled data. Thanks for your acknowledgement. 2. The Label Refinement is an important module in our proposed Cross-modality Noisy Supervision framework, which is one of the major contributions of our method. Thanks for your acknowledgement. 3. Due to space constraints, we have included more qualitative results, including failure cases, in the supplementary materials. In the final version, we intend to present a more comprehensive selection of representative results within the main paper. Thanks for your understanding. 4. We would like to emphasize that our method exhibits comparable runtime and memory consumption to other state-of-the-art approaches during the inference stage. This similarity arises from the shared utilization of 2D and 3D backbones with other methods. It's important to note that the vision foundation models (CLIP and SAM) are exclusively employed during the training phase. In our final version, we will provide further elaboration on these technical details. Thank you for highlighting this aspect. Once again, we express our gratitude for your feedback. We hope that our response addresses all of your concerns. If you identify any additional improvements within our paper, please do not hesitate to share them with us.

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC