Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene Understanding

Complex 3D scene understanding has gained increasing attention, with scene encoding strategies playing a crucial role in this success. However, the optimal scene encoding strategies for various scenarios remain unclear, particularly compared to their image-based counterparts. To address this issue, we present a comprehensive study that probes various visual encoding models for 3D scene understanding, identifying the strengths and limitations of each model across different scenarios. Our evaluation spans seven vision foundation encoders, including image-based, video-based, and 3D foundation models. We evaluate these models in four tasks: Vision-Language Scene Reasoning, Visual Grounding, Segmentation, and Registration, each focusing on different aspects of scene understanding. Our evaluations yield key findings: DINOv2 demonstrates superior performance, video models excel in object-level tasks, diffusion models benefit geometric tasks, and language-pretrained models show unexpected limitations in language-related tasks. These insights challenge some conventional understandings, provide novel perspectives on leveraging visual foundation models, and highlight the need for more flexible encoder selection in future vision-language and scene-understanding tasks. Code: https://github.com/YunzeMan/Lexicon3D

Paper

Similar papers

Peer review

Reviewer vFsT6/10 · confidence 4/52024-07-01

Summary

This paper presents a comprehensive analysis of leveraging visual foundation models for complex 3D scene understanding. The authors unify three kinds of feature representation: image, video and 3D in a unified paradigm and analyze the effectiveness of those representations on different kinds of 3D tasks.

Strengths

1. The study is thorough and meaningful. I think it provides interesting finds to the 3D vision community. 2. Experiments are extensive and comprehensive. The chosen tasks and methods are representative.

Weaknesses

1. This work lacks a unified conclusion. The experimental observations are independent (in some extent like an experimental report rather than a research paper). It is better for the authors to summarize the core underlying principles, or design some models according to the observation, which may make this work more applicable. 2. Since indoor 3D perception is mainly applied in embodied AI system, I suggest the author further study 3D scene understanding in an online setting [1, 2] rather than the current offline setting, which can be directly adopted in real-world robotic tasks. [1] Fusion-aware point convolution for online semantic 3d scene segmentation, CVPR 2020 [2] Memory-based Adapters for Online 3D Scene Perception, CVPR 2024

Questions

No

Rating

6

Confidence

4

Soundness

3

Presentation

4

Contribution

3

Limitations

Limitation is well addressed.

Reviewer yLXP5/10 · confidence 4/52024-07-10

Summary

This paper explores the importance of scene encoding strategies in the context of 3D scene understanding, an area gaining significant attention. The authors investigate the optimal encoding methods for various scenarios, addressing the lack of clarity compared to image-based approaches. They conduct an in-depth study of different visual encoding models, evaluating their strengths and weaknesses across multiple scenarios. The study examines seven foundational vision encoders, including image-based, video-based, and 3D models, across four key tasks: Vision-Language Scene Reasoning, Visual Grounding, Segmentation, and Registration. The findings reveal that DINOv2 performs exceptionally well overall, video models are particularly effective for object-level tasks, diffusion models excel in geometric tasks, and language-pretrained models have unexpected limitations in language-related tasks. These results challenge existing assumptions, provide fresh insights into the use of visual foundation models, and underscore the need for adaptable encoder selection in future vision-language and scene-understanding research.

Strengths

1, The paper is well-written and easy to understand. 2, It surveys most of the current visual foundation models.

Weaknesses

1, I do not see any new insights from this work. Numerous previous studies [1, 2, 3] have already demonstrated that leveraging foundation models can improve 3D understanding. 2, This work simply projects visual foundation models from different views onto the 3D point cloud and fine-tunes them. There is no specific design to better integrate/distill these features. 3, The latest 3D foundation model, Uni3D, is not discussed in this work. 4, Inference time is a significant issue, as the input requires multiple-view images. [1] Multi-View Representation is What You Need for Point-Cloud Pre-Training [2] Bridging the Domain Gap: Self-Supervised 3D Scene Understanding with Foundation Models [3] CLIP2Scene: Towards Label-efficient 3D Scene Understanding by CLIP [4] Uni3D: Exploring Unified 3D Representation at Scale

Questions

See weakness.

Rating

5

Confidence

4

Soundness

2

Presentation

3

Contribution

1

Limitations

Yes

Reviewer yLXP2024-08-08

Thanks for the clarification. I will increase my score to 5.

Authorsrebuttal2024-08-08

Thank you for your positive feedback!

We appreciate the reviewer for the positive feedback. Your constructive comments and suggestions are indeed helpful for improving the paper. Also, many thanks for raising the score. We will continue to improve our work and release the code. If the reviewer has any follow-up questions, we are happy to discuss!

Reviewer 4P6H6/10 · confidence 5/52024-07-12

Summary

This paper examines various scene encoding methods for 3D scene understanding, encompassing image, video, and 3D models. It explores four distinct tasks: registration, scene reasoning, visual grounding, and segmentation. The experimental results indicate that different encoding techniques excel in different tasks, underscoring the importance of selecting appropriate encoders for enhanced understanding in 3DVL.

Strengths

1. This paper provides thorough analysis on different 2D and 3D performance on scene understanding. Currently, we have not seen these features compared in one framework 2. This probing framework and the insights from the experimental results are crucial for the 3D vision-language community.

Weaknesses

1. Details of using 3D feature field for these tasks should be discussed. What is the baseline model for 3D grounding, QA and registration. These information should be elaborated in appendix. 2. The combination of 3D and 2D feature should be studied in different tassk.

Questions

1. In 2D visual grounding, results using image feature and swin3D features are worse than M3DRef, which model do you use to conduct visual grounding? 2. See weakness above, details of these experiments should be included to make these results more convincing.

Rating

6

Confidence

5

Soundness

3

Presentation

3

Contribution

3

Limitations

More 3D encoders should be considered, like point transformer, pointnet++, and sparse conv UNet.

Reviewer q2Jg8/10 · confidence 4/52024-07-30

Summary

This paper conducts a large-scale study to answer the unexplored question: which method (among image-based, video-based, and 3D foundation models) performs the best in 3D scene understanding? The results show that DINOv2 demonstrates superior performance, video models excel in object-level tasks, diffusion models benefit geometric tasks, and language-pretrained models show unexpected limitations in language-related tasks.

Strengths

The paper is well-structured overall. The investigated question is very interesting to me. I also like the extensive experiments involved in this paper.

Weaknesses

(1). What about more advanced object-centric encoders like Segment Anything (SAM) for complex 3D scene understanding? Are they better or worse than LSeg? (2). Will the results be different when a different probing method is used (e.g., a pyramid network to aggregate multi-scale features from the foundation model)?

Questions

Please see 'Weaknesses'.

Rating

8

Confidence

4

Soundness

3

Presentation

3

Contribution

3

Limitations

The authors have discussed the limitations of this paper, and this paper has no direct negative societal impact.

Area Chair 2Gix2024-08-08

Dear reviewer, The author-reviewer interaction period has started. Please read the responses provided by the authors, respond to them early on in the discussion, and discuss points of disagreement. Thanks

Area Chair 2Gix2024-08-08

Dear reviewer, The author-reviewer interaction period has started. Please read the responses provided by the authors, respond to them early on in the discussion, and discuss points of disagreement. Thanks

Area Chair 2Gix2024-08-08

Dear reviewer, The author-reviewer interaction period has started. Please read the responses provided by the authors, respond to them early on in the discussion, and discuss points of disagreement. Thanks

Area Chair 2Gix2024-08-08

Dear reviewer, The author-reviewer interaction period has started. Please read the responses provided by the authors, respond to them early on in the discussion, and discuss points of disagreement. Thanks

Reviewer vFsT2024-08-08

The authors' rebuttal has solved most of my concerns. I will raise the score from 5 to 6.

Authorsrebuttal2024-08-08

Thank you for your positive feedback!

We appreciate the reviewer for the positive feedback. Your constructive comments and suggestions are indeed helpful for improving the paper. Also, many thanks for raising the score. We will continue to improve our work and release the code. If the reviewer has any follow-up questions, we are happy to discuss!

Reviewer q2Jg2024-08-13

Thank you for the rebuttal. My concerns are addressed by the authors. Therefore, I will keep my score.

Authorsrebuttal2024-08-13

Thank you for your positive feedback!

We appreciate the reviewer for the positive feedback. Your constructive comments and suggestions are indeed helpful for improving the paper. We will continue to improve our work and release the code. If the reviewer has any follow-up questions, we are happy to discuss!

Reviewer 4P6H2024-08-14

Thank the authors for the rebuttal

I have read the rebuttal and other reviews. Most of my concerns have been solved, so I will maintain my original score as 6: Weak Accept.

Authorsrebuttal2024-08-14

Thank you for your positive feedback!

We appreciate the reviewer for the positive feedback. Your constructive comments and suggestions are indeed helpful for improving the paper. We will continue to improve our work and release the code. If the reviewer has any follow-up questions, we are happy to discuss!

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC