Weaknesses
1. My major concern lies in the evaluation. SemanticKITTI is a small-scale dataset and does not adequately address dynamic objects in the creation of ground truth SSC labels. The authors are suggested to conduct experiments on larger and more diverse datasets, such as KITTI-360, nuScenes, and Waymo, to better validate their approach. The authors are suggested to consider using SSCBench [1], which employs the same format as SemanticKITTI. Cross-dataset evaluation on SSCBench would significantly strengthen the paper. Furthermore, Occ3D [2] could also be utilized for evaluation, though it uses a different format than SemanticKITTI.
2. My second concern pertains to the computational overhead introduced by the proposed method. The test-time adaptation involves multiple modules, such as creating binary occupancy labels and training two different models at each timestamp, which can significantly increase computational demands. I recommend that the authors conduct a detailed computational analysis, including an evaluation of the computation-performance tradeoff and the frames per second (FPS) achieved.
3. The literature survey is notably incomplete. Several important SSC works are missing from the related work section, including both datasets and methods [1-7]. A more comprehensive review of the existing literature is necessary to provide proper context and background for your study. Meanwhile, the idea of Line of Sight has been studied in prior works, such as 3d object detection [8] and point cloud registration [9,10]. The authors are suggested to add the missing references.
4. The current approach is tailored to LiDAR data, which may limit its applicability to other sensor modalities or combined sensor data. The authors are suggested to add more details about how to use their method in the camera-based SSC methods.
[1] Li, Y., Li, S., Liu, X., Gong, M., Li, K., Chen, N., Wang, Z., Li, Z., Jiang, T., Yu, F. and Wang, Y., 2023. Sscbench: A large-scale 3d semantic scene completion benchmark for autonomous driving. arXiv preprint arXiv:2306.09001.
[2] Tian, X., Jiang, T., Yun, L., Mao, Y., Yang, H., Wang, Y., Wang, Y. and Zhao, H., 2024. Occ3d: A large-scale 3d occupancy prediction benchmark for autonomous driving. Advances in Neural Information Processing Systems, 36.
[3] Cao, A.Q. and De Charette, R., 2022. Monoscene: Monocular 3d semantic scene completion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 3991-4001).
[4] Shi, Y., Li, J., Jiang, K., Wang, K., Wang, Y., Yang, M. and Yang, D., 2024, March. PanoSSC: Exploring Monocular Panoptic 3D Scene Reconstruction for Autonomous Driving. In 2024 International Conference on 3D Vision (3DV) (pp. 1219-1228). IEEE.
[5] Li, Y., Yu, Z., Choy, C., Xiao, C., Alvarez, J.M., Fidler, S., Feng, C. and Anandkumar, A., 2023. Voxformer: Sparse voxel transformer for camera-based 3d semantic scene completion. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 9087-9098).
[6] Zhang, Y., Zhu, Z. and Du, D., 2023. Occformer: Dual-path transformer for vision-based 3d semantic occupancy prediction. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 9433-9443).
[7] Huang, Y., Zheng, W., Zhang, Y., Zhou, J. and Lu, J., 2023. Tri-perspective view for vision-based 3d semantic occupancy prediction. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 9223-9232).
[8] Hu, P., Ziglar, J., Held, D. and Ramanan, D., 2020. What you see is what you get: Exploiting visibility for 3d object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 11001-11009).
[9] Ding, L. and Feng, C., 2019. DeepMapping: Unsupervised map estimation from multiple point clouds. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 8650-8659).
[10] Chen, C., Liu, X., Li, Y., Ding, L. and Feng, C., 2023. Deepmapping2: Self-supervised large-scale lidar map optimization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 9306-9316).