Summary
This paper aims to estimate the depth map in a dynamic scene environment.
To this end, the paper proposed 1) two separate module architectures that estimate rigid scene flow and residual scene flow and 2) a motion initialization method to enable a stable learning process.
However, the technical and performance comparisons with the previous self-supervised depth estimation, self-supervised scene flow estimation, self-supervised motion segmentation, and self-supervised optical flow are insufficient.
The current manuscript needs a lot of modification.
Weaknesses
W1. Technical and performance comparisons with the recent self-supervised scene flow estimation
The proposed method needs to describe its originality and superiority compared to recent scene flow estimation on the KITTI, nuScene, or Waymo datasets.
- Xiang, Xuezhi, et al. "Self-supervised learning of scene flow with occlusion handling through feature masking." Pattern Recognition 139 (2023): 109487.
- Jiao, Yang, Trac D. Tran, and Guangming Shi. "Effiscene: Efficient per-pixel rigidity inference for unsupervised joint learning of optical flow, depth, camera pose and motion segmentation." Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2021.
- Hur, Junhwa, and Stefan Roth. "Self-supervised monocular scene flow estimation." Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2020.
- Hur, Junhwa, and Stefan Roth. "Self-supervised multi-frame monocular scene flow." Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2021.
- Guizilini, Vitor, et al. "Learning optical flow, depth, and scene flow without real-world labels." IEEE Robotics and Automation Letters 7.2 (2022): 3491-3498.
W2. Technical and performance comparisons with the recent self-supervised motion segmentation
The proposed method needs to describe its originality and superiority compared to recent motion segmentation on the KITTI, nuScene, or Waymo datasets.
- Liu, Liang, et al. "Unsupervised Learning of Scene Flow Estimation Fusing with Local Rigidity." IJCAI. 2019.
- Jiao, Yang, Trac D. Tran, and Guangming Shi. "Effiscene: Efficient per-pixel rigidity inference for unsupervised joint learning of optical flow, depth, camera pose and motion segmentation." Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2021.
- Xiang, Xuezhi, et al. "Self-supervised learning of scene flow with occlusion handling through feature masking." Pattern Recognition 139 (2023): 109487.
W3. Technical and performance comparisons with the recent self-supervised optical flow estimation
The proposed method needs to describe its originality and superiority compared to recent optical flow estimation on the KITTI, nuScene, or Waymo datasets.
- Teed, Zachary, and Jia Deng. "Raft: Recurrent all-pairs field transforms for optical flow." Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part II 16. Springer International Publishing, 2020.
- Jonschkowski, Rico, et al. "What matters in unsupervised optical flow." Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part II 16. Springer International Publishing, 2020.
- Zhao, Wang, et al. "Towards better generalization: Joint depth-pose learning without posenet." Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2020.
W4. Technical and performance comparisons with the recent self-supervised depth estimation
The proposed method needs to describe its originality and superiority compared to recent self-supervised depth estimation on the KITTI, nuScene, or Waymo datasets.
- Watson, Jamie, et al. "The temporal opportunist: Self-supervised multi-frame monocular depth." Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2021.
- PLADE-Net: Towards Pixel-Level Accuracy for Self-Supervised Single-View Depth Estimation with Neural Positional Encoding and Distilled Matting Loss, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2021.
- Guizilini, Vitor, et al. "Multi-frame self-supervised depth with transformers." Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2022.