General Response (Cont'd)
**Q4 - Limitations (`834L`, `fKvz`, `K12Q`)**
We identify two limitations of the current inference pipeline. First, it is sensitive to prompts. For text-conditioned tasks, minor variations in textual scene descriptions can lead to large quality differences in the output. Second, a typical failure mode is the image parsing error. For image-conditioned tasks, input images are parsed with the backbone visual language model. With the same input image, parsing results have high variance across multiple inference runs. Failure case results are updated on the [project page](https://sclg-page.github.io/#failure).
While the current paper provides a viable inference method for the proposed representation that leverages the commonsense knowledge and code-writing capability of LMs for scene generation and editing tasks, we fully agree with reviewers `834L`, `fKvz`, and `K12Q` that addressing the weaknesses inherited from LMs would provide further improvement in robustness and output quality for downstream tasks. We leave these as exciting future directions to improve the inference of the Scene Language, and emphasize that the representation itself is the focus and the core contribution of this paper.
___
References
1. Gao, G., Liu, W., Chen, A., Geiger, A., & Schölkopf, B. (2024). GraphDreamer: Compositional 3D scene synthesis from scene graphs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 21295-21304).
2. Zhang, R., Isola, P., Efros, A. A., Shechtman, E., & Wang, O. (2018). The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 586-595).
3. Bahmani, S., Skorokhodov, I., Rong, V., Wetzstein, G., Guibas, L., Wonka, P., ... & Lindell, D. B. (2024). 4D-fy: Text-to-4D generation using hybrid score distillation sampling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 7996-8006).
4. Ilharco, G., Wortsman, M., Carlini, N., Taori, R., Dave, A., Shankar, V., Namkoong, H., Miller, J., Hajishirzi, H., Farhadi, A., & Schmidt, L. (2021). OpenCLIP (0.1). Zenodo.
5. Huang, Z., He, Y., Yu, J., Zhang, F., Si, C., Jiang, Y., ... & Liu, Z. (2024). VBench: Comprehensive benchmark suite for video generative models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 21807-21818).
6. Teed, Z., & Deng, J. (2020). RAFT: Recurrent all-pairs field transforms for optical flow. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part II 16 (pp. 402-419). Springer International Publishing.