Weaknesses
1. Writing: The use of the term "module" is vague and inconsistent throughout the paper. Does it refer to a computational block, a neural network component, or a specific feature processing unit? This ambiguity undermines the reader's understanding of the architecture.
Phrases like "manages noise independently" and "compensates limitations" lack specificity. How the proposed modules achieve these outcomes is poorly explained, leaving the mechanism unclear. The writing is repetitive and verbose, making it difficult to parse the main contributions and technical details. For example, the description of the Channel Compression Layer (CCL) is overly long, yet still leaves the reader unclear about its exact role. Key claims, such as computational efficiency and noise robustness, are repeated multiple times without sufficient empirical backing or theoretical analysis.
2. Novelty: The primary claim of novelty, a "parallel connection" architecture, is incremental and not supported by any theoretical insights or innovative algorithmic developments. This feature feels more like an engineering optimization rather than a fundamental research contribution. Many components of the architecture (e.g., PointPillars, Sparse ResNet, attention mechanisms) are off-the-shelf methods, and the way they are combined lacks originality. The authors fail to show how these combinations lead to insights or advancements that are relevant to ICLR.
3. Relevance: The paper’s focus on a specific application (collaborative perception for autonomous vehicles) lacks sufficient theoretical contributions or broader applicability to the machine learning community. The proposed method is narrowly tailored to V2X scenarios, which limits its relevance for a broader audience. The paper does not introduce new machine learning techniques or paradigms but rather applies existing concepts in a slightly modified architecture.
4. Related works: The paper could benefit from a more comprehensive literature review, as some highly relevant works about collaborative perception are missing [1-6]. Including a broader range of recent studies would provide a stronger context for the contributions and better situate the proposed framework within the current state of research.
[1] Li, Y., Ren, S., Wu, P., Chen, S., Feng, C. and Zhang, W., 2021. Learning distilled collaboration graph for multi-agent perception. Advances in Neural Information Processing Systems, 34, pp.29541-29552.
[2] Li, Y., Ma, D., An, Z., Wang, Z., Zhong, Y., Chen, S. and Feng, C., 2022. V2X-Sim: Multi-agent collaborative perception dataset and benchmark for autonomous driving. IEEE Robotics and Automation Letters, 7(4), pp.10914-10921.
[3] Huang, S., Zhang, J., Li, Y. and Feng, C., 2024. Actformer: Scalable collaborative perception via active queries. ICRA 2024.
[4] Yang, D., Yang, K., Wang, Y., Liu, J., Xu, Z., Yin, R., Zhai, P. and Zhang, L., 2024. How2comm: Communication-efficient and collaboration-pragmatic multi-agent perception. Advances in Neural Information Processing Systems, 36.
[5] Su, S., Li, Y., He, S., Han, S., Feng, C., Ding, C. and Miao, F., 2023, May. Uncertainty quantification of collaborative detection for self-driving. In 2023 IEEE International Conference on Robotics and Automation (ICRA) (pp. 5588-5594). IEEE.
[6] Su, S., Han, S., Li, Y., Zhang, Z., Feng, C., Ding, C. and Miao, F., 2024. Collaborative multi-object tracking with conformal uncertainty propagation. IEEE Robotics and Automation Letters.