We appreciate the reviewer’s constructive feedback and the opportunity to address their concerns. Below, we respond to each point raised in detail:
1. Conflation of Dimensional Collapse and Neural Collapse:
We thank the reviewer for pointing out this issue. In the revised version of the paper, we have corrected all references to appropriately use the term dimensional collapse where applicable.
2. Clarity Regarding Orthogonal Structures vs. Simplex ETF:
We appreciate the reviewer’s insightful comments regarding the Simplex ETF and its optimal cosine values. As discussed in Section 3.2, while the Simplex ETF theoretically achieves superior angular separation, we observed in practice that it is rarely achieved by merely minimizing the loss. Instead, the optimization often results in hyperplane-based formulations with zero-mean embeddings.
Building on this observation, our proposed method, CLOP, encourages embeddings to align with different subspaces. This results in full-rank embeddings, rather than degenerate hyperplanes. To further substantiate our claims, we have added new experiments comparing the use of orthonormal structures versus Simplex ETF prototypes. These results are presented in Figures 6 and 7 of the revised paper.
3. Limited Experiments on Small-Scale Datasets:
We acknowledge the reviewer’s concern about the generalizability of our findings given the limited dataset scope. In the revised version, we have included additional experiments on the ImageNet-1K dataset (Figures 6 and 7), addressing concerns about scalability and supporting the broader applicability of our approach.
4. Performance Improvements on Larger Batch Sizes:
We acknowledge the reviewer’s observation regarding the performance improvements under larger batch sizes. Our aim is to demonstrate that under small batch sizes, our method performs relatively well compared to competing approaches with larger batch sizes, making it particularly suitable for scenarios constrained by computational memory resources. For larger batch sizes, the increased diversity of positive and negative pairs naturally reduces the risk of collapse, which lessens the relative impact of our method. Nonetheless, we observe that our approach remains competitive in these settings.
We thank the reviewer again for their valuable feedback. We hope that the revisions adequately address the concerns raised.