Weaknesses
The major weakness of this article in the context of submission to ICLR is relevancy. The main contribution concerns a theoretical analysis of the bounds of representation capacity, which appears pertinent for ICLR. While there is little discussion on learning, theoretical analysis on the information capacity of a learnable representation could certainly fit within the scope of ICLR. However, there are two axes for concern with relevancy: the relevancy of VSAs, and the relevancy of the provided analysis.
The pertinence of VSAs to learning representations is described in the introduction relying on previous publications on VSA. Of those references, many are unpublished and a single author (D Kleyko) is cited twelve times, with some citations being simply for demonstration of application cases "biological times series data (Kleyko et al., 2018a; 2019b; Burrello et al., 2019)." There are two sentences in the current article relating VSAs to other, more common, forms of representation learning, notably deep learning, and the example provided is unclear (see Questions). From my understanding, VSAs are similar to very wide dense neural network layers, and are in some cases binarized. If this is the case, relating VSAs to existing work on high-dimensional neural network layers could lead to better understanding and greater relevancy of the approach. See, for example:
Jacot, A., Gabriel, F., & Hongler, C. (2018). Neural tangent kernel: Convergence and generalization in neural networks. Advances in neural information processing systems, 31.
Arora, S., Du, S. S., Hu, W., Li, Z., Salakhutdinov, R. R., & Wang, R. (2019). On exact computation with an infinitely wide neural net. Advances in neural information processing systems, 32.
The second relevancy concern is with the results of the analysis. This article exhaustively covers over 15 different bounds for VSAs, and pages 5 through 9 are almost exclusively used for the computation of these bounds. As such, the findings are never discussed or analyzed. Do the proven bounds differ from or build on existing understanding of this bound? Does the understanding of these bounds increase the utility of VSAs, or give insight into how their representations should be formed? Section 3.2 goes in the right direction in its discussion of Hopfield networks, however it stops short of analyzing the proven bounds. See, for example, the following works which study information capacity properties in order to improve or analyze representation learning:
Cheng, H., Lian, D., Gao, S., & Geng, Y. (2018). Evaluating capability of deep neural networks for image classification via information plane. In Proceedings of the European Conference on Computer Vision (ECCV) (pp. 168-182).
Gabrié, M., Manoel, A., Luneau, C., Macris, N., Krzakala, F., & Zdeborová, L. (2018). Entropy and mutual information in models of deep neural networks. Advances in Neural Information Processing Systems, 31.
Von Oswald, J., Niklasson, E., Randazzo, E., Sacramento, J., Mordvintsev, A., Zhmoginov, A., & Vladymyrov, M. (2023, July). Transformers learn in-context by gradient descent. In International Conference on Machine Learning (pp. 35151-35174). PMLR.