Summary
The authors introduce a generalization of RKHS for $C^*$ algebra valued kernels, called RKHM; they build networks by composing sequentially elements taken from a collection of RKHMs, one RKHM per layer. They prove generalization bounds for those networks. These networks output matrices at each layer.
Strengths
I believe that a strength of the paper is that the generalization bound obtained in this paper for RKHM is better than the known ones for vvRKHS. It is unclear to me if RKHMs generalize vvRKHSs: maybe considering RKHM in the commutative $C^*$ of diagonal matrices is a way to represent a vvRKHS as an RKHM. If it is the case then the result of this paper is a better generalization bound for vvRKHS.
Weaknesses
I feel that the paper is sometimes difficult to read. For example the name '$\mathcal{A}$-valued positive definite kernel' does not refer to the space $\mathcal{X}$ that characterizes the domain of the kernel; when defining deep RKHM knowing this information could help make explicit what the domain of $k_j$ as being $A_{j-1}\times A_{j-1}$.
Two remarks in this direction are on some notations:
- line 83: shouldn't the content of the brace in the definition of $M_{k,0}$ be $\sum_{i=1}^{n} \phi(x_i) c_i \vert n\in \mathbb{N}, (c_i\in A, i\leq n), (x_i\in \mathcal{X}, i\leq n) $
- equation line 195, maybe the notation: $(f_j\in \mathcal{M}_j)_j$, could recall that the optimization is over all the RKHMs that define the networks.
Questions
In 'Connection with neural tangent kernel', is the aim of this paragraph to define a neural tangent kernel for deep RKHM?
In 'Comparison to CNN', how long does the training take and what is the memory consumption of RKHM and CNN?
Rating
6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.
Confidence
3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.
Limitations
In Conclusion and Limitations, I feel that the statement 'connections with existing studies such as CNNs and neural tangent kernel' is a bit of a strong statement as the authors explain in section 6.1 that CNN and RKHM do not really relate to one another and it is unclear to me how deep the connection to neural tangent kernel is.