Summary
This theory paper fits within a general framework in which one tries to get information on training of deep learning models using the formalism of the so-called Neural Tangent Kernel.
Specifically, the topic is smallest eigenvalue control for the NTK kernel, and the authors study the minimum eigenvalue under the assumption that one has datasets of unit norm and assume that the data are "well spread" i.e. they have controlled separation constants (and sometimes controlled covering, meaning the data are "uniformly spread"). The obtained bounds depend on this separation constant and on the input and output space dimensions.
The technique uses the so-called hemisphere transform and basic harmonic analysis on the sphere. These methods have not been used before for this particular problem.
Strengths
The studied problem is arguably relevant for dynamical study of NN evolution.
The techniques used are innovative within this field.
Weaknesses
The main weakness is that requirement on the data distribution to be delta-separated is not as "harmless" or "general" as the authors claim (furthermore, I did not find a justification of this claim in the paper; the authors just state that $\delta$-separation is "milder" than previous work requirements, without explaining why and without verifying that).
In practice, it is not trivial to ensure that a sample from a data distribution is well separated in the sense of theorem 8, or a Delone set with controlled constants, making it uniformly separated in the sense of Theorem 1. The assumption of iid data is in practice easier to justify, and checking for delta-separation may be itself a hard problem.
Questions
Main question:
A step for a good comparison to previous work is in having formulated Corollary 2, that is an "iid data analogue" of the main result of thm 1. However it is not clear how the bounds from previous works compare to this result. I suggest to put some effort to explicit this comparison in the most explicit way possible.
Other minor observations and questions:
1) the notion of "$\delta$-separatedness" is a terminology used in the community for the case that points are at minimum distance larger or equal than $\delta$. The notion is not the same as used in this paper, and defined in line 44. Also, at 3 instances in the paper the notion of "collinearity" is used, which is a bit misleading: any two points are collinear. So I suggest to replace "collinearity" with something more explicit such as "being on the same line through the origin" and that the notion of "$\delta$-separated" is either called by a different name, or that it be emphasized the difference with the usual notion.
2) In the paragraph before line 39, there is an instance in which $\mathbb R^{d\times n}$ should be replaced by $\mathbb R^{d_0\times n}$.
3) Section 2 has a large overlap with the introduction. Could it be shortened or merged?
4) lines 141-142, about data in $\mathbb S^1$: this sentence is not clear to me, and it is not clear how passing from $\mathbb S^1\subset \mathbb R^2$ to $\mathbb S^1\times\{0\}\subset \mathbb R^3$ affects the constant $\delta'$ from Theorem 1.
5) line 243-244 and lines 313-315: the fact that data are required to be $\delta$-separated has to be mentioned, since it restricts generality.
Limitations
The main concerns were mentioned in the "weaknesses" part.