Summary
The manuscript investigates the mean-field Langevin dynamics (MFLD) and improves the particle approximation error by removing the dependency on the log-Sobolev inequality (LSI) constant. This is of relevance, as the LSI constant in general might deteriorate with the regularization coefficient, limiting the impact of prior bounds.
The Authors illustrate the applicability of their result with three examples.
1. An improved convergence of the MFLD,
2. sampling guarantees for the final stationary distribution,
3. uniform-in-time propagation of chaos w.r.t. the Wasserstein distance for the MFLD.
Strengths
- Improving the estimates of the particle approximation of the MFLD is of interest given its appearance in the learning problem of mean-field neural networks, which has been an interesting research topic pursued by different research groups during the last years. The paper is of purely theoretical nature and improves upon the state-of-the-art by removing the dependence of the approximation error constant on the LSI constant. The technical tools to do so, seem novel.
- The paper is well-written and -structured with the objectives and contribution clearly stated and pursued.
- In my opinion, the content of the paper could be also of interest beyond the scope of the MFLD, as propagation of chaos results with favorable constants are of interest in a wide variety of fields. The Authors may want to consider commenting on this.
Weaknesses
Apart from some minor questions addressed below, there is one point of critique that I would like to raise.
Namely, the Authors do not really motivate why one should expect that the particles approximation _does not_ depend on the LSI constant. This, in my opinion, would improve the reading experience of the paper.
Questions
- line 61: How does the LSI constant $\alpha$ deteriorate with the regularization parameter $\lambda$? Could you provide a sketch to give some more information? This could be also worth to be included and referenced in the manuscript.
And some more minor comments:
- In line 28, the Authors may want to add one further line of work, namely "Mean field analysis of neural networks: A law of large numbers" and "Mean field analysis of neural networks: A central limit theorem" by J Sirignano, K Spiliopoulos, to exhaustively cover the literature.
- In regards of (2), the Authors might want to mention that $\nabla \frac{\delta \mathcal{F}}{\delta \mu}$ is also known as the Wasserstein gradient $\nabla_W \mathcal{F}$.
- lines 35, 36: Could you add references here?
- lines 99, 100: Write $\mathcal{P}(\mathbb{R}^d)$.
- line 100: Not sure what you mean with "it follwos that" here
- line 157: $N$ not $d$
- line 234: identical to
- lines 266-268: This sentence sounds a bit complicated.
- line 301: Discussion
Limitations
The Authors do point out some limitations in the Conclusions, which are reasonable and in my opinion justifiable.
They could further emphasize some limitations of their result following from Assumption 3. In particular the uniform boundedness of $h$ seems to be a restriction, but since the class of covered neural networks is still fair, I would not consider it as a substantial limitation. Yet, it could be highlighted.