Analysis of the Rate of Convergence of an Over-Parametrized Deep Neural Network Estimate Learned by Gradient Descent

Estimation of a regression function from independent and identically distributed random variables is considered. The <inline-formula> <tex-math notation="LaTeX">$L_{2}$ </tex-math></inline-formula> error with integration with respect to the design measure is used as an error criterion. Over-parametrized deep neural network estimates are defined which are based on a special network topology, which use a special random initialization and where all the weights are learned by the gradient descent. It is shown that the expected <inline-formula> <tex-math notation="LaTeX">$L_{2}$ </tex-math></inline-formula> error of these estimates converges to zero with the rate close to <inline-formula> <tex-math notation="LaTeX">$n^{-1/(1+d)}$ </tex-math></inline-formula> in case that the regression function is Hölder smooth with Hölder exponent <inline-formula> <tex-math notation="LaTeX">$p \in [{1/2,1}]$ </tex-math></inline-formula>. In case of an interaction model where the regression function is assumed to be a sum of Hölder smooth functions where each of the functions depends only on <inline-formula> <tex-math notation="LaTeX">$d^{*}$ </tex-math></inline-formula> of of d components of the design variable, it is shown that these estimates achieve the corresponding <inline-formula> <tex-math notation="LaTeX">$d^{*}$ </tex-math></inline-formula>-dimensional rate of convergence.

Paper

References (56)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC