Regularity and Tailored Regularization of Deep Neural Networks, with application to parametric PDEs in uncertainty quantification

<p> In this paper we consider deep neural networks (DNNs) with a smooth activation function as surrogates for high-dimensional functions that are somewhat smooth but costly to evaluate. We consider the standard (non-periodic) DNNs as well as propose a new model of periodic DNNs which are especially suited for a class of periodic target functions when Quasi-Monte Carlo lattice points are used as training points. The primary contribution of this paper is the derivation of explicit bounds for all mixed derivatives of DNNs with respect to their input parameters. The bounds depend on the neural network parameters as well as the choice of activation function, with explicit constants. These bounds are fully general and remain independent of both the target function and the training data. By imposing restrictions on the network parameters to match the regularity features of the target functions, we prove that DNNs with <inline-formula content-type="math/mathml"> <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" alttext="upper N"> <mml:semantics> <mml:mi>N</mml:mi> <mml:annotation encoding="application/x-tex">N</mml:annotation> </mml:semantics> </mml:math> </inline-formula> tailor-constructed lattice training points can achieve the generalization error (or <inline-formula content-type="math/mathml"> <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" alttext="upper L 2"> <mml:semantics> <mml:msub> <mml:mi>L</mml:mi> <mml:mrow class="MJX-TeXAtom-ORD"> <mml:mn>2</mml:mn> </mml:mrow> </mml:msub> <mml:annotation encoding="application/x-tex">L_{2}</mml:annotation> </mml:semantics> </mml:math> </inline-formula> approximation error) bound <inline-formula content-type="math/mathml"> <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" alttext="monospace t monospace o monospace l plus script upper O left-parenthesis upper N Superscript negative r slash 2 Baseline right-parenthesis"> <mml:semantics> <mml:mrow> <mml:mrow class="MJX-TeXAtom-ORD"> <mml:mi mathvariant="monospace">t</mml:mi> <mml:mi mathvariant="monospace">o</mml:mi> <mml:mi mathvariant="monospace">l</mml:mi> </mml:mrow> <mml:mo>+</mml:mo> <mml:mrow class="MJX-TeXAtom-ORD"> <mml:mi class="MJX-tex-caligraphic" mathvariant="script">O</mml:mi> </mml:mrow> <mml:mo stretchy="false">(</mml:mo> <mml:msup> <mml:mi>N</mml:mi> <mml:mrow class="MJX-TeXAtom-ORD"> <mml:mo> − </mml:mo> <mml:mi>r</mml:mi> <mml:mrow class="MJX-TeXAtom-ORD"> <mml:mo>/</mml:mo> </mml:mrow> <mml:mn>2</mml:mn> </mml:mrow> </mml:msup> <mml:mo stretchy="false">)</mml:mo> </mml:mrow> <mml:annotation encoding="application/x-tex">\mathtt {tol} + \mathcal {O}(N^{-r/2})</mml:annotation> </mml:semantics> </mml:math> </inline-formula> , where <inline-formula content-type="math/mathml"> <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" alttext="monospace t monospace o monospace l element-of left-parenthesis 0 comma 1 right-parenthesis"> <mml:semantics> <mml:mrow> <mml:mrow class="MJX-TeXAtom-ORD"> <mml:mi mathvariant="monospace">t</mml:mi> <mml:mi mathvariant="monospace">o</mml:mi> <mml:mi mathvariant="monospace">l</mml:mi> </mml:mrow> <mml:mo> ∈ </mml:mo> <mml:mo stretchy="false">(</mml:mo> <mml:mn>0</mml:mn> <mml:mo>,</mml:mo> <mml:mn>1</mml:mn> <mml:mo stretchy="false">)</mml:mo> </mml:mrow> <mml:annotation encoding="application/x-tex">\mathtt {tol}\in (0,1)</mml:annotation> </mml:semantics> </mml:math> </inline-formula> is the tolerance achieved by the training error in practice, and <inline-formula content-type="math/mathml"> <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" alttext="r equals 1 slash p Superscript asterisk"> <mml:semantics> <mml:mrow> <mml:mi>r</mml:mi> <mml:mo>=</mml:mo> <mml:mn>1</mml:mn> <mml:mrow class="MJX-TeXAtom-ORD"> <mml:mo>/</mml:mo> </mml:mrow> <mml:msup> <mml:mi>p</mml:mi> <mml:mrow class="MJX-TeXAtom-ORD"> <mml:mo> ∗ </mml:mo> </mml:mrow> </mml:msup> </mml:mrow> <mml:annotation encoding="application/x-tex">r = 1/p^{*}</mml:annotation> </mml:semantics> </mml:math> </inline-formula> , with <inline-formula content-type="math/mathml"> <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" alttext="p Superscript asterisk"> <mml:semantics> <mml:msup> <mml:mi>p</mml:mi> <mml:mrow class="MJX-TeXAtom-ORD"> <mml:mo> ∗ </mml:mo> </mml:mrow> </mml:msup> <mml:annotation encoding="application/x-tex">p^{*}</mml:annotation> </mml:semantics> </mml:math> </inline-formula> being the <italic>summability exponent</italic> of a sequence that characterises the decay of the input variables in the target functions, and with the implied constant independent of the dimensionality of the input data. We apply our analysis to popular models of parametric elliptic partial differential equations (PDEs) in uncertainty quantification. In our numerical experiments, we restrict the network parameters during training by adding tailored regularization terms, and we show that for an algebraic equation mimicking the parametric PDE problems the DNNs trained with tailored regularization perform significantly better. </p>

Paper

Similar papers

© 2026 NYSGPT2525 LLC