Inductive Bias of Multi-Channel Linear Convolutional Networks with Bounded Weight Norm

We provide a function space characterization of the inductive bias resulting\nfrom minimizing the $\\ell_2$ norm of the weights in multi-channel convolutional\nneural networks with linear activations and empirically test our resulting\nhypothesis on ReLU networks trained using gradient descent. We define an\ninduced regularizer in the function space as the minimum $\\ell_2$ norm of\nweights of a network required to realize a function. For two layer linear\nconvolutional networks with $C$ output channels and kernel size $K$, we show\nthe following: (a) If the inputs to the network are single channeled, the\ninduced regularizer for any $K$ is independent of the number of output channels\n$C$. Furthermore, we derive the regularizer is a norm given by a semidefinite\nprogram (SDP). (b) In contrast, for multi-channel inputs, multiple output\nchannels can be necessary to merely realize all matrix-valued linear functions\nand thus the inductive bias does depend on $C$. However, for sufficiently large\n$C$, the induced regularizer is again given by an SDP that is independent of\n$C$. In particular, the induced regularizer for $K=1$ and $K=D$ (input\ndimension) is given in closed form as the nuclear norm and the $\\ell_{2,1}$\ngroup-sparse norm, respectively, of the Fourier coefficients of the linear\npredictor. We investigate the broader applicability of our theoretical results\nto implicit regularization from gradient descent on linear and ReLU networks\nthrough experiments on MNIST and CIFAR-10 datasets.\n

Paper

References (39)

Scroll for more · 27 remaining

Similar papers

© 2026 NYSGPT2525 LLC