We present a novel Deep Neural Network (DNN) architecture for non-linear\nsystem identification. We foster generalization by constraining DNN\nrepresentational power. To do so, inspired by fading memory systems, we\nintroduce inductive bias (on the architecture) and regularization (on the loss\nfunction). This architecture allows for automatic complexity selection based\nsolely on available data, in this way the number of hyper-parameters that must\nbe chosen by the user is reduced. Exploiting the highly parallelizable DNN\nframework (based on Stochastic optimization methods) we successfully apply our\nmethod to large scale datasets.\n