Watch your Up-Convolution: CNN Based Generative Deep Neural Networks are Failing to Reproduce Spectral Distributions
Generative convolutional deep neural networks, e.g. popular GAN\narchitectures, are relying on convolution based up-sampling methods to produce\nnon-scalar outputs like images or video sequences. In this paper, we show that\ncommon up-sampling methods, i.e. known as up-convolution or transposed\nconvolution, are causing the inability of such models to reproduce spectral\ndistributions of natural training data correctly. This effect is independent of\nthe underlying architecture and we show that it can be used to easily detect\ngenerated data like deepfakes with up to 100% accuracy on public benchmarks.\n To overcome this drawback of current generative models, we propose to add a\nnovel spectral regularization term to the training optimization objective. We\nshow that this approach not only allows to train spectral consistent GANs that\nare avoiding high frequency errors. Also, we show that a correct approximation\nof the frequency spectrum has positive effects on the training stability and\noutput quality of generative networks.\n