On the Impact of Lossy Image and Video Compression on the Performance of Deep Convolutional Neural Network Architectures
Recent advances in generalized image understanding have seen a surge in the\nuse of deep convolutional neural networks (CNN) across a broad range of\nimage-based detection, classification and prediction tasks. Whilst the reported\nperformance of these approaches is impressive, this study investigates the\nhitherto unapproached question of the impact of commonplace image and video\ncompression techniques on the performance of such deep learning architectures.\nFocusing on the JPEG and H.264 (MPEG-4 AVC) as a representative proxy for\ncontemporary lossy image/video compression techniques that are in common use\nwithin network-connected image/video devices and infrastructure, we examine the\nimpact on performance across five discrete tasks: human pose estimation,\nsemantic segmentation, object detection, action recognition, and monocular\ndepth estimation. As such, within this study we include a variety of network\narchitectures and domains spanning end-to-end convolution, encoder-decoder,\nregion-based CNN (R-CNN), dual-stream, and generative adversarial networks\n(GAN). Our results show a non-linear and non-uniform relationship between\nnetwork performance and the level of lossy compression applied. Notably,\nperformance decreases significantly below a JPEG quality (quantization) level\nof 15% and a H.264 Constant Rate Factor (CRF) of 40. However, retraining said\narchitectures on pre-compressed imagery conversely recovers network performance\nby up to 78.4% in some cases. Furthermore, there is a correlation between\narchitectures employing an encoder-decoder pipeline and those that demonstrate\nresilience to lossy image compression. The characteristics of the relationship\nbetween input compression to output task performance can be used to inform\ndesign decisions within future image/video devices and infrastructure.\n