CNN Filter for Super-Resolution with RPR Functionality in VVC

High resolution (HR) videos are easy to go beyond the band-width limit during transmission after encoding. Downsampling followed by upsampling is a well-known strategy for compressing HR videos with limited bandwidth. Versatile video coding (VVC) provides a reference picture resampling (RPR) functionality to consider the limited bandwidth. In this paper, we propose a convolutional neural network (CNN) filter for super-resolution (SR) with the RPR functionality in VVC. We design a lightweight SR network that combines CNN with the RPR functionality in VVC, called lightweight network of multi-level mixed scale and depth information with attention mechanism (LMSDANet). For LMSDANet, we provide a lightweight block of multi-mixed scale and depth information with attention (LMSDAB) to extract multi-scale and convolutional layer depth information while enhancing the representation of features. Compared with VTM-11.0_NNVC-2.0 anchor, LMSDANet achieves {-9.16% (Y), 17.03% (U), -7.61% (V)} and {-4.14% (Y), 6.34% (U), -2.25% (V)} BD-rate changes (average on A1 and A2) in AI and RA configurations, respectively.

Paper

Full text

PDF

CNN Filter for Super-Resolution with RPR Functionality in VVC

Semantic Scholar · Computer Science · 2023

Abstract

High resolution (HR) videos are easy to go beyond the band-width limit during transmission after encoding. Downsampling followed by upsampling is a well-known strategy for compressing HR videos with limited bandwidth. Versatile video coding (VVC) provides a reference picture resampling (RPR) functionality to consider the limited bandwidth. In this paper, we propose a convolutional neural network (CNN) filter for super-resolution (SR) with the RPR functionality in VVC. We design a lightweight SR network that combines CNN with the RPR functionality in VVC, called lightweight network of multi-level mixed scale and depth information with attention mechanism (LMSDANet). For LMSDANet, we provide a lightweight block of multi-mixed scale and depth information with attention (LMSDAB) to extract multi-scale and convolutional layer depth information while enhancing the representation of features. Compared with VTM-11.0_NNVC-2.0 anchor, LMSDANet achieves {-9.16% (Y), 17.03% (U), -7.61% (V)} and {-4.14% (Y), 6.34% (U), -2.25% (V)} BD-rate changes (average on A1 and A2) in AI and RA configurations, respectively.

Similar papers

© 2026 NYSGPT2525 LLC