Video deblurring using transformer-based deep neural networks

In this study, a video deblurring approach using transformer-based deep neural networks is proposed, which consists of feature extraction module, feature alignment module, temporal fusion transformer, and reconstruction transformer. Feature extraction module is employed to extract frequency and spatial domain features of video frames. Feature alignment module is employed to bi-directionally propagate feature maps and utilize optical flow to guide the alignment of feature maps. Temporal fusion transformer is employed to build correspondences across neighboring frames and fully utilize sharp details in neighboring frames. The fused feature map is fed into reconstruction transformer to gradually reconstruct fine-grained details and generate the residual frame. Finally, each deblurred frame is obtained by adding the residual frame to the corresponding input frame. Based on the experimental results obtained in this study, in terms of two objective performance metrics and subjective evaluation, the performance of the proposed approach is better than those of four comparison approaches.

Paper

Full text

PDF

Video deblurring using transformer-based deep neural networks

Semantic Scholar · Computer Science · 2026

Abstract

In this study, a video deblurring approach using transformer-based deep neural networks is proposed, which consists of feature extraction module, feature alignment module, temporal fusion transformer, and reconstruction transformer. Feature extraction module is employed to extract frequency and spatial domain features of video frames. Feature alignment module is employed to bi-directionally propagate feature maps and utilize optical flow to guide the alignment of feature maps. Temporal fusion transformer is employed to build correspondences across neighboring frames and fully utilize sharp details in neighboring frames. The fused feature map is fed into reconstruction transformer to gradually reconstruct fine-grained details and generate the residual frame. Finally, each deblurred frame is obtained by adding the residual frame to the corresponding input frame. Based on the experimental results obtained in this study, in terms of two objective performance metrics and subjective evaluation, the performance of the proposed approach is better than those of four comparison approaches.

Similar papers

© 2026 NYSGPT2525 LLC