FVSR-CT: A CNN- and Transformer-based Model for Better Face Video Super-Resolution Results

Shaohua Jia, Wan-Chi Siu, Pengyu Liu, Kebin Jia · 2025

Research on face video super-resolution has made significant strides at 2x and 4x magnification, but there is comparatively less work on higher magnification tasks. Leveraging the spatial processing capabilities of Convolutional Neural Network (CNN) and the long-range dependency modeling of Transformers, this paper presents Face Video Super-resolution CNN Transformer (FVSR-CT), an effective CNN- and Transformer-based model designed for high-magnification face video super-resolution tasks. However, designing an appropriate CNN- and Transformer-based model for high-magnification face video super-resolution is challenging due to the lack of sufficient spatial information in the input frames, difficulties in inter-frame feature alignment, and the high computational costs associated with high-resolution spatial modeling. To address these challenges, FVSR-CT advocates using Multi-channel Spatial Encoding to extract and enhance information from the input frames, employing Inter-frame Point-wise Masked Attention to establish inter-frame alignment, and implementing a Low-rank Decomposition Reconstruction Method to optimize the parameters of the attention mechanism. Compared to existing advanced models, our proposed method achieves highly competitive results. Such as CNN- and Transformer-based model can serve as a baseline for video super-resolution and other video reconstruction tasks.

Read the paper · More papers on PaperTik