GOP-Based Deep Preprocessing for Video Coding
Daichi Arai, Shunsuke Iwamura, Kazuhisa Iguchi, Atsuro Ichigaya · 2024
Neural network-based video preprocessing techniques have recently shown remarkable improvements in video codec performance. However, conventional preprocessing methods tend to prioritize perceptual quality over peak signal-to-noise ratio (PSNR), a key standard for video quality assessment. In this study, We propose a novel deep preprocessing method based on a group of pictures (GOP) structure, specifically aimed at enhancing the rate-distortion performance in terms of PSNR. This approach involves developing a video compression model that employs the GOP structure of the target video codec and training a preprocessing model through joint optimization with the video compression model. Experimental results demonstrate that our GOP-based deep preprocessing method not only improves PSNR but also elevates other quality metrics, including VMAF, across various codecs like MPEG-2, HEVC, and VVC. Additionally, ablation studies highlight the critical role of GOP structures in enhancing encoding efficiency based on PSNR.