Multi-Domain Spatio-Temporal Deformable Fusion model for video quality enhancement
Garibaldi da Silveira Júnior, Gilberto Kreisler, Bruno Zatt, Daniel Palomino, Guilherme Correa · 2024
Lossy video compression introduces artifacts that can degrade the perceived visual quality of the video. Improving the quality of compressed videos involves mitigating these artifacts through filtering techniques. Deep neural network (DNN) models have emerged as powerful tools for this task, demonstrating effectiveness in artifact reduction. However, traditional approaches typically evaluate these models using videos compressed by a single coding standard, limiting their applicability across diverse codecs. To address this limitation, this study proposes a novel multi-domain architecture built upon the Spatio-Temporal Deformable Fusion technique. This innovative approach enables the development of models capable of enhancing videos compressed by various codecs, ensuring consistent performance across different standards. Experimental results showcase the efficacy of the proposed method, yielding significant improvements in average Peak Signal-to-Noise Ratio (PSNR) for videos compressed with HEVC, VVC, VP9, and AV1, with enhancements of 0.764 dB, 0.448 dB, 0.736 dB, and 0.228 dB, respectively. The code of our MD-STDF approach is available at https://github.com/Espeto/md-stdf