Frame Differences Matter in Quality Assessment of Compressed Videos
Xinyi Wang, Angeliki Katsenou, David Bull · 2025
With the rapid growth of User-Generated video Content (UGC) exchanged between users and sharing platforms, the need for video quality assessment in the wild is increasingly evident. UGC is typically acquired using consumer devices and undergoes multiple rounds of compression (transcoding) before reaching the end user. Therefore, traditional quality metrics that employ the original content as a reference are not suitable. In this paper, we propose a No-Reference Video Quality Assessment (NR-VQA) model, ReLaX-VQA, that uses frame differences to select “key” spatio-temporal fragments along with various spatial features of the sampled frames. These are then used to better capture spatiotemporal variability in the quality of neighbouring frames. Furthermore, the model benefits from abstraction achieved by employing layer-stacking techniques in deep features from Residual Networks and Vision Transformers. Extensive testing demonstrates that ReLaX-VQA outperforms existing NR-VQA methods (that are based solely on video data) over most datasets, achieving an average SRCC of 0.8658 and PLCC of 0.8873. Code, trained models, and supplementary data are available: https://github.com/xinyiW915/ReLaX-VQA.