Saliency Dataset and Predictive Model for Areas of Interest in VVC Perceptual Coding

Jorge Kessler-Martín, Pablo Fernández-Lagos, David García-Lucas, Gabriel Cebrián-Márquez, Belén Ríos, Guillermo Vigueras, Antonio Jesús Díaz-Honrubia · 2024

Video coding standardization organizations have invested significant efforts in achieving greater compression factors over the years. Approved in 2020, the Versatile Video Coding (VVC) standard reduces the bit rate needed to encode a sequence by half compared to its predecessor. However, users today have increasingly demanding requirements, leading to a significant rise in video traffic on the Internet. In this context, perceptual video coding aims to reduce video bit rate by decreasing the objective quality while maintaining the subjective quality. This work presents a novel dataset designed for training models to predict video saliency, i.e., areas in the video to which viewers are more likely to pay attention. The dataset is publicly available. Furthermore, this work also proposes a machine learning model that classifies each Coding Tree Unit (CTU) as salient or not, and adjusts its quality accordingly. The results show that this model has an accuracy of 95% and correctly classifies as salient 98% of the CTUs that are actually salient.

Read the paper · More papers on PaperTik