Per title video quality encoding with CRF estimation based on scenes using DNN
Francisco Micó-Enguídanos, J. Aguado, Miguel García · 2024
In the digital era, video content dominates network traffic, with HTTP Adaptive Streaming commonly used by streaming services. Recently, novel video encoding schemes based on Deep Neural Networks (DNN) have been proposed. One of the proposals used features extracted from fixed length downscaled video segments. The DNN used as input the features and the desired Video Multimethod Assessment Fusion (VMAF), to estimate adaptively the Constant Rate Factor (CRF) to be applied to the original segment to attain the target VMAF. In this work we analyze the generalization capabilities of this trained DNN when instead of fixed length segments we consider scenes, that have variable duration. We show that the CRF is adapted per scene and per video and the final encoded videos have the requested VMAF. According to our tests, the encoded videos have an average absolute error of 1.34, being less than 2 and therefore not perceived by the end user. Besides, the size of the encoded videos decreases, up to 12% in the best case, compared to those obtained estimating the CRF from fixed length segments.