3D convolution-based video compression model without motion estimation
Zhichen Liu, Jian-ping Luo · 2023
In recent years, there are more and more researches in the field of video compression based on deep learning. Most of these researches maintain the same framework, including motion estimation and compensation modules. We propose a different method, which does not need an explicit motion estimation module. Instead, video compression is achieved by extracting and compressing the residuals between two adjacent frames. The residual is extracted by a 3D pyramid network. Based on the robust representation ability of deep features in various applications, the whole coding and compression process is transferred from pixel space to feature space to improve compression efficiency. The transformation process is realized by two networks: feature extraction and frame reconstruction. Our video compression framework has a simple structure and a smaller number of parameters than most subsequent proposed models. The experiment result shows that the performance of our PSNR model is better than H.264, and is comparable to H.265 in some cases, while our MS-SSIM model is better than H.265 in all cases.