End-to-End Deep Video Compression Based on Hierarchical Temporal Context Learning
Kejun Wu, Zhenxing Li, You Yang, Qiong Liu, Xiao-Ping Zhang · IEEE Transactions on Multimedia · 2025
Emerging learning-based video compression suffers from error propagation in long group of pictures (GOP), yielding limited coding performance. To address this problem, a novel end-to-end Deep Video Compression method based on Hierarchical Temporal Context Learning (DVCH) is proposed in this paper. DVCH aims to fully exploit temporal contexts and suppress error propagation for better coding performance. It first divides video frames into several hierarchies with different compression qualities. The frames in lower hierarchies have high compression quality, and serve as reference frames. To mine high-quality reference information, we propose a Hierarchical Temporal Context Learning (HTCL) network as the fundamental module of our DVCH. The informative temporal context features from hierarchical prediction structure can be extracted by the network. Motion vectors (MVs) between the to-be-coded frame and its reference frames are estimated by the MV Learning module and used to align the extracted contexts. The contexts are fed into Context Coding module to generate the prediction of the decoded frame. Moreover, a multi-stage training strategy is developed to solve the imbalanced training challenge. Experimental results demonstrate that the proposed DVCH exceeds x264 and other end-to-end video compression methods, regardless of objective, subjective, error propagation suppression, GOP sizes, and sequence length evaluations. As much as 49.27% bitrate savings and 2.52 dB PSNR gains can be achieved in large GOP.