Cross-Frame Integrated Prediction for Feature-Space Video Compression

Hongxin Qiu, Zhidao Zhou, Kai Lu, Fan Liang · 2024

Learned video compression in the feature domain employs implicit motion compensation to acquire predicted features or compute feature residuals, effectively minimizing spatiotemporal redundancy in the reconstructed frames. In this paper, we propose a cross-frame integrated prediction (CIP) network for feature-space video compression. Leveraging multiple features in motion estimation and compensation, our approach enables more context-aware prediction. Specifically, on the one hand, we introduce a global feature extraction (GFE) module in motion estimation to extract the information of the current feature and multiple reference features, providing a high-quality offset map for deformable motion compensation. On the other hand, we utilize an attentional feature fusion (AFF) module for multiple predicted features in motion compensation, which is beneficial for preserving crucial details and adapting to diverse scenes and content. By passing the final predicted feature to the residual compression and frame reconstruction, we achieve a single end-to-end video compression framework avoiding laborious multi-stage training. Comprehensive experimental results show that the proposed method not only maintains a low number of model parameters but also achieves significant performance improvement in video compression tasks, especially in the case of high resolution and high bitrate.

Read the paper · More papers on PaperTik