Learned Video Compression with Spatial Correlation Priors and Hierarchical Temporal Attention

Qian Huang, Wenchao Shan, Qian Xu, Zaipeng Xie, Yiming Wang · 2025

Accurately predicting the probability distribution of quantized latent representations is a critical challenge for entropy models in learned video compression (LVC). Existing mainstream LVC methods typically adopt ready-made entropy models based on image compression, which fail to fully exploit the information of spatial-temporal correlation. To address this issue, we propose a spatial correlation priors and hierarchical temporal attention (SCP-HTA) model, which exploits the spatial correlation information from the current video frames and refine the temporal information from the context. First, we extract the spatial correlation of the current frame to guide the generation of masks, enabling the frame to leverage more information during encoding and decoding process. Additionally, to obtain more accurate temporal information, we introduce a hierarchical temporal attention module at channel level when we generate the context. Experimental results demonstrate that the proposed SCP-HTA model achieve 15.76% bitrate saving in PSNR and 61.82% in MS-SSIM on average across all test datasets when compared with VTM-13.2 (LDP).

Read the paper · More papers on PaperTik