TM2SP: A Transformer-Based Multi-Level Spatiotemporal Feature Pyramid Network for Video Saliency Prediction

Chenming Li, Shiguang Liu · IEEE Transactions on Circuits and Systems for Video Technology · 2025

This paper proposes an end-to-end video saliency prediction network model, termed TM2SP-Net (Transformer-based Multi-level Spatiotemporal Feature Pyramid Network). Leveraging the strong encoding learning capability of Video Swin Transformer for video data, we design a Multi-level Spatiotemporal Feature Pyramid Network (MLSTFPN) that effectively detects and enriches salient regions and spatial details across different scales. In particular, a pre-trained image saliency detection encoder is employed to extract salient features from each frame, serving as prior knowledge to guide the multi-scale spatiotemporal feature fusion and decoding processes. Additionally, we introduce an Inception Gate-Controlled Fusion (IGCF) and Layered Self-Attention Aggregation Fusion (LSAF) mechanisms to efficiently merge spatiotemporal features across various stages. Finally, extensive experiments conducted on the DHF1K, Hollywood-2, UCF-Sports, and six audio-visual saliency datasets demonstrate the superiority of our method over existing state-of-the-art approaches.

Read the paper · More papers on PaperTik