Improved Visual Saliency Prediction Based on Video Swin Transformers
ChaeEun Woo, SuMin Lee, Soo-Min Park, Serin Choi, Jekyung Ryu, Byung-Hyung Kim · Journal of Korea Multimedia Society · 2024
In this paper, we propose a Video Swin Transformer Saliency Network (VST-SalNet). The proposed model utilizes the Video Swin Transformer as its backbone to effectively learn the spatiotemporal fea- tures of video data and is designed to handle long-range spatiotemporal dependencies. Additionally, it integrates high-level semantic information and low-level details through the application of a feature pyramid structure. This structure enables multi-scale feature fusion and refines spatial details across resolutions. In turn, the model enhances spatial resolution by effectively handling objects of various sizes, preserving semantic information, and minimizing information loss. Experimental results on DHF1K, Hollywood-2, and UCF Sports datasets, evaluated using metrics such as SIM and CC, confirm that VST-SalNet outperforms the state-of-the-art models.