Multi-dimensional Attention Feature Alignment for High Resolution Saliency Detection
Biao Tong, Xiaoning Song · 2023
In the context of saliency object detection, semantic information within images plays a crucial role. However, convolutional structures are constrained by their receptive fields, limiting their ability to model long-range dependencies effectively. On the other hand, Vision Transformer (ViT) networks excel at capturing global contextual information within images but come with the drawback of significant computational demands, particularly when handling high-resolution photos. This paper introduces a dual-branch multi-dimensional attention feature pyramid fusion network to address this contradiction. This network leverages a convolutional encoder to extract fine-grained details from high-resolution images and a ViT encoder to extract semantic information from corresponding low-resolution images. Additionally, it employs pre-trained weights generated using the MAE self-supervision method to enhance ViT structure’s understanding of salient objects and feature extraction capabilities. Within the decoding structure, a multi-dimensional attention fusion enhancement module is designed to weigh and integrate features across multiple dimensions simultaneously. This module effectively suppresses noise and reinforces prominent features. Furthermore, a pyramid feature refinement module is introduced to fine-tune features during the multi-scale fusion process, focusing on enhancing edge details. Experimental results on datasets, including HRSOD, DAVIS, UHRSD, DUT-OMRON, and DUTS, demonstrate the network’s superior performance.