Research on Salient Object Detection Algorithm Based on Image Depth Information

Peijie Zeng, Wei Zhou · 2024

This paper proposes a novel RGBD segmentation network structure named AsymFusion-FPDNet, aiming to improve the accuracy of RGBD segmentation by fully utilizing the rich semantic information of RGB images. Although significant progress has been made in existing RGBD segmentation methods, most of them primarily focus on the utilization of depth image information and the fusion of RGB and depth information, while overlooking the semantic information in RGB images. Therefore, based on the asymmetric multimodal structural network AsymFormer, we introduce dilated convolution blocks after the final downsampling operation of the RGB image feature extraction network. By gradually expanding the receptive field through dilated convolutions with different dilation factors, we extract more extensive feature information to enhance the representation capability of the network. Additionally, we incorporate a feature pyramid module and a CSPLayer_2Conv structure into the RGB image feature extraction network structure to improve the model's feature extraction ability in complex scenes, particularly by fusing spatial details and semantic information to enhance the model's comprehensive understanding and capture of targets. Experimental results show that AsymFusion-FPDNet achieves excellent performance on the NYUv2 dataset, with a MIoU of 56.91% and a PA of 79.78%, representing an improvement of 2.81% in MIoU and 1.29% in PA compared to the advanced AsymFormer model. This validates the effectiveness and advancement of the proposed model for RGBD segmentation tasks.

Read the paper · More papers on PaperTik