Depth-aware Trimodal Network for RGB-T Salient Object Dection

Zisen Zhang, Wei Li · 2022

RGB-T salient object detection (SOD) aims to locate and segment the most attractive object in color-thermal image pairs. Depth information has been proved beneficial in RGB-D salient object detection during recent years. To take advantage of the structural information contained in the depth map, we build a depth-aware trimodal network (DATNet) to introduce depth modal into RGB-T SOD. First, we generate depth maps from RGB frames through existing monocular depth estimation methods, depth stream constitutes the third input to our DATNet. We use three independent transformer backbone net to extract hierarchical features. Our multimodal interaction module (MIM) is designed to achieve three modal feature enhancement and complementation. Notably, we treat the three modals unequally. Multimodal enhanced fusion module (MEFM) aims to refine the features after MIM and fuse them in stages. Finally, we use multiple supervisions in progressive decoding to obtain high-quality saliency maps. Comprehensive experiments on three public RGB-T SOD datasets show that the proposed network surpasses 8 state-of-the-art RGB-T SOD methods in terms of five metrics.

Read the paper · More papers on PaperTik