Study on Deep Learning-Based Multimodal Fusion for Significant Target Detection
Rui Wang, Rong Dai, Yuanqing He · 2024
With the rapid development of computer vision and deep learning technologies, Salient Object Detection (SOD) has been widely applied in fields such as image understanding, video analysis, and robot vision. Traditional target detection methods mainly rely on single mode data. However, due to the diversity of image and video data, single mode has limitations in dealing with complex scenes. This article summarizes the core research achievements in this field over the past five years, and provides a detailed review of multi-modal saliency object detection tasks based on RGB-D (depth image data), RGB-T (thermal imaging data), and VDT (visual depth thermal imaging). The existing research is systematically organized, summarized, and evaluated. Finally, the article also conducted in-depth discussions and prospects on the current challenges and future development trends in this field.