Multimodal collaborative saliency object detection network using MCSDNet
Ying Yi, Yutao Hu, Jian Sun, Changping Li · Displays · 2026
Salient object detection (SOD) is a core preprocessing task for image acquisition and process. Simulating the human visual attention mechanism to identify salient objects from complex scenes is crucial for fields such as image segmentation and autonomous driving perception. However, existing methods are limited by global topological associations or are susceptible to rough boundaries and ambient noise. To address these issues, this paper proposes a multimodal collaborative salient detection network (MCSDNet) where a convolutional neural networks (CNN) local contrastive features-based three-stream complementary architecture, CapsNet structure modeling, and boundary-guided salient region were first introduced, for achieving cross-modal feature collaboration. Moreover, a dynamic interaction mechanism is designed to fuse features through cross-modal attention and eliminate noise using a boundary semantic alignment module. A phased optimization strategy was adopted in this work to directly constrain the contour distribution with the use of boundary enhancement loss. Simulation results show that our proposed MCSDNet outperforms existing methods with 13 leading indicators on five datasets (e.g., ECSSD), in particular, the obtained mean M values decreased by 12.9%, 5.4%, 3.6%, 8.3%, and 6.4%, respectively, significantly improving the target integrity and boundary accuracy in complex scenes.