Weakly Supervised Salient Object Detection with Dual-path Mutual Reinforcement Network
Yanling Ou, Yanjiao Shi · 2024
Traditional salient object detection (SOD) methods heavily rely on large-scale pixel-level datasets, making them both time-consuming and expensive. However, it is a significant challenge to effectively integrate long-range dependencies and local feature information for weakly supervised salient object detection (WSOD). In this paper, we propose a novel dual-path mutual reinforcement network (DMRN) for WSOD, which includes a CNN branch and a ViT branch in the encoder to effectively capture local details and global features, respectively. A mutual fusion module (MFM) is introduced to perform efficient feature fusion and establish long-range contextual dependencies. Additionally, an attention-based cross-feature module (ACFM) is designed to enhance feature interactions between branches, ensuring accurate localization of salient objects. To further enhance the model, two auxiliary prediction modules (APM) are used to generate pseudo labels from different perspectives, compensating for sparse annotations. Experimental results on five benchmark datasets show that the proposed method outperforms 11 existing weakly supervised methods and shows its superior performance in capturing salient features and achieving accurate object detection.