Location, Neighborhood, and Semantic Guidance Network for RGB-D Co-Salient Object Detection
Wujie Zhou, Bingying Wang, Xiena Dong, Caie Xu, Fangfang Qiang · IEEE Transactions on Artificial Intelligence · 2025
Red–green–blue-depth (RGB-D) deep learning-based co-salient object detection (Co-SOD) automatically detects and segments common salient objects in images. However, this computationally intensive model cannot be run on mobile devices. To help overcome this limitation, this article proposes a localization, neighborhood, and semantic guidance network (LNSNet) with knowledge distillation (KD), called LNSNet-S*, for RGB-D Co-SOD to minimize the number of parameters and improve the accuracy. Apart from their backbone networks, the LNSNet student (LNSNet-S) and teacher (LNSNet-T) models use the same structure to capture similarity knowledge in category, channel, and pixel-point dimensions to train an LNSNet-S with KD for superior lightweight performance. For optimization, a positioning path progressive activation uses hierarchical transformers to fuse features from low to high levels, generating class activation localization maps using the fused bimodal information to obtain location information. The high-level neighborhood-guidance information is then used to guide the low-level features. Next, a multisource semantic enhancement embedding module progressively fuses multiscale cross-modal semantic information guided by class-activated localization information. A class-based progressive triplet loss facilitates the transfer of category, channel, and pixel-point information. Extensive experiments demonstrated the effectiveness and robustness of the novel LNSNet-S* in different sizes, and significant improvements were observed. The smallest LNSNet-S* model reduced the number of parameters by more than 92% compared to that of LNSNet-T, requiring only 15.9 M parameters. The code has been publicly released at https://github.com/Wang-5ying/LNSNet.