ScribbleSAM: Weakly supervised salient object detection and localization in remote sensing images using Transformer and Segment Anything Model

Tae Hun Lee, Jae Yeol Lee · Journal of Computational Design and Engineering · 2025

Abstract Optical remote sensing images (RSIs) significantly differ from natural scene images (NSIs) concerning object types, orientations, sizes, and complex backgrounds. Recently, scribble annotations have been introduced as a more efficient solution to reduce the burden of pixel-level labelling for salient object detection (SOD) in RSIs. However, scribble annotations provide less information, leading to lower performance. To address these challenges, this paper proposes a new weakly supervised salient object segmentation model, ScribbleSAM, for RSIs with scribble annotations by combining Segment Anything Model (SAM) and a Transformer model. The proposed model can significantly improve segmentation performance and accuracy compared to previous weakly supervised models. As pre-processing, a simple but effective method is proposed to automatically generate pseudo masks using SAM with selective points of the foreground and background scribbles and key feature points of the foreground scribble. Then, the scribble annotations and generated pseudo-SAM masks are used together to train the dual branches of ScribbleSAM. The scribble annotations are used to train the scribble branch to capture the overall shape of the salient object, and the pseudo-SAM masks are used to train the SAM branch to refine the model's predictions of finer details. Both branches can complement each other. Extensive quantitative and qualitative experiments show that the proposed model outperforms existing weakly supervised models in RSIs and is competitive even with fully supervised models on two RSI SOD datasets. Furthermore, ablation studies verify the effectiveness and generalization of the proposed model.

Read the paper · More papers on PaperTik