Lightweight Spatial-Frequency Collaborative Interaction Network for RGB-D Salient Object Detection
Yitong Lu, Ziguan Cui · Sensors · 2026
RGB-D salient object detection (SOD) aims to segment the most prominent objects from the background with a pair of given RGB and depth images. Existing RGB-D methods usually rely on heavy backbones to achieve high accuracy, while current lightweight methods struggle to maintain competitive performance. To break this intractable trade-off between effectiveness and model complexity, we propose a Lightweight Spatial-Frequency Collaborative Interaction Network (SFCINet), a unified and highly efficient framework. The core of SFCINet resides in the synergy between spatial-domain features and frequency-domain global priors. Specifically, we introduce the Spatial-Frequency Synergy (SFS) module, which shifts the perspective to a joint complex Fourier domain. By adaptively learning and optimizing the decoupled amplitude and phase components, it effectively isolates clutter to yield a purified global frequency-synergized prior, which modulates the spatial branches to eliminate cross-modal discrepancies for subsequent feature fusion while supplementing global information during decoding. To alleviate the interference caused by cross-modal representation discrepancies, we design the Cross-Guidance Interaction (CMGI) module, which employs a reciprocal anchoring mechanism. It guides the counterpart to mutually filter irrelevant noise and select task-relevant information, achieving fusion in an efficient manner. Finally, we present a Calibrated Hierarchical Decoder (CHD), which injects frequency-synergized global priors into the hierarchical decoding process. It re-establishes the connection between the frequency and spatial domains, ultimately achieving global-local consistency. Extensive experiments demonstrate that SFCINet delivers superior performance over state-of-the-art methods.