Transformer-CNNs Network with Reverse Convolution Feature Step-by-Step Weighted Dense MLP Group and Binary Mask Guidance Feature Fusion of Salient Map

Ziyuan Luo, Qiang Dai, Haiyu Liao, Xiaohui Luo · 2024

Salient object detection (SOD) is crucial in computer vision applications. Many advanced RGB-D SOD models employ direct operations for feature fusion and multi-stage/multi-scale decoding for saliency maps. However, problems like feature differences between RGB and depth images at various scales lead to issues such as inaccurate target localization, information loss, and redundant information, making mode identification more difficult. To solve these, ReCoSF-Net, which combines transformers and CNNs, is proposed. It consists of four modules: GIE-MLPs (using interconnected MLPs with dense connections and reverse weighting for pattern learning), SFGD (using the prior decoder's saliency map to guide bimodal feature fusion), OFER (combining the last saliency map with original features for edge refinement), and LSAM (identifying salient parts in long-sequence features after SFGD's fusion).

Read the paper · More papers on PaperTik