Hybrid Cross Modality Feature Fusion
Liaomo Zheng, Zhipeng Zhang, Shiyu Wang, Songjie Zhou · 2024
This paper presents a novel RGB-D saliency object detection model, which utilizes the MAE model for RGB feature extraction and the ViT model for depth feature extraction. By employing our Hybrid Cross Modality Feature Fusion module, the model effectively integrates RGB and depth features. Experimental results demonstrate that the proposed module significantly enhances detection performance across various public datasets, especially on datasets with lower-quality depth images. Additionally, the superiority of the proposed model is further validated by comparing it with multiple feature fusion schemes.