Light Field Saliency Detection Based on Multi-modal Fusion
Ben Jiang, Dan Xu, Jinlong Shi · 2022
Compared with RGB images, light field images contain more abundant visual information, which is helpful to accurately detect salient objects in complex scenes. However, most of the existing light field saliency detection methods use single light field data or do not fully consider the differences and complementarities between different light field data, resulting in insufficient multi-modal fusion. To address these issues, a multi-modal feature fusion network is proposed, which makes full use of the rich visual information in the light field images to realize the accurate saliency object detection. The proposed network consists of two parallel subnets, which are used to process the micro-lens image array and all-foucs image respectively. Then the light field refinement module is used to refine the feature map extracted from the micro-lens array stream, and finally the multi-modal feature fusion is realized by the light field attention module to predict saliency objects more accurately. In order to verify the effectiveness of proposed method, extensive comparison with several existing light field saliency detection algorithms is carried on both Lytro-Illum and LFSD datasets. Experimental results show that the proposed method is superior to others in all evaluation metrics on Lytro-Illum dataset, and has desired generalization abilities on LFSD dataset.