RGB-D Face Recognition via Spatial and Channel Attentions

Libiao Jiang, Junwei Zhang, Changyu Li, Jianwei Zhou · 2021

After decades of extensive research, the field of 2D face recognition is fruitful. However, 2D face recognition is sensitive to changes in pose, facial expression and illumination. 3D face recognition has great potential due to the inherent invariance of pose and illumination changes, but these 3D face images are still difficult to acquire due to several issues such as cost and accessibility. In contrast, RGB-D images captured by low-cost depth sensors such as Kinect are relatively easy to acquire, and research on such low-quality inputs is still limited. In this thesis, an end-to-end multimodal fusion method based on spatial and channel attention is proposed to effectively fuse two image modalities, RGB and depth, to enhance RGB-D face recognition. Our work is specified as follows: firstly, three branches combined with the spatial and channel attention modules are established on the basis of ResNet18, so that the features of three modalities (RGB, depth map, and their fusion modalities, respectively) are obtained. The features under these three modalities are then fused, fed into a shared layer, and the fused features are further learned for deeper discriminative features and passed through a spatial attention vectorization module at the end. The method achieves good results on the Intellifusion RGBD dataset, while we perform ablation trials, which well reveal the contribution of each component of our method to the final performance.

Read the paper · More papers on PaperTik