Efficient Audio-Visual Speaker Recognition Via Deep Multi-Modal Feature Fusion

Yufei Wang · 2021 17th International Conference on Computational Intelligence and Security (CIS) · 2021

Audio-visual speaker recognition has received an increasing attention in recent years due to the growing security demands. In this paper, we present an efficient deep multi-modal feature fusion approach to fuse the face and audio data for reliable speaker recognition. Within our proposed approach, we first utilize the convolutional neural networks (CNN) to extract the inherent facial features and employ a group of Mel-Frequency Cepstral Coefficients (MFCC) to characterize each audio sample. Accordingly, the facial features learned by CNN are compatible with audio features at high-level appropriately. Subsequently, we propose several deep multi-modal learning schemes, i.e., deep feature-level fusion and deep decision-level fusion, to fuse the face and audio modalities at different level such that the speaker identity can be well identified. The experimental results have shown that our proposed audio-visual speaker recognition approach can produce better performance than single modality, and the deep feature-level fusion yields comparative and even better results than the state-of-the-art counterparts.

Read the paper · More papers on PaperTik