Multimodal Person Identification through the Fusion of Face and Voice Biometrics
Marwa Saleh, Ismail I. Jouny · 2022
Person identification is used in a variety of everyday applications from access control and security cameras to social media and face ID on smartphones. There are still, however, various limitations to this technology that do not make it completely reliable. This paper focuses on face and voice biometrics in particular and takes advantage of their fusion to produce a more accurate final decision. The Michigan State University (MSU-AVIS) dataset is used which contains 50 subjects whose faces are captured from different angles and whose voices are recorded as they read from a random script. For the face dataset, videos are first split into frames which are passed into the Viola-Jones algorithm for face detection. A convolutional neural network (CNN) is trained to identify the subjects based on their facial features. For the voice dataset, each audio file is split into frames and any silence or unvoiced speech is removed. Two features are extracted: pitch and Mel frequency cepstrum coefficients which are then passed into another CNN. Finally, outputs of both CNNs are fused at the decision level by choosing the output of higher confidence in each case.