Multi-level Speaker Authentication: An Overview and Implementation

V Amrutha, K H Anagha, Amshu Kamal K, Balachandra Kumaraswamy · 2020

User authentication using biometrics is done based on the physical and behavioral characteristics of humans. A wide range of systems needs accurate personal recognition schemes to either identify or verify an individual's identity while requesting their services. The work presents a model that extracts speech and facial features of individuals to build an authentication system consisting of speaker verification, facial recognition and digit recognition. Frame level Mel Frequency Cepstral Coefficients (MFCC) and delta MFCC features are extracted for speech based authentication. The Speaker verification system is built around a Gaussian Mixture Model (GMM) classifier with likelihood ratio tests for verification. OpenCV's Haar feature-based Cascade Classifiers have been used to extract the face area as facial features. Facial recognition is modeled with a Deep Convolutional Network. Digit recognition is centered around the spectrograms which are the image representations for the spoken digits. Convolution Neural Networks are used to classify the digits by capturing the similarities between the spectrograms of each of the digits. The multilevel model with three stages of authentication systems represents a secure solution for access control problems.

Read the paper · More papers on PaperTik