An Overview on Automatic Speaker Verification System Techniques

A Ajila · International Journal for Research in Applied Science and Engineering Technology · 2020

Automatic Speaker Verification (ASV) is a widely used voice-based biometric system. Like every other biometric system, ASV is also vulnerable to malicious attacks. Replay attacks in ASV systems are a greater threat compared to other attacks. Many types of research have been conducted to find the countermeasures for the replay attacks. Feature extraction and classification are the two fundamental processes of speaker verification. A wide variety of features and classifiers were implemented in the past studies to check their performances. The feature extraction methods discussed in this paper are Mel-Frequency Cepstral Coefficients (MFCC), Linear Prediction Coefficients (LPC), Linear Prediction Cepstral Coefficients (LPCC), Discrete Wavelet Transform (DWT) and Perceptual Linear Prediction (PLP). Some of the classifiers that are discussed are Gaussian Mixture Model (GMM), Gaussian Mixture Model-Universal Background Model (GMM-UBM), Hidden Markov Model (HMM), Support Vector Machine (SVM), Dynamic Time Wrapping (DTW) and deep learning networks like Convolutional Neural Network (CNN) and Recurrent Neural Network (RNN). This paper intends to focus on the various attacks in the ASV system along with these two fundamentals of the system.

Read the paper · More papers on PaperTik