Pop Noise Detection Using Group Delay Cepstral Coefficients
Arth Shah, Prathav Kevadiya, Hemant A. Patil · 2024
This proposed study leverages the distinctive pop noise, which is involuntarily generated when pronouncing specific phonemes, such as plosive, fricative, nasal, and affricate sounds. We employ Group Delay Cepstral Coefficients (GDCC) as the key features for the Voice Liveness Detection (VLD) task. To measure the effectiveness of GDCC, we compared the performance using different spectral features, namely, Mel Frequency Cepstral Coefficients (MFCC), Linear Frequency Cepstral Coefficients (LFCC), and Gammatone Frequency Cepstral Coefficients (GFCC), employing them with three types of classifiers, namely, Convolutional Neural Networks (CNN), Convolutional Recurrent Neural Networks (CRNN), and Localized Convolutional Neural Network (LCNN). We achieved an accuracy of 79.55% for the GDCC feature vector as its highest, which is comparatively much higher than other feature vectors, thereby proposing a new optimal approach for the VLD task.