Linear Frequency Residual Cepstral Features for Replay Spoof Detection on ASVSpoof 2019
Priyanka Gupta, Hemant A. Patil · 2022 30th European Signal Processing Conference (EUSIPCO) · 2022
Playing a pre-recorded speech to gain illegal access to Automatic Speaker Verification (ASV) system is one of the easiest attacks to execute but difficult to detect. Such attacks are called as replay attacks. Designing robust ASV systems from such attacks motivates to explore signal processing framework for Spoof Speech Detection (SSD). This paper exploits excitation source-based information in the form of Linear Frequency Residual Cepstral Coefficients (LFRCC) feature set for SSD task. In the source-filter model of speech production, the excitation source is also known to contain speaker-specific information. In this context, the residual obtained from Linear Prediction (LP) of speech is exploited in cepstral domain to detect replay attack. Improvements in results are obtained by choosing appropriate order of LP, as the order of LP controls the amount of information carried by the residual. Experiments performed on ASVSpoof 2019 Physical Access (PA) dataset using Gaussian Mixture Model (GMM) and Convolutional Neural Network (CNN) show that the optimal LP order is 8 which gives EER on the evaluation set as 17.30% and 15.21% using GMM and CNN classifiers, respectively.