Novel Spectral Root Cepstral Features for Replay Spoof Detection
Prasad A. Tapkir, Ankur T. Patil, Neil Shah, Hemant A. Patil · 2018
Replay poses a greater threat to the Automatic Speaker Verification (ASV) system than any other spoofing attacks, as it neither require any specific expertise nor a sophisticated equipment. In this paper, we propose a novel countermeasure by modeling the replayed speech as a convolution of genuine speech with additional impulse responses (due to the microphone, loudspeaker, recording and replay environment). In particular, we propose the new feature set, namely, Magnitude-based Spectral Root Cepstral Coefficients (MSRCC) and Phase-based Spectral Root Cepstral Coefficients (PSRCC), that performs better than the baseline system (CQCC), on ASVspoof 2017 challenge database, which gives 29.18% Equal Error Rate (EER) on the evaluation set. The proposed feature set detects the effect of these additional impulse responses, in the quefrency-domain. Experiments performed on evaluation set using MSRCC and PSRCC, with Gaussian Mixture Model (GMM) as a classifier gives 18.61% and 24.35% EER, respectively. On the other hand, Convolutional Neural Network (CNN) classifier gives 24.50% and 26.81% EER, respectively. The score-level fusion of MSRCC and PSRCC gives reduced EER of 10.65% using GMM and 17.76% using CNN classifier, indicates the complementary information captured by the proposed feature sets.