Time-frequency Resolution Optimization Features on Spoof Detection

Zetian Li, Jianguo Wei, Qilong Sun · 2021

With the improvement of synthesized speech and converted speech technology, computers can simulate speech that is difficult for human ears to distinguish, and speaker recognition systems are extremely vulnerable to such fraudulent speech attacks. Therefore, the research on deceptive speech detection attacks is of great significance for improving the reliability of speaker recognition systems. In this article, we study the characteristics of the Automatic Speaker Verification Deception and Countermeasure Challenge (ASVSpoof2019) dataset based on Fourier Transform-based Linear Frequency Cepstral Coefficients (LFCC) and Constant Q-Transformed Constant Q Cepstral Coefficients (CQCC). And analyzed these features in the time resolution and frequency resolution optimization method and energy expression optimization method. The back-end classifier uses Gaussian Mixture Model (GMM) and deep residual network as the back-end classifier. The result of using equal error rate (EER) on the GMM classifier is 6.98%, which is an increase of 13.72% compared with baseline LFCC.

Read the paper · More papers on PaperTik