Cross-Teager Energy Cepstral Coefficients for Replay Spoof Detection on Voice Assistants
Rajul Acharya, Harsh Kotta, Ankur T. Patil, Hemant A. Patil · 2021
Voice assistants (VAs) are highly vulnerable to replay attacks, where the impostor plays pre-recorded voice samples to gain an unauthorized access to personalised devices. To that effect, we present an optimal microphone-channel selection scheme using Cross-Teager Energy Operator (CTEO) for spoofed speech detection (SSD) task. Here, a channel refers to the speech signal obtained from a single microphone among the microphone array. The key idea of this work is optimal channel selection based on maximum cross-energies from a multichannel input, which is suitable for SSD task. This newly proposed feature set is named as Cross-Teager Energy Cepstal Coefficients (CTECCmax). The reason be hind maximizing the cross-energies is to identify the distortions in replay speech signal which is added due to intermediate devices. This key idea is also cross-validated by selecting the least estimated cross-energies as feature set CTECCmin. The noticeable improvement in the performance is observed for CTECCmaxover CTECCminfor two classifiers, namely, Gaussian Mixture Model (GMM) and Light Convolutional Neural Network (LCNN).