Significance of Quadrature and In-Phase Components for Synthetic Spoofed Speech Detection

Priyanka Gupta, Piyushkumar K. Chodingala, Hemant A. Patil · 2022 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) · 2022

For synthetic and voice converted Spoofed Speech Detection (SSD), Instantaneous Frequency (IF)-based features either exploit Hilbert Transform (HT) or Teager-based Energy Separation Algorithm (ESA) to estimate IF. However, HT-based approach leads to poor resolution in time-domain, and ESA-based approach leads to the lack of relative phase information. To that effect, we propose CFCCIF-QESA feature set, which encompasses excellent time resolution as well as relative phase information. Hence, we illustrate the significance of incorporating quadrature-phase component along with in-phase component for SSD of synthetic and voice converted spoof. The proposed feature set is evaluated using various performance metrics, namely, EER, MCC, F-measure, J-statistic, Jaccard Index, and Hamming loss. CFCCIF-QESA achieves a relative decrease in Average EER (AEER) by 22.53% w.r.t. the performance of CFCIF-ESA, for all the attacks in the dataset. Furthermore, results w.r.t. S10 attack (i.e., MaryTTS using unit selection synthesis), which is the most difficult to detect attack, show that CFCCIF-QESA achieves a 29.16% relative decrease in EER as compared to the CFCCIF-ESA. Furthermore, using model-level measures such as, Kullback Leibler Divergence (KLD) and Jenson Shannon Divergence (JSD), we also show the better discriminative ability of our proposed feature set CFCCIF-QESA than the existing CFCCIF-ESA. Finally, the analysis of the latency period for CFCCIF-QESA and CFCCIF-ESA is presented, and it shows better suitability of CFCCIF-QESA w.r.t. deployment in practical SSD systems.

Read the paper · More papers on PaperTik