Feature Genuinization based Residual Squeeze-and-Excitation for Audio Anti-Spoofing in Sound AI

Ruchira Ray, Sanka Karthik, Vinayak Mathur, Prashant Kumar, G. Maragatham, Sourabh Tiwari, Rashmi T Shankarappa · 2021

Voice modality in human-machine interaction has gained popularity in the last decade due to advances in voice technology. All digital devices support voice as input while employing voice assistants. It is the most used way of interaction in headless digital appliances and IoT devices. Emerging audio spoofing techniques pose a significant threat to Automatic Speaker Verification (ASV). False wakeup of voice assistants and their response on recorded audio replay imposes security concerns and customer's hesitancy. As applications of ASV and replay detection are ubiquitous, it is essential to make these systems robust. We propose a two-stage hybrid model: genuinization transformer to efficiently differentiate between the distribution of synthetic and genuine speech and non-speech audio, followed by Residual Squeeze-and-Excitation networks (ResSEnet) to learn relevant latent features and classify audio input as spoofed and bonafide. To handle both speech and non-speech audio sounds effectively, we use log-mel features. The proposed model is evaluated using the ASVspoof 2019 Logical Access (LA) dataset. Experimental results show that our proposed model significantly elevates performance compared to the baseline and state-of-the-art models.

Read the paper · More papers on PaperTik