Fake Speech Detection in Domain Variability Scenario

Rishith Sadashiv T N, Ayush Agarwal, S. R. Mahadeva Prasanna · 2024

Domain variability refers to a condition where the train and test speech data are from different environments. This seems to be challenging to deal with in fake speech detection task. This work extends earlier work done on fake speech detection using features from openSMILE toolkit and a combination of several machine learning classifiers. The first extension is to evaluate earlier work on a larger ASVspoof 2019 LA database. The second extension is to expand the feature size to the entire 88 dimensional features from openSMILE toolkit which shows significantly improved performance for the larger database. The next extension of multistyle training helps in dealing with domain variability scenario. The final contribution employs a two stage approach where the first stage detects the domain and the second stage performs fake vs bonafide speech classification. Both multistyle training and two stage approach seem to handle fake speech detection in domain variability condition with both handcrafted features and state-of-the-art deep learning models.

Read the paper · More papers on PaperTik