Audio Deepfake Detection: End-to-End training with powerful pretrained ASR

Kh Hamad Mansoor, Mehreen Alam · 2024

The advancements in audio deepfake generation raise significant concerns about its potential misuse, with implications ranging from personal reputation damage to global repercussions. While efforts are being made to mitigate this threat, the results so far are not particularly impressive. Many proposed solutions falter when tested on datasets that differ substantially from the ones on which they were trained. In our approach, we utilize Meta's MMS-300m pretrained ASR model as a feature extractor and train it end-to-end (E2E) alongside various classifiers (ResNet-18, MesoNet, AASIST, MLP, and SLP). We train on a small subset of the widely recognized ASVSpoof2021 DF dataset and conduct cross-dataset evaluations on the In-The-Wild (ITW) dataset, a standard benchmark. The E2E training of the robust ASR model yields a significant improvement, with all our models surpassing the current state of the art. Our best-performing model achieves an Equal Error Rate (EER) of 0.0342%, representing an impressive 55.48% improvement.

Read the paper · More papers on PaperTik