Efficient Classification of Partially Faked Audio Using Deep Learning

Abdulazeez AlAli, George Theodorakopoulos, Ahmed Emad · 2025

The rise of synthetic and manipulated audio content, especially partial fake speech, presents significant challenges for verifying audio authenticity. Partial fake speech refers to segments of audio in which only certain parts have been altered or synthesized, making it more difficult to detect compared to fully synthetic speech. This paper introduces a novel detection model specifically designed to identify partial fake speech. Our approach incorporates Wav2Vec 2.0 as a feature extractor, along with max pooling, conformer blocks, attention-based pooling, and fully connected layers. Experimental results on two datasets demonstrate the model’s effectiveness in detecting partial fake speech. Our models outperforms existing methods in terms of Equal Error Rate (EER), achieving 0% on the RFP dataset and 2.99% on the ASVSpoof 2019 LA dataset.

Read the paper · More papers on PaperTik