Review of Existing Methods for Generating and Detecting Fake and Partially Fake Audio

Abdulazeez AlAli, George Theodorakopoulos · 2024

Using deep-learning technologies, both text-to-speech (TTS) and voice conversion (VC) methods can generate fake speech effectively, making it challenging to differentiate between real and fake speech. Accordingly, researchers have employed deepfake detection solutions to distinguish them. These solutions can achieve high detection accuracy and exhibit robustness against unseen data, which are data that differ from those used in initial model training. The emergence of partially fake (PF) audio, which combines real and fake speech, presents a new challenge for deepfake detection. This tutorial presents a comprehensive overview of TTS, VC, and PF generation and detection methods and analyses the characteristics of publicly available datasets for each type. Furthermore, it highlights directions for PF detection that can pave the way for valuable research in fake speech detection.

Read the paper · More papers on PaperTik