WavLM and Omni-Scale CNNs: Enhancing Boundary Detection in Partially Spoofed Audio

Menghan Li, Zhihua Huang · 2024

Partially spoofed/fake audio, in which segments of utterances are replaced with synthetic or natural audio clips, has emerged as a new form of deep audio forgery, posing potential severe threats to societal security. To address this issue, we employ a deep learning-based frame-level detection system, introducing a frame-level detection approach. We explore the effectiveness of WavLM in waveform boundary detection, utilizing WavLM as a feature extractor for original audio samples. Acoustic features and frame-level embeddings are concatenated, with an OS block embedded within the frame-level feature extraction process, in conjunction with CNN-1D to form the ResNet-OS network. This system can detect partially deceived audio and pinpoint the manipulated segments, effectively integrating multi-scale convolution to consider the interplay between local and global information in time series classification tasks. Experimental results show that this approach offers substantial generalization capabilities and robustness compared to traditional frame-level detection techniques.

Read the paper · More papers on PaperTik