Robust Multilingual Audio Deepfake Detection Through Hybrid Modeling
Candy Olivia Mawalim, Yutong Wang, Aulia Adila, Shogo Okada, Masashi Unoki · 2025
The increasing sophistication of AI-generated human voice poses a significant threat, demanding robust detection systems that can generalize effectively across diverse linguistic environments and synthesis techniques.In response to the SAFE Challenge, this paper introduces a novel approach to multilingual audio deepfake detection.Our primary contribution lies in the comprehensive study of deepfake detection using a multilingual speech corpus encompassing 17 languages and a broad spectrum of synthesis methods and acoustic conditions, designed to enable more realistic and challenging evaluations.To optimally utilize this diverse data, we propose a hybrid detection model that synergistically combines the strengths of end-to-end RawNet and AASIST architectures with language-agnostic representations learned from a multilingual selfsupervised learning model.Additionally, we explore the efficacy of RawBoost data augmentation in enhancing robustness against realworld noise.Our experimental evaluation demonstrates promising generalization in generated audio detection, achieving approximately 73% balanced accuracy across multilingual data and unseen synthesis algorithms.