DualStreamNet : Robust Audio-Video Deep Fake Media Detection Using Complimentary Information Fusion
Shreyas Sheeranali, Raghavendra Ramachandra, Sushma Krupa Venkatesh · 2025
Deepfake technology employs sophisticated machine-learning techniques to create highly convincing video and audio recordings of individuals doing or saying things that they never actually did or said. These falsified media pieces have the potential to deceive and manipulate viewers, posing significant risks to their privacy, security, and trust in digital media. In this paper, we present a novel method DualStreamNet for reliable Audio-Video (AV) fake media detection which exploits the complementary information. The proposed DualStreamNet includes independent detector for video and audio modality. We introduced a novel video fake detection framework using a SlowFast encoder as the backbone, and a novel architecture based on a 3D CNN with skip connections. We also introduced novel features to reliably detect audio fakes using a Continuous Wavelet Transform (CWT) Filter Bank that was further processed using the ResNet50 architecture. Finally, the decisions from the video and audio detectors are combined using the logical OR rule to make the final decision. Extensive experiments were performed on two publicly available audio-video fake datasets: FakeAVCeleb and SWAN-DF. The obtained results indicate the improved detection accuracy of the proposed method compared to existing methods.