A Comprehensive Comparative Analysis of Deepfake Detection Techniques in Visual, Audio, and Audio-Visual Domains
Ahmed Ashraf Bekheet, Amr S. Ghoneim, Ghada Ahmed Khoriba · 2024
In recent years, the rise of social media platforms has made them vital channels for sharing news, where audio and visual content play a crucial role in enhancing the credibility of news content. However, significant Artificial Intelligence (AI) progress has introduced new techniques and tools for manipulating multimedia content. These advancements have made it easier to create fabricated digital media, leading to a harmful impact on sharing misinformation, especially in fake news. Consequently, an urgent need arises to explore prevailing methodologies for detecting fake images, audio, and videos, accompanied by a comprehensive exposition of their strengths and limitations. Our survey addresses these methodologies and conducts a rigorous comparative analysis of diverse approaches using various metrics and datasets. We categorize these approaches into visual-based, audio-based, and audio-visual-based deepfake detection methods, encompassing techniques employed across domains. Additionally, we examine notable datasets utilized in detecting image, video, and audio deepfakes, offering insights into their attributes and appropriateness for evaluation purposes. Our findings highlight the effectiveness and limitations of current detection methods, providing a roadmap for future research in multimodal deepfake detection. This includes exploring emerging facets of video manipulation, such as text overlays and motion patterns, investigating advanced deep learning architectures like Transformers, and emphasizing the need for extensive, diverse, and publicly accessible datasets to enhance the robustness and validation of detection methods.