Multimodal Deepfake Detection Across Cultures and Languages

Abhinav Dhall · 2025

The phenomenal growth of generative AI methods has made high-quality multimodal synthetic content generation possible. While synthetic data generation has many valuable applications, it has also resulted in deepfakes. These deepfakes are increasingly used to spread misinformation and disinformation. To address this challenge, we need deepfake detectors, which can be deployed at scale in real-world settings and for different usecases. Such detectors should be trained on large and diverse datasets and use methods, which generalize well to data generated from unseen methods and also provide explainable results. Furthermore, they must be effective across different languages and cultural contexts to ensure broad inclusivity. The 1 Million Deepfakes Detection Challenge is designed to provide a large-scale benchmark for detecting and localizing deepfakes. The dataset currently includes more than two million samples. Results from the participating methods highlight the limitations of existing approaches, particularly in accurately localizing the manipulated segments. In many societies, people mix languages and dialects during conversation. From the perspective of deepfakes, such data is harmful as it is non-trivial to detect for observers. Most deepfake datasets are monolingual, contain clean audio and are focused on Western languages. We need multilingual, code switching, dialect diverse audio and video datasets with realistic artifacts. ArEnAV dataset is one such resource consisting of Arabic-English language based code-switching. ArEnAV is a large-scale audio-visual benchmark centered on Arabic-English code switching across MSA, Egyptian, Levantine, and Gulf, totaling over 765 hours and 387,072 clips, generated using 4 TTS and 2 lipsync models. Furthermore, most existing deepfake datasets focus on simple face swaps or object changes, but they miss more complex and realistic edits that can change the meaning of an image. To address this, we propose MultiFakeVerse, a large dataset comprising of deepfakes generated using language based reasoning. A Vision Language Model first identifies the main person in an image, decides on a realistic manipulation (such as changing an object they hold or altering information about them), and then generates the edited image with a diffusion model. When tested on MultiFakeVerse, current detection methods that work well on traditional deepfakes perform much worse on these more higher-level, meaning-altering manipulations.

Read the paper · More papers on PaperTik