Multimodal Forgery Detection in Videos Using a Tri-Network Model

Padma Pooja Chandran, Shruti Sriram, Chitra Babu · 2024

The advent of social media has tremendously revolutionized information dissemination. However, it has also opened up fresh avenues for deception, notably through multimodal forgery. Moreover, cutting-edge Machine Learning (ML) algorithms like Generative Adversarial Networks (GANs) have opened up possibilities to create highly realistic forged content (i.e., images, audio and video) to propagate disinformation through social networks (i.e., Facebook, Instagram, Twitter, etc.), posing a serious menace to society. To address this serious threat, this paper proposes a multi-modal forgery detection architecture that comprises a Video Swin Transformer for video forgery detection and a Swin Transformer for audio forgery detection. It further proposes an audio-visual network using a lip-syncing model, Wav2lip, to generate synthetic lips for the original audio and compare it with the original lip features extracted from the video. The proposed architecture aims to achieve a greater accuracy for detecting audio and video deepfakes. Experimentation was performed on the FakeAVCeleb dataset to evaluate the performance of two prominent transformer architectures, namely Vision Transformer and Swin Transformer, within the context of the audio network. Experimental results underscore the effectiveness of the Swin Transformer in the audio network, with a superior accuracy of 99.88%.

Read the paper · More papers on PaperTik