A Novel Deepfake Detection Framework using Audio-Video-Textual Features
S. Asha, P. Vinod, Irene Amerini, Varun G. Menon · Research Square · 2022
Abstract Recent advances in deep learning and computer vision have spawned a new class of forgeries known as deepfakes. Deepfakes raise many legal and ethical concerns, which makes the need to distinguish them demanding. Emerging trends in deepfakes are (i) creation of realistic facial deepfakes, (ii) usage of synthesized human voices and fabricated subtitles in addition to facial threats, and (iii) crafting bogus videos that contain illegal content in between original sequences. To efficiently handle these deepfakes, we propose an ensemble deepfake detection framework named the D-Fence layer. It comprises two unimodal classifiers to identify facial and vocal tainted inputs. Additionally, it incorporates two cross-modal classifiers between related domains, Video-Audio and Audio-Text to recognize deepfakes. Since existing datasets lack fake content across multiple modalities, we construct a multi-modal faux dataset using original videotapes from several existing datasets. To ensure the effectiveness in spotting fake videos generated using modern approaches, we present a novel facial threat ’Bogus-in-the-middle’ attack, which inserts fake video frames between the real ones and ’Downsampling attack’ to craft fraudulent audio. A comparative study of our D-Fence layer against various state-of-the-art multi-modal deepfake detection systems is also conducted. The proposed ensemble architecture outperforms existing unimodal classifiers and efficiently detects deepfakes under several adversarial conditions with a detection accuracy of 89.5%.