Deepfake Detection with Wavelet-Integrated Convolutional Networks
Supriyo Sadhya, Xiaojun Qi · 2024
Deep learning techniques have made it much easier to generate realistic fake content by superimposing or replacing existing images, videos, or audio with highly realistic alternative content. These manipulations often involve the faces or voices of individuals, creating convincing but fabricated representations. Due to the potential misuse of deepfakes for malicious purposes including spreading misinformation, creating fraudulent content, stealing people’s identity, and manipulating public opinion, the development of detection techniques and policies to mitigate harmful effects of deepfakes has become an important research area. In this paper, we combine both spatial and wavelet features to develop a simple yet effective model to detect deepfakes. Specifically, we pass the input color image through the first convolutional layer and employ a one-level wavelet transform to decompose the channel-wise sum of each batch of features. We then thresh-hold the subbands and reconstruct a single channel feature map, which is concatenated with the original batch of features and passed onto the subsequent layers to capture the facial manipulations using both spatial and frequency features. The expanded wavelet transformed features are fed into the VGG19 backbone to help detect deepfakes with an improved detection performance. We perform both within and cross domain evaluations to compare the performance of the proposed model and state-of-the-art peer models in terms of Area Under Curve (AUC) and Equal Error Rate (EER) metrics. Our extensive experimental results demonstrate that the proposed wavelet-integrated VGG19 model offers a more robust solution than the peer wavelet-integrated Xception model and both VGG19 and Xception baseline models to combating the proliferation of fake multimedia content on digital platforms.