Enhancing Deepfake Detection: Spatial-Temporal Preprocessing and Self-Attention ResI3D Model
Sang-Ho Son, J. Lee, K. C. Min, Wooju Kim · 2023
Deepfake technology is the outcome of employing deep learning techniques to overlay the face of one individual onto the video of another. As deep learning technology advances rapidly, the proliferation of high-quality deepfakes for malicious digital activities is notably on the rise. With growing concerns about the misuse of deepfake technology, there is an increasing demand for research into deep learning-based methodologies to detect and counteract it. While Deepfake detection using deep learning has been a subject of prior research, these approaches primarily rely on images hence not utilizing temporal information. Additionally, research combining CNN and RNN has inherent limitations. It operates with compressed data, resulting in the loss of spatial information and the utilization of the inherent temporal characteristics in pixel-to-pixel temporal data. In this study, we propose a detection model that harnesses the inherent attributes of video data through self-attention on both the spatial and temporal axes, using the ResI3D model along with the Non-Local Block. Additionally, we conducted experiments during the preprocessing phase to validate and implement methods that facilitate the model's effective learning of both temporal and spatial information. As a result, our model demonstrated enhanced performance when compared to existing deepfake video detection models.