Spatial Vision Transformer: A Novel Approach to Deepfake Video Detection
Phạm Minh Thuấn, Lam Thu Bui, Pham Duy Trung · 2024
Deepfake, a technology that leverages artificial intelligence to create manipulated videos, has emerged as a significant threat to security and privacy. In recent years, Transformer-based models, such as Vision Transformer (ViT) and Convolution Vision Transformer (CViT), have demonstrated high efficacy in detecting deepfake videos. However, optimizing the Convolutional blocks in CViT remains a challenge. This paper proposes an enhanced CViT model by replacing the traditional ConvBlocks and incorporating Spatial and Channel Reconstruction Convolution (SCConv) blocks. Experiments conducted on the Deepfake Detection Challenge (DFDC) dataset indicate that the new model not only improves detection accuracy but also reduces false detection rates, underscoring its potential in deepfake detection.