Advancing Deepfake Detection: Integrating Time Series Processing with Vision Transformers

Srabonti Deb, Tasmia Jannat, Mohiuddin Ahmed · 2024

Deepfake technology allows the creation of convincing fake videos, where individuals appear to be saying or doing things they never did. This emerging capability represents an alarming evolution in the landscape of misinformation with significant implications for personal privacy, security, and the integrity of media content. While previous works have explored various methods for deepfake detection, they often struggle with accuracy and adaptability across different manipulation types. This research introduces a novel method that employs Vision Transformer(ViT) architecture to detect and classify deepfake videos. ViTs, initially developed for image processing, have been modified to handle video data by processing frames in batches, capturing temporal patterns, and detecting inconsistencies over time. The proposed model processes sequences of frames, utilizing the self-attention mechanism to identify manipulations based on spatial and temporal discrepancies across frames. The model was trained and evaluated on the Faceforensics++ dataset encompassing various types of manipulations including DeepFake, Face2Face, FaceSwap, NeuralTextures, and FaceShifter. Through rigorous evaluation, the model demonstrates exceptional accuracy, precision, recall, and F1 scores, outperforming existing methods in deepfake detection. The proposed method demonstrated a detection performance that reached 83.63% accuracy. The inclusion of batch processing of frames in our model enabled us to detect and address temporal irregularities and artifacts that are typically associated with video manipulation. This study contributes to the advancement of digital forensics tools offering a strong solution for identifying and combating the rapid growth of deepfake content in multimedia contexts.

Read the paper · More papers on PaperTik