Cascaded ResNet50-LSTM Architecture for Deepfake Video Identification
Jayanti Rout, Minati Mishra · 2025
With the increasing sophistication of AI-generated content, detecting fake videos has become a critical challenge in digital media forensics. Such fake videos can cause loss to individuals and organizations in terms of fiscal matters, goodwill, societal status, etc. Traditional detection methods such as frame-by-frame analysis, motion inconsistencies, etc., were timeconsuming, subjective, and prone to human bias. This research presents a hybrid deep learning framework that combines Convolutional Neural Networks (CNNs) for spatial feature extraction with Long-Short-Term Memory (LSTM) networks for modeling temporal relationships. The proposed model is designed to efficiently differentiate between authentic and AI-generated videos by leveraging both spatial and sequential patterns within video frames. The model is trained and evaluated on the Celeb-DF (v2) dataset. The preprocessing phase involves frame extraction, resizing, and normalization to enhance feature representation. The experimental findings indicate that the proposed model delivers robust performance in the test set: 94. 98% accuracy, 98. 06% precision, 96. 08% recall, 97. 06% F1 score and 98. 32% AUC-ROC. Further evaluation using a confusion matrix and loss accuracy curves highlights the robustness of the model in detecting AI-generated content. This research contributes to improving deepfake detection techniques, ensuring improved security and reliability in multimedia authentication systems.