WaveFuseNet: Landmark-Guided Deepfake Detection with CNN–LSTM–ViT Feature Fusion and Comparative Frequency-Domain Analysis
Firdous Sadaf M. Ismail, Tausif Diwan, Nileshchandra Kalbarao Pikle · Cybernetics & Systems · 2026
In today’s digital era, the line between real and fabricated content has blurred, with deepfakes emerging as potent tools for misinformation, identity theft and manipulation undermining forensic reliability and public trust. Existing detection methods often fail in real-world settings due to dependence on global features and limited interpretability. To overcome these gaps, we present WaveFuseNet, a landmark guided deepfake detection framework evaluated under multiple frequency-domain transforms that is region-aware and frequency-sensitive. The framework evaluates four spectral trans forms DWT, FFT, STFT and Gabor independently on key facial regions (eyes, nose, mouth) identified via MediaPipe landmarks. A real-only ColorJitter augmentation introduces natural variation while retaining synthetic artifacts. For each transform-specific experiment, extracted features are processed through a hybrid encoder combining ResNet18 for spatial learning, LSTM for temporal coherence and a pretrained Vision Transformer (ViT) for global attention. On the Celeb-DF v2 dataset, WaveFuseNet achieves 98.84% accuracy and 96.3% F1-score, demonstrating high robustness for forensic and media verification tasks. Ablation results reveal DWT and Gabor outperform other transforms by exposing subtle edge and texture inconsistencies, while FFT and STFT provide complementary frequency cues. By systematically evaluating these orthogonal frequency representations under a unified CNN–LSTM–ViT backbone WaveFuseNet ensures interpretable and resilient detection against evolving deepfake manipulations.