DEEP-TRUST: DEEPFAKE DETECTION VIA HYBRID CNN-ELA-GAN

Nallamilli V K Reddi, Karangi Sindhu, Shaik Ahmed Aaquelah, Mohammad Azarunnisa, Pradhan Nikhil · International Journal of Engineering Applied Sciences and Technology · 2025

—Deepfakes pose a growing threat to digital security and trust, demanding robust methods to detect AI-generated manipulations. Traditional approaches like XceptionNet and Error Level Analysis (ELA), while foundational, struggle with evolving generative architectures like diffusion models and fail to balance accuracy with interpretability. Static forensic methods also lack adaptability to dynamic adversarial attacks. This study introduces Deep-Trust, a hybrid framework integrating Convolutional Neural Networks (CNNs), Error Level Analysis (ELA), and Generative Adversarial Networks (GANs) to expose and localize deepfake artifacts. We analyzed 5,000+ images from benchmark datasets (Face Forensics++, Celeb-DF) and identified 12 critical forensic features, including compression anomalies, texture inconsistencies, and spectral distortions. A twostage preprocessing pipeline was designed: first, ELA amplifies pixel-level compression artifacts by recompressing images at varying JPEG quality levels, and second, a GAN-based adversarial training module generates synthetic deepfakes to harden the detector against unseen manipulations. Unlike conventional models, the CNN-ELA-GAN framework dynamically optimizes feature weights and hyperparameters through adversarial training, enhancing both detection accuracy and computational efficiency. The dual-branch CNN processes ELA maps and raw images in parallel, fusing low-level forensic cues with high-level semantic features. Gradient-weighted Class Activation Mapping (Grad-CAM) further localizes tampered regions (e.g., distorted eyes or synthetic hair textures) with humaninterpretable heatmaps. Evaluated on 10,000+ samples, Deep-Trust achieved 98.7% accuracy and 28 FPS inference speed on NVIDIA A100 GPUs, outperforming XceptionNet (94.2% accuracy) and ELA-only baselines (82.1% accuracy). The optimized 12-feature configuration reduced training time by 33% (from 120s to 80s per epoch) while maintaining robustness against adversarial attacks. This work demonstrates the synergy of forensic analysis, deep learning, and adversarial training for combatting deepfakes, offering a scalable solution for real-time social media moderation, digital forensics, and secure authentication systems.

Read the paper · More papers on PaperTik