Multi-Modal Deepfake Detection using AI: Combining Audio, Visual, and Metadata Cues to Enhance Detection Accuracy and Robustness
United States of America, Sai Santhosh Polagani · International Journal of Innovative Research in Science Engineering and Technology · 2025
Quick developments in deepfake technologies have created serious worries about media and cyber issues, as well as the trust of the public. Depending on a single style of data evidence such as video or sound, traditional methods are ineffective against advancing deepfake creation. The framework put forward in this study connects audio, visual and metadata information to help improve both the accuracy and the durability of detection algorithms. Using voice ticks, eye movements, body gestures and file information, we train machines to spot realistic content from altered videos or images with great accuracy. An architecture for the hybrid neural network combines CNNs, RNNs and attention mechanisms to uncover strong connections between multimodal data. Experience on standard datasets has shown that the proposed model with multiple inputs performs better than a single input model, especially when only small or weak changes take place. As a result of this research, it is clear that using all aspects together is essential to fight deepfakes and can help with real-time, reliable use in digital forensic and content verification systems.