Real-time audio-visual deepfake detection using multimodel pipeline in video conference
Ananya Jha, B. Siddharth Bhat, Abhay CM, Chirag Hegde, Ramamoorthy Srinath · 2025
Video conferencing has become a cornerstone of communication in professional, educational, and social contexts globally. Due to the emergence of sophisticated real-time deepfake technology that is open-source and easily accessible to anyone, the trust in the authenticity of participants during live video calls has diminished. As opposed to offline forensic analysis, detecting deepfakes in real-time is inherently more difficult. In this work, we present an approach to real-time deepfake detection that takes place in the context of video conferencing: a twostep verification process designed to verify user authenticity before the call and through additional continuous on-call monitoring, enabling participants to verify authenticity through optional video and audio detection. The framework employs four models, which are active iris pattern analysis, facial hue correlation, ResNeXt- LSTM model, and Mel spectrogram analysis. Complementary strengths of the modules enable reliable detection despite real-world variances, ensuring secure communication. Experiments on custom datasets demonstrate high accuracy under controlled conditions, highlighting challenges in scenarios like dim lighting, complex environments, and noisy audio.