Spectrogram-Based CNN Framework for Live Deepfake Voice Detection in Google Meet
Mayank Jyani, Pramod Singh Rathore · 2025
The rapidly growing development of AI-generated voice synthesis is increasing the threat of deepfake audio on real-time communication platforms. Traditional deepfake detection systems, which are mostly based on post-processed data techniques such as MFCC and SVM-fall behind the sophistication of modern synthetic voices, resulting in high false positive rates and limited applicability. This research has proposed a new, real-time deep-fake voice detection framework that integrates with Google Meet via a Chrome Extension. The system uses spectrogram-based feature extraction and a deep Convolutional Neural Network (CNN) to process live audio streams in real-time via the FastAPI and WebSocket protocols, providing instant detection and alerts. Comparative analysis shows that this model is more accurate than traditional models and also reduces false positives, with a classification performance AUC score reaching 0.90. In addition, the study plans a future extension that will include facial deepfake detection, which could make online meetings even more secure. This work is a significant contribution to the field of AI ethics and cybersecurity, as it offers a scalable, deployable, and effective solution that could be useful against real-time audio-based impersonation threats.