Detecting AI-Generated Speech Manipulation through CNN-BiLSTM Hybrid Networks

Rahul Dixit, Arth Agrawal, Anuja Dixit, Sweta Tripathi · 2025

Audio forgery, the malicious manipulation or generation of audio content, poses significant threats to privacy, security, and digital trust. Various audio forgeries are prevalent today, including copy-move forgery, splicing, noise injection, Text-to-Speech (TTS), time stretching, pitch shifting, voice conversion (VC), resampling, and deep forgery. Among these, TTS and VC attacks emerge as particularly concerning due to their ability to mimic genuine voices, enabling identity theft and misinformation convincingly. While existing research predominantly focuses on detecting copy-move forgeries, this paper addresses the relatively underexplored challenges of detecting Text-to-Speech and voice conversion forgeries. A hybrid CNN-Bidirectional LSTM model is proposed, which uses convolutional layers for spectral feature extraction and bidirectional LSTM layers for temporal pattern recognition. To enhance robustness, audio augmentations are applied during preprocessing. The model is built and tested on the ASVspoof 2019 LA dataset, achieving an accuracy of 96.65% in identifying advanced audio forgeries. This study bridges a critical gap in audio forgery detection and provides a foundation for safeguarding audio authenticity against emerging threats.

Read the paper · More papers on PaperTik