MFCC-MGD-PE Multi-Feature with CNN-LSTM Hybrid Network for Synthetic Speech Detection

Lili Zhang, Kejing Wei, Xiuli Huang · Journal of Physics Conference Series · 2025

Abstract As synthetic speech technology advances rapidly, the threat to voice security grows. To effectively address this challenge, this paper propose a synthetic speech detection framework that combines multiple features with CNN-LSTM model. This framework integrates three complementary features: Mel-Frequency Cepstral Coefficients (MFCC), Modified Group Delay (MGD), and Permutation Entropy (PE), to form feature vectors. It uses CNN-LSTM hybrid neural network for feature learning and classification. Experiments were conducted on the ASVspoof 2019 dataset and synthetic speech samples from various mainstream synthetic software. The experimental results show that the method performs well in recall, accuracy and ERR, which fully demonstrates the advantages of three features in capturing speech information from different perspectives and effectively distinguish between natural speech and synthetic speech, moreover, offering higher accuracy and robustness. This study offers an effective solution to bolster voice security.

Read the paper · More papers on PaperTik