Compressed Speech Steganalysis Through Deep Feature Extraction Using 3D Convolution and Bi-LSTM

Kesen Li, Feipeng Gao, Jie Chi Yang · IEEE Access · 2025

Voice-over-Internet-Protocol (VoIP) based speech steganography techniques provide convenience for covert communication while posing significant threats to network security. Accurately detecting hidden information in voice signals is of critical significance for cybersecurity. In this paper, we focus on comprehensive feature extraction to enhance detection accuracy across varying embedding rates. We propose a steganalysis approach that combines 3D convolution with Bi-LSTM, incorporating an attention mechanism. Firstly, the encoded speech data is partitioned into three-dimensional data blocks, and 3D convolution is used to capture the correlation between consecutive frames. Subsequently, Bi-LSTM is employed to extract the contextual features. Finally, the extracted features are fused and fed into a classification network to determine whether secret information is embedded within the speech stream. We design corresponding experiments on a publicly available dataset comprising 41 hours of Mandarin speech and 72 hours of English speech encoded with the G.723 standard. The proposed method demonstrates superior performance, achieving accuracy improvements of approximately 2%, 3%, 3%, and 4% over state-of-the-art methods.

Read the paper · More papers on PaperTik