DRN-BiLSTM-Based Algorithm for Sound Source Localization in Reverberant Environments

Yiyi Cai, Xinmin Ren, Yu Wei Zhu, Yang Chen · 2025

Sound source localization (SSL) is a critical branch of array signal processing, where improving accuracy in reverberant environments remains a persistent challenge. To achieve high-precision and efficient localization under such conditions, we propose a neural network architecture that combines dilated residual convolution (DRN) with bidirectional long short-term memory (BiLSTM). The model employs generalized cross-correlation with phase transform (GCC-PHAT) sequences from microphone arrays as input features. The DRN layers first extract multi-scale spatial features, while the BiLSTM module captures temporal dependencies in acoustic signals. Experimental results show that the proposed DRN-BiLSTM model consistently outperforms traditional methods and existing deep learning approaches in various reverberant conditions, achieving a $26.2 \%$ to $36.7 \%$ reduction in mean absolute error (MAE) and a $\mathbf{3 2. 1 \%}$ to $\mathbf{3 3. 2 \%}$ reduction in median absolute error (MedAE), while also maintaining high localization accuracy in previously unseen environments. These improvements significantly enhance the accuracy of three-dimensional sound source localization.

Read the paper · More papers on PaperTik