Attention-Integrated CNN for Precise Sound Source Localization with Microphone Arrays

Elham Yazdankhah, Salman Karimi · 2024

Sound source localization (SSL) in noisy reverberant environments using microphone arrays presents a significant challenge that has garnered considerable research interest. Recent advancements in deep learning have demonstrated improved performance in these classification tasks by rethinking traditional approaches to sound source localization. This paper presents a novel convolutional neural network (CNN) model to determine sound sources’ direction of arrival (DOA) in noise and reflection conditions. The model combines time-domain features (root mean square energy) and frequency-domain features (sound intensity, spectral contrast, Mel-frequency cepstral coefficients) extracted from audio signals. Three types of experiments were conducted to evaluate the model’s robustness: assessing the impact of additive noise, examining the effects of reverberation time, and analyzing how room conditions influence microphone arrays. The experimental results reveal that the proposed model outperforms existing base systems in the literature, achieving the highest accuracy in DOA estimation.

Read the paper · More papers on PaperTik