A Study of Sound Source Localization Based on Improved Time Delay Features and the Convolutional Neural Network
Huitao Feng, Hong Zhang, Zhanxu Shen, Zhixin Qiu, Xiaojun Zhang, Zhi Tao · 2024
Precise sound source localization technology is of great practical significance in many fields. Among them, the delay features extracted based on the generalized cross-correlation method are widely used. However, the applicability of these features varies significantly in different environments. To address this issue, this paper proposes a sound source localization method based on improved generalized cross-correlation and neural networks. By using features extracted from both MLP weighted and SCOT-HB weighted methods as initial values, and employing controllable beam-forming success rate features and a convolutional neural network with dual-channel input for iterative learning, sound source localization is achieved. We conducted sound source localization error tests on multiple models, including the one proposed in this paper, using various task scenarios in the LOCATA database. The results show that even in the noisy environment with silent fragments, the root mean square error of the positioning angle of the model is still less than 5 degrees, and its positioning accuracy is much higher than that of other models of the same kind.