A Method of Sound Event Localization and Detection Based on Three-Dimension Convolution

Pengcheng Mei, Jibin Yang, Qiang Zhang, Xiang Yi Huang · 2022 7th International Conference on Image, Vision and Computing (ICIVC) · 2022

Deep Learning methods represented by convolutional neural networks can jointly realize Sound Event Detection (SED) and Sound Source Location (SSL). However, due to the noise and reverberation in real scenes, the accuracy of direction estimation is still dissatisfactory. Since three-dimensional convolution can carry out convolution calculation in time, frequency and channel domains for multichannel input simultaneously, it can learn more inter-channel and intra-channel features and effectively solve the above problems compared to two-dimensional convolution. Inspired by it, a method based on three-dimension convolution feature extraction called SELD3Dnet is proposed. The amplitude and phase characteristics of input multi-channel audio are calculated, and the deep feature representation is extracted through multiple 3D convolutional structures. Finally, the category and spatial location of sound events are estimated by recurrent neural networks and fully connection layers. Comparative experiments are conducted on TUT2018 datasets, and the results show that the proposed method improves the F1 metric by 13.9% and the frame recall metric by 21.1% on average under various types of real scene data subset ov1, ov2, ov3, which can validate the performance of the proposed method.

Read the paper · More papers on PaperTik