Data Diversity for Improving DNN-based Localization of Concurrent Sound Events
Daniel Krause, Archontis Politis, Konrad Kowalczyk · 2021 29th European Signal Processing Conference (EUSIPCO) · 2021
Sound source localization (SSL) is an actively researched topic in the field of multichannel audio signal processing with numerous practical applications. Since it is used in different acoustic contexts, ensuring a good generalization of the techniques and models to various acoustic signals and environments is of great importance. In this paper, we aim to investigate the influence of different types of sources on the training process of a model based on a deep neural network (DNN). We present several training datasets, containing different mixtures of noise, speech and sound events, and perform a comparative study in which we test the trained models on distinct target signals. The problem is analyzed in the context of localizing two simultaneously active sound sources. Additionally, two data augmentation methods are incorporated into the framework to verify their impact on model generalization. The results of experiments performed for two concurrent sources show the localization accuracy of models trained with diverse data types. In particular, training using speech and sound event data or mixtures thereof are shown to result in localization accuracy increase for a variety of source types under test.