Ensembling Residual Networks for Multi-Label Sound Event Recognition with Weak Labeling
Hadjer Ykhlef, Zhor Diffallah, Adel Allali · 2022
Sound event recognition is concerned with the development of systems that are able to identify and distinguish events. In realistic settings, sound events originate from different sources and often overlap, which can make the design of such systems more challenging. To mend with this, state-of-the-art recognition systems substantially rely on training multilabel deep neural networks. This process usually requires a large set of labeled audio data. However, most existing datasets are usually small in size or large but unlabeled. Moreover, hand-labeling is a very costly and time-consuming process. In this paper, we design a recognition system that learns from both the labeled and the unlabeled audio clips following the semi supervised learning paradigm. Our system operates in three main stages: (1) Train a baseline, a Residual Network (ResNet), on the labeled data. (2) Use the baseline to generate pseudo-labels of the unlabeled data. (3) Resume training the baseline on both the labeled and the unlabeled data, along with the inferred pseudo-labels. To demonstrate the efficacy, we have conducted experimental comparison on FSDKaggle2019 dataset made of sound clips with annotations of varying reliability. We have tested the improvement over the baseline, and have performed comparison with a multitask-ResNet model trained using the unverified labels. In addition, we have studied ensembling various variants of our approach. The experimental results indicate the superiority of our system over the other alternatives.