HODGEPODGE: Sound Event Detection Based on Ensemble of Semi-Supervised Learning Methods

Ziqiang Shi, Liu Liu, Huibin Lin, Rujie Liu, Anyan Shi · 2019

In this paper, we present a method called HODGEPODGE 1 for large-scale detection of sound events using weakly labeled, synthetic, and unlabeled data in the Detection and Classification of Acoustic Scenes and Events (DCASE) 2019 challenge Task 4: Sound event detection in domestic environments.To perform this task, we adopted the convolutional recurrent neural networks (CRNN) as our backbone network.In order to deal with the small amount of tagged data and the large amounts of unlabeled indomain data, we aim to focus primarily on how to apply semisupervise learning methods efficiently to make full use of limited data.Three semi-supervised learning principles have been used in our system, including: 1) Consistency regularization applies data augmentation; 2) MixUp regularizer requiring that the predictions for a interpolation of two inputs is close to the interpolation of the prediction for each individual input; 3) MixUp regularization applies to interpolation between data augmentations.We also tried an ensemble of various models, which are trained by different semi-supervised learning principles.Our approach significantly improved the performance of the baseline, achieving a event-based f-measure of 42.0% compared to 25.8% of the baseline on the official evaluation dataset.Our submissions ranked third among 18 teams in the task 4.

Read the paper · More papers on PaperTik