Dynamic Thresholding on FixMatch with Weak and Strong Data Augmentations for Sound Event Detection
Tanmay Khandelwal, Rohan Kumar Das · 2022 13th International Symposium on Chinese Spoken Language Processing (ISCSLP) · 2022
Recent state-of-the-art (SOTA) semi-supervised learning methods have shown great promise in improving sound event detection (SED) performance when the labeled data is scarce. Using a combination of consistency regularization, pseudo-labeling, and data augmentation techniques, the model predictions are constrained to be noise invariant. The recently proposed FixMatch achieved SOTA results on SED tasks. However, it uses a pre-defined constant threshold throughout the training process to generate the pseudo-labels, thus failing to account for the learning difficulties for each class and the model learning stage. To address this issue, we propose a dynamic thresholding method as an extension to FixMatch for generating pseudolabels based on the model’s predictions on weakly augmented features. This method retains the generated pseudo-labels based on the dynamic threshold value. The model is then trained to predict the generated pseudo-label when fed with a strongly augmented version of the same feature. On DCASE 2022 Task 42022 dataset, our method helped us in improving the SED system performance by 34.22% compared to the baseline in terms of polyphonic sound event detection score.