Multi-Branch Convolutional Macaron net for Sound Event Detection
Teck Kai Chan, Cheng Siong Chin · IEEE/ACM Transactions on Audio Speech and Language Processing · 2021
Sound Event Detection remains a challenging task due to the lack of strongly labeled data. While the use of weakly labeled and unlabeled data can alleviate this issue, most states of the art utilized the Mean Teacher approach, which requires training two identical models in a semi-supervised manner. Such methodology can have two critical limitations. Firstly, it can be computationally expensive if a very deep model is designed. Secondly, a model designed might only be optimal for either audio tagging or frame-level prediction but not both. Thus, using the Mean-Teacher approach may only allow a model to perform at its maximum potential for one of the tasks. However, the aforementioned issues can be circumvented by designing two different models where the less complex model provides the frame-level prediction while the more complex model provides the audio tags. To increase the accuracy of the models, we propose the use of Squeeze and Excite, meta-ACON, an improved Transformer encoding layer, a triple instance-level pooling approach (i.e., multi-branch pooling), and an improved cyclic learning scheme. Based on such a framework, the best system can achieve an event-based F1-score of 48.5%. By ensembling the top 5 models, the event-based F1-score increases to 50.4%. The proposed framework can achieve a minimum margin of over 12% against the baseline system while being competitive to the other state of the arts.