Sound Event Detection Using Attention and Aggregation-Based Feature Pyramid Network
Ji Won Kim, Geon Woo Lee, Hong Kook Kim, Nam Kyun Kim · 2022 27th Asia Pacific Conference on Communications (APCC) · 2022
This paper proposes a sound event detection (SED) model using an EfficientNet-B2 and an attention and aggregation-based feature pyramid network (A2-FPN). In particular, the EfficientNet-B2 is first obtained from the pretrained model on the basis of the pretraining, sampling, labeling, and aggregation (PSLA) framework. Then, the A2-FPN module is applied to the outputs of the layers of the EfficientNet-B2 to deal with the different time and frequency resolutions from acoustic features. The aggregated feature map from the A2-FPN module is used as input features to two bidirectional gated recurrent unit layers. Specifically, the proposed A2-FPN-based SED model is trained by the mean-teacher approach to utilize weakly labeled and unlabeled data. Finally, the proposed A2-FPN-based SED model is applied to the detection and classification of acoustic scenes and events (DCASE) 2021 Challenge Task 4. Consequently, it is shown that the polyphonic sound event detection score (PSDS) 1 and 2 of the proposed A2-FPN-based SED model are the higher of 0.03 and 0.172, respectively, than those of the DCASE 2021 Challenge Task 4 baseline.