Deep Semantic Encoder-Decoder Network for Acoustic Scene Classification with Multiple Devices
Xinxin Ma, Yunfei Shao, Yong Ma, Wei-Qiang Zhang · Asia-Pacific Signal and Information Processing Association Annual Summit and Conference · 2020
In this paper, we proposed Mini-SegNet, a simplified encoder-decoder SegNet model to capture deep semantic information in sound events. The semantic information can effectively discriminate the acoustic segments in different scenes. We also applied spectrum correction to combat mismatched frequency response. In order to prevent over-fitting, we adopted mixup augmentation, ImageDataGenerator and temporal crop augmentation for data augmentation. Our best single system achieved an average accuracy of 65.15% on different devices in the DCASE2020 Development dataset, more than 10% improvement over the baseline system. The results indicate that our approach can achieve good classification performance, without use of any supplementary data from outside the official challenge dataset.