Transformer-based Environmental Sound Classification Modeling by Jointing Multi-class Classification and Similarity Clustering
Xing Zhou, Ming Lei Zhao · Research Square · 2022
Abstract As environmental sound signal is not as regular as speech, it has varying temporal structures and is more difficult to distinguish. Previous research on Environmental Sound Classification (ESC) has designed sophisticated methods to extract feature from raw waveforms directly, but this may not be generalized well across different ESC tasks. We proposed an end-to-end audio scene classification network which is only based on the log-mel feature. First, we used Transformer network to encode signals that can capture crucial temporal information from self-attention. Then we combined Multi-class classifier with similarity clustering so as to maximize the distance between different classes. At last, we visualized the Transformer’s ability to locate important temporal information. The performance on ESC10 and ESC50 showed that our architecture reached an average accuracy of 95.3% and 84.2%, respectively. That was an achievement of new state-of-the-art performance with only log-mel input. Meanwhile, that is nearly equivalent to the best performance of the model based on raw waveform or combined feature method.