Acoustic Scene Recognition Based on Convolutional Neural Networks
Fengiiao Sun, Mingiiang Wang, Qihang Xu, Xiaogung Xuan, Xin Zhang · 2019
Audio scene recognition is a process of automatically determining the scene around the device by extracting the features of scene audio signals. It is more about the perception and understanding of non-speech signals, and has a profound guiding significance for the machine to make more intelligent choices. To solve this problem, this paper proposes an audio scene recognition method based on convolutional neural network. Firstly, short-time Fourier transform and Mel filter bank are used to transform the audio signal into log-mel spectrum. Then, log-mel fragments are trained by using CNN neural network, and the features are extracted. Finally, softmax was used to identify and classify CNN features. This method is used to test the data set of IEEE DCASE 2018. Experimental results show that this method has a high recognition rate.