Audio Segmentation and Classification Approach Based on Adaptive CNN in Broadcast Domain

Sun Jingzhou, Wang Yongbin, Xiaosen Chen · 2019

Audio segmentation and classification is the basis of broadcast audio processing in the broadcast domain. A number of methods have been developed over the years in order for the broadcast station to segment and classify the audio into 7 classes, which are: male speech, female speech, speech with noise, speech with music, noise, music, and silence. Some of the major methods include SVM, GMM and HMM among others, however the overall accuracy of these existing methods is not satisfactory. This paper, therefore proposes an adaptive CNN method which classifies clips directly based on audio sample points. To our knowledge, this is the first work to do segmentation and classification in broadcast domain by Adaptive CNN. Each convolutional layer is equivalent to the clip feature extraction. The study uses the method of Adaptive CNN by increasing the amount of computation to ensure that the most suitable features can be extracted. In other words, Adaptive CNN increases the process of feature selection compared to CNN. A large number of experiments were carried out and compared with other approaches and the results show that the Adaptive CNN technique achieved a significant error reduction.

Read the paper · More papers on PaperTik