Deep Multi-view Features from Raw Audio for Acoustic Scene Classification
Arshdeep Singh, Padmanabhan Rajan, Arnav V. Bhavsar · 2019
In this paper, we propose a feature representation framework which captures features constituting different levels of abstraction for audio scene classification.A pre-trained deep convolution neural network, SoundNet, is used to extract the features from various intermediate layers corresponding to an audio file.We consider that the features obtained from various intermediate layers provide the different types of abstraction and exhibits complementary information.Thus, combining the intermediate features of various layers can improve the classification performance to discriminate audio scenes.To obtain the representations, we ignore redundant filters in the intermediate layers using analysis of variance based redundancy removal framework.This reduces dimensionality and computational complexity.Next, shift-invariant fixed-length compressed representations across layers are obtained by aggregating the responses of the important filters only.The obtained compressed representations are stacked altogether to obtain a supervector.Finally, we employ the classification using multi-layer perceptron and support vector machine models.We comprehensively perform the validation of the above assumption on two public datasets; Making Sense of Sounds and open set acoustic scene classification DCASE 2019.