Environment Sound Classification using stacked features and convolutional neural network

Shilpa Gupta, Varun Srivastava, Deepika Kumar · 2024

Environmental Sound Classification (ESC) finds a vital application in wildlife conservation, audio-video systems, music instrument classification, automatic speech recognition systems, sound event detection etc. Many conventional model training methods depending upon an enormous amount of annotated data, have been proposed in the literature for the same. The proposed technique focuses on the use of CNN for classifying short audio clips of environmental sounds. Existing models use auditory features like either Log-Mel spectrogram (LM) or Mel Frequency Cepstral Coefficient (MFCC) etc. for the classification or improvement, we have stacked all the different features visualized using the Librosa libraffiry such that it combines all the feature information into one image. The accuracy of the network is evaluated on the ESC-50 dataset of environmental and urban recordings. The stacked features seemed to perform better for the dataset chosen. Three stacked features are provided as input in form of channels to a transfer learned model, which outperforms the CNN models that trained from scratch. The highest precision and recall are obtained for Log-Mel Scale Spectrogram, Spectral Contrast and chroma features which is 95.91% and 95.81% respectively.

Read the paper · More papers on PaperTik