Acoustic Scene Classification from Binaural Signals using Convolutional Neural Networks

Rohith Mars, Pranay Pratik, Srikanth Nagisetty, Chongsoon Lim · 2019

In this paper, we present the details of our proposed framework and solution for the DCASE 2019 Task 1A -Acoustic Scene Classification challenge.We describe the audio pre-processing, feature extraction steps and the time-frequency (TF) representations employed for acoustic scene classification using binaural recordings.We propose two distinct and light-weight architectures of convolutional neural networks (CNNs) for processing the extracted audio features and classification.The performance of both these architectures are compared in terms of classification accuracy as well as model complexity.Using an ensemble of the predictions from the subset of models based on the above CNNs, we achieved an average classification accuracy of 79.35% on the test split of the development dataset for this task.In the Kaggle's private leaderboard, our solution was ranked 4 th with a system score of 83.16% -an improvement of ≈ 20% over the baseline system.

Read the paper · More papers on PaperTik