Audio Environment Identification

Rashmi R. Patil, Rashmika Patole, Priti P. Rege · 2019

An audio recorded in the claimed environment is subject to a number of possible manipulations and thereby, the necessity of a reliable system to identify the audio environment is high. This paper first focuses on combining the state-of-the-art features and later the use of an appropriate classifier amongst Support Vector Machine (SVM), Artificial Neural Network (ANN) to help improve the classification accuracy. Also the use of handcrafted standard featured as an input to Convolutional Neural Network (CNN) is shown. In this paper, a comparative study is done between the standard MFCCs+Δs, GFCCs+Δs and MFCCs+Δs, GFCCs+Δs obtained after dereverberation. Also, reverberation feature called Decay rate distribution (DRD) from Reverberation time (RT) are combined along with the above mentioned ones to look forward to further improvement in accuracy and obtain useful observation results. The improvement in accuracy with the combination features used is demonstrated using Dataset1 generated synthetically using classifiers ANN and SVM, amongst which ANN perform better than SVM with the highest obtained accuracy of 91.67% for the combined feature set MFCCs+Δs+DRD. Dataset 2 from the DCASE challenge 2018 is used for standard MFCCs and GFCCs as features and CNN as classifier.

Read the paper · More papers on PaperTik