Multi Label Sound Classification using Deep Learning Models
Tasnim Akter Onisha, Jongyeop Kim, Jongho Seol · 2024
Accurate and automated sound classification enables a strong groundwork for diverse advanced deep learning applications within the audio and music domain. This study focuses on the application of Convolutional Neural Networks (CNN) and combined LSTM (Long Short-Term Memory) and GRU (Gated Recurrent unit) models for instrument classification from audio signals, contributing to intelligent audio processing systems. Our proposed model exclusively utilizes the Mel-frequency cepstral coefficients (MFCCs) extraction from the audio data for preprocessing. A large and complex dataset, including Nineteen instrument classes are used for training and evaluation. These experimental results demonstrate promising performance, with our proposed CNN architecture achieving an impressive accuracy of 97%, and the LSTM-GRU model achieves a lower accuracy of 80%, compared to the CNN model on the multi-label sound classification task for instruments classes, but its ability to model temporal dependencies add valuable insights into the dynamics of instrument audio sequences. These findings provide valuable insights for researchers and practitioners in audio signal processing and machine learning.