Multi-Class Urban Sound Classification with Deep Learning Architectures
Mita Avadhani, Jahnavi, Anupama P Bidargaddi, S Thushara · 2024
In the area of urban sound classification, our re-search delves into the challenge of correctly categorizing diverse urban soundscapes. Employing deep learning models like Arti-ficial neural networks (ANNs), Convolutional neural networks (CNNs), Recurrent neural networks (RNNs), and Long Short-Term Memory (LSTM) networks, we aim to decode complex patterns within audio data for optimal classification. Utilizing Mel-frequency cepstral coefficients (MFCCs) for feature extraction, we explore the strengths and trade-offs of each architecture. Our study not only advances urban sound classification but also emphasizes the advantages and considerations of each model, providing practical insights for real-world applications in urban audio analysis. The Artificial Neural Network (ANN) showcases an impressive accuracy of 94.79 %, emphasizing its robust per-formance in capturing key sound features. Following closely, the Convolutional Neural Network (CNN) achieves a notable accuracy of 93.64 %, highlighting its proficiency in spatial feature extraction. The Recurrent Neural Network (RNN) achieves an accuracy of 65.14%, offering insights into its suitability for specific sound classification tasks. Additionally, the Long Short-Term Memory (LSTM) architecture emerges as a significant contributor with an accuracy of 86.09 %, underlining its effective-ness in capturing complex temporal dependencies. These results highlight the diverse strengths of each architecture and LSTM in particular performs exceptionally well in classifying urban sounds, which enhances the possibility of practical applications in this field.