Enhanced Audio Signal Classification with Explainable AI: Deep Learning Approach in Time and Frequency Domain Analysis

A. Emily Jenifer, K Abirami, M. R. Rajeshwari · Procedia Computer Science · 2025

The music genre classification is highly used in the recommender system and content organization. It is challenging for humans to manually categorise millions of pieces of audio signals into distinct or related classes Consequently, a machine learning model can accurately categorise them. This paper implements the audio signal classification with a song genre classification dataset. Time domain feature extractions, like amplitude envelope and crest factor, and frequency domain features, like spectrogram and spectral centroids, are extracted and processed in 2D lightweight CNN for effective classification. The GTZAN dataset was used to analyze this model. The 1000 audio files in the GTZAN dataset are categorized into 10 music genres: Pop, Disco, Metal, Classical, Rock, Reggae, Blues, Country, Hip-Hop, and Jazz. Each music file has a duration of 30 seconds with rich and diverse instrumental, vocal styles, tempos, and rhythms. By combining both time and frequency domains of those audio files, the proposed method leverages the strengths of both types of analysis and effectively by capturing both local and global audio patterns. Thus, the custom-made lightweight model performs better, has 93.75% accuracy, and outperforms the existing classification models. Further, to enhance the interpretability of the model classification, the features contributing to the individual music genres are analyzed with the Local Interpretable Model-Agnostic Explanations (LIME) explainable artificial intelligence (XAI) model. The results of LIME highlight the most influential segment of the audio, which aids in classification.

Read the paper · More papers on PaperTik