Applications and Comparative Analysis of Machine Learning and Statistical Learning Signal Processing in Music Genre Classification
Yue Ping Wu, Zhu Qing Cheng, Kaige Zhou, Baixuan Chen, Guanchao Tong · 2025
Music genre classification is a crucial task in audio signal processing and machine learning, widely applied in music recommendation systems, streaming platforms, and digital libraries. This study presents a comprehensive framework integrating advanced preprocessing techniques, feature extraction, and a comparative analysis of traditional machine learning models and deep learning architectures. The preprocessing pipeline employs the Short-Time Fourier Transform (STFT) to extract time-frequency domain information, complemented by the computation of Mel-Frequency Cepstral Coefficients (MFCC) and Mel spectrograms as essential features. Data augmentation and normalization are utilized to enhance the robustness and generalization of the models. The framework evaluates traditional machine learning methods, including Support Vector Machines (SVM), Naive Bayes, Random Forest, and XGBoost, alongside deep learning architectures such as Convolutional Neural Networks (CNN), Long Short-Term Memory (LSTM) networks, and Gated Recurrent Units (GRU). The performance of these models is evaluated using the GTZAN Music Genre Dataset, a widely used benchmark in the field of music genre classification, which provides a rich and diverse set of music genres for testing. Experimental results assessed using metrics like accuracy, F1 score, and confusion matrix, demonstrate that deep learning models excel at leveraging time-frequency features, while traditional models perform effectively with smaller datasets. By combining Fourier Transform-based preprocessing with robust modeling strategies, this research offers a systematic approach to improving the music genre classification performance and provides valuable insights for advancements in music signal processing.