Music Classification Model Development Based on Audio Recognition using Transformer Model

Andreas Aditya Alvaro Harryanto, Kevin William Gunawan, Rio Nagano, Rhio Sutoyo · 2022

There are several music genres, such as blues, classical, disco, and more. Music genres can be predicted by supervised machine learning using extracted features. This research explores the Transformer deep learning model for audio classification using music genres as labels. The dataset used in this paper is the GTZAN dataset, which contains music in a waveform audio file (WAV) format. The features of this dataset are extracted using Mel-frequency Cepstral Coefficients (MFCC). This study finds that the Transformer model's performance gives higher accuracy than other neural network architectures. Furthermore, the Transformer model's performance is compared with related research results that utilize various traditional machine learning methods with the same dataset. The best result achieved was the accuracy of 75.1% for the Transformer model with the best fine-tune parameters.

Read the paper · More papers on PaperTik