Speeding Up Music Similarity

Elias Pampalk · 2005

This paper describes (1) the submission to the ISMIR’04 genre classification contest and (2) the submission to the MIREX’05 (Music Information Retrieval eXchange) audio-based genre classification and artist identification tasks. The main difference between the submissions is the reduction of computation time in the order of magnitudes. This paper concludes with a discussion of the relationship between genre classification and artist identification, the relationship between similarity and classification, and references to related MIREX’05 submissions. 1 IMPLEMENTATION OVERVIEW Features are extracted from 22kHz mono wav input (2 minutes from the center of each piece are used for further analysis). For the 2004 submission these features are cluster models of MFCC spectra. The 2005 submission additionally uses fluctuation patterns and two descriptors derived from them: Gravity and Focus. For each piece in the test set the distance to all pieces in the training set is computed. A nearest neighbor classifier is used. There is no training other than storing the features of the training data. Each piece in the test set is assigned the genre label (or artist’s name) of the piece closest to it. 1.1 M2K Specific The functions are implemented in Matlab 7 and submitted with an M2K wrapper. The 2004 submission requires the Netlab Toolbox and the signal processing toolbox. The 2005 submission does not require any additional toolboxes. The same functions are used for the genre classification and artist identification tasks. 1.2 Computation Time The CPU times given in Table 1 are measured on a 1.3GHz Intel Centrino laptop. The 2004 submission does not fulfill the MIREX’05 time constraints (72 hours per task). For example, it takes 10 days to compute the (symmetric) distance matrix on a collection with 3000 pieces. The 2005 submission completes this in less than 4 hours. 2

Read the paper · More papers on PaperTik