Using Autotagging for Classification of Vocals in Music Signals

Nuno Pinto Hespanhol Lopes dos Santos · Open Repository of the University of Porto (University of Porto) · 2013

Modern society has drastically changed the way it consumes music. During these last recent years, listeners have become more demanding in how many songs they want to have accessible and require to access them faster than ever before. The modern listener got used to features like automatic music recommendation and searching for songs that, for example, have female vocals'' and ambient'' characteristics. This has only been possible due to sophisticated autotagging algorithms. However, there has been an increasing belief in the research community that these algorithms often report over optimistic results. This work approaches this issue, in the context of automatic vocal detection, using evaluation methods that are rarely seen in literature. Three methods are conducted for the evaluation of the classification model developed: same dataset validation, cross dataset validation and filtering. The cross dataset experiment shows that the concept of vocals is generally specific per dataset rather than universal as expected. The filtering experiment, which consists of iteratively applying a random filterbank, shows drastic performance drops, in some cases, from a global f-score of 0.72 to 0.27. However, these filters have been showed not to affect the human ear's ability to distinguish vocals, by conducting a listening experiment with over 150 candidates. Additionally, a comparison between two binarization algorithms - maximum and dynamic threshold - is performed and shows no significance difference. The results are reported on three datasets that have been widely used within the research community, on which a mapping from its original tags to the vocals domain was performed and which is made available to other researchers.

Read the paper · More papers on PaperTik