Classification of Explicit Music Content Based on Lyrics, Music Metadata, and User Annotation
Egivenia Egivenia, Gabriella Ryanie Setiawan, Stella Shania Mintara, Derwin Suhartono · 2021
Over the recent years, there has been numerous studies on explicit music classification in an attempt to classify explicit music which are inappropriate for younger audiences. While several previous studies have been conducted using music lyrics and metadata with different types of classifiers, user annotations have not been used before. This study attempts to implement user annotation into the dataset to see the comparison through the accuracy score. The incorporation of user annotation is addressed in this study by implementing Random Forest and Support Vector Machine classifiers to classify explicit music by using a dataset that consisted of the music lyrics, metadata, and user annotations that we have collected from 200 songs. The resulting models of Random Forest and Support Vector Machine were able to classify music into 2 categories: explicit and non-explicit, with a f1-score of 0.809 and 0.791, respectively, that succeeded in outperforming the results of previous similar research.