Deep Learning for Explicit Content Classification in Music Lyrics

Raissa Camilla Maringka, Green Arther Sandag · 2024

The rapid advancement of technology has transformed how music is accessed and consumed, particularly among younger audiences, making explicit content—such as themes of violence, sexuality, and coarse language—a growing concern for their development. Most music platforms lack automated systems to effectively filter such content, creating a need for more reliable detection methods. This research develops machine learning models, specifically LSTM and BERT, to detect explicit content in English-language music lyrics. The evaluation results show that while the LSTM model achieves an 88% accuracy with strong precision, recall, and F1-score, it still presents some False Negatives and False Positives. In contrast, the BERT model demonstrates superior performance with 94 % accuracy and high precision, recall, and F1-score values. These findings highlight the effectiveness of BERT in content detection, offering a significant advancement over existing methods and providing a robust solution for music platforms to safeguard younger users from harmful content, ensuring a safer and more responsible digital listening environment.

Read the paper · More papers on PaperTik