Deep Autoencoder-based Framework for Robust Singer Identification in Music Analysis
Sangeetha Rajesh, N. J. Nalini · 2024
This paper introduces a robust classification system identify music clips based on singers using Mel-Frequency Cepstral Coefficients (MFCC) and Chroma Energy Normalized Statistics (CENS) with Stacked Autoencoder (SAE). The dataset consists of Indian language songs from ten renowned playback singers, spanning different musical eras. Identifying singers in polyphonic music, where instrumental accompaniment adds complexity, is a significant challenge. This paper aims to propose a system capable of effectively extracting vocal features for accurate identification. The primary objective is to evaluate the system's identification performance using metrics such as identification rate and Equal Error Rate (EER), focusing on MFCC and CENS feature representation and SAE-based feature learning. The Stacked Autoencoder offers key advantages in feature learning by effectively reducing the dimensionality of input data while capturing important vocal characteristics. Its ability to learn hierarchical feature representations makes it well-suited for handling the complexity of polyphonic music, where interrelated time and frequency features are crucial for identifying singers. Experimental results demonstrate an average identification rate of 96.9% for polyphonic music. The findings highlight the system's potential to accurately identify singers despite the challenges posed by instrumental accompaniment, showcasing the benefits of SAE in music signal analysis.