Evaluation Singing Art Voice Quality Using Siamese Gated Recurrent Extreme Learning Networks

Yang Dong · 2025

The quick evolution of artificial intelligence and deep learning has provided new chances for singing art voice quality assessment, where exact evaluation of the vocal performance is critical to training, music education, and intelligent feedback systems. Current methods tend to depend on hand-engineered acoustic features, rule-based processing, or shallow machine learning models that do not handle intricate temporal-spectral dependencies present in singing signals and exhibit weak generalization and weak adaptability to varied voice types. Conventional methods like MFCC-based classification, simple CNN architectures, and recurrent networks have demonstrated some progress but are faced with high computation, overfitting, and inadequate robust attention mechanisms. To cope with the above challenges, a Siamese Gated Recurrent Extreme Learning Networks (SGRELN) is proposed by incorporating gated recurrent units, extreme learning machines and Siamese modules for emphasizing prominent vocal features. Experimental results clearly show noteworthy superiority, where RAT-CNN attains 98.8% accuracy and 97.7 % F1-score compared to traditional models. Comparative mechanism emphasizes the efficacy, strength, and interpretability of the model proposed, with greater reliability in real applications of singing quality evaluation in the real world with superior performance, scalability, and domain adaptability.

Read the paper · More papers on PaperTik