Improving the Performance of Automatic Lyrics Transcription by Transfer Learning from Speech Domain
Himanshu Sharma, Ramakrishnan Angarai Ganesan · 2024
Automatic lyrics transcription (ALT) in the singing domain poses unique challenges because of lack of adequate data and the inherent difficulty of transcribing the sung lyrics. How-ever, its counterpart problem in the speech domain, automatic speech recognition has made remarkable progress in recent years, driven by the availability of extensive datasets and self-supervised learning paradigms. This work exploits the similarities between speech and singing to build a high-performing ALT system. Specifically, we develop a solution based on transfer learning, which exploits these similarities by modifying the wav2vec 2.0 model, originally designed for speech recognition, to a singing dataset. Our solution achieves an absolute WER reduction of 10.48% compared to the best results in the literature on DALI vl dataset, both methods being trained and tested exclusively on the DALI vl dataset only. Better results available in the literature have used high amounts of training data combining DSing and DALI v2 datasets.