Lyric recognition in monophonic singing using pitch-dependent DNN
Dairoku Kawai, Kazumasa Yamamoto, Seiichi Nakagawa · 2017
One of the difficulties in sung speech recognition is the small distance in an acoustic space between phonemes in sung speech. Therefore we considered clustering the speech based on a pitch (fundamental frequency F0) and creating a larger distance between the phonemes. In addition, we considered a two-stage training method of DNN-HMM: the first stage is trained by using conventional acoustic features like MFCCs, and the second stage is re-trained by augmented features with a pitch feature. We expected to train pitch information more explicitly in the second stage of the training and obtained a relative improvement of 9% as expected.