Hierarchical Classification Networks for Singing Voice Segmentation and Transcription

Zih-Sing Fu, Li Su · Zenodo (CERN European Organization for Nuclear Research) · 2019

Identifying the onset and offset time of a musical note is a challenging step for singing voice transcription, as the soft onset/offset, portamento, and vibrato phenomena are rich in singing voice signals. In this work, we investigate how to utilize local data representation with pattern recognition for onset and offset detection of singing voice. We consider onset and offset detection as a hierarchical classification problem, where every local data representation as input is classified into one of all the possible event states in monophonic singing, namely the silence, activation, and transition states, and the transition state is further classified into the onset and offset states. An objective function based on this hierarchical taxonomy nicely guides the model to capture the complicated temporal dynamics of note sequences. Multi-channel data representations containing spectral differences and pitch saliency are employed to reflect the patterns of note transition in singing voice signals. The proposed method implemented with residual networks provides improved performance over prior art in onset and offset detection. Moreover, by integrating with a pitch detection framework, the proposed method also outperforms previous singing voice transcription methods. This result emphasizes the importance of note segmentation in singing voice transcription.

Read the paper · More papers on PaperTik