Vocal Melody Extraction using Attention U-net and Voice Detection Model

Bing Zhu, Hong Wei Ding, Hui Wang, Huaichang Du · 2020

The melody extraction of polyphonic music is always an important research topic. Most melody extraction algorithms are based on computing a saliency function of pitch candidates or separating the melody source from the mixture. With the popularity of deep learning, data-driven methods based on deep neural networks, are gaining more and more attentions in the research of melody extraction. In this paper, our data-driven approach based on deep learning to treat melody extraction as a pixel-wise multi-task classification problem in image processing. The proposed model is composed of the advanced semantic segmentation Attention U-net that predicts the pitch salience map of the singing melody and voice detection model that estimate a voice activity of the polyphonic music and use a Softmax and an Argmax layer to get the final melody estimation. Experiments show that the proposed method can achieve better experimental results with small training dataset.

Read the paper · More papers on PaperTik