Residual drum sound estimation for RPCA singing voice extraction

Shiori Mikami, Arata Kawamura, Youji Iiguni · 2017

In this paper, we propose an efficient singing voice extraction method from a music sound consisting of a singing voice and a drum sound. The proposed method is based on robust principal component analysis (RPCA) which is a technique to separate a given matrix into a sparse matrix and a low rank matrix. Dealing with a spectrogram of the music sound as the given matrix, RPCA gives an extracted singing voice spectrogram as a sparse spectrogram. We improve the capability of RPCA method by removing a residual drum sound from the sparse spectrogram. The residual drum sound is estimated as spectral components which repeatedly arise in the sparse spectrogram. Removing the estimated residual drum sound spectrogram from the sparse spectrogram gives an improved singing voice spectrogram. Simulation results show that the proposed method can improve 10dB of GNSDR as compared with the conventional RPCA singing voice extraction method.

Read the paper · More papers on PaperTik