NMF based speech and music separation in monaural speech recordings with sparseness and temporal continuity constraints
Ming Tu, Xie Xiang, Yishan Jiao · Advances in intelligent systems research/Advances in Intelligent Systems Research · 2013
Abstr act.This paper proposes a semi-supervised approach of speech and music separation in monaural speech recordings based on non-negative matrix factorization (NMF).Considering the scenario that the genre of background music is known, music basis vectors are randomly picked from the magnitude of short time fourier transform (STFT) of training music, while speech basis vectors are estimated by executing NMF on the magnitude of STFT of polluted speech signal.Moreover, we apply sparseness and temporal continuity constraints to speech and music respectively and evaluate how different constraints can influence the separation performance.The test set contains 10 Mandarin speech utterances from 10 speakers mixed with music in different speech-music ratios (SMR).The baseline is semi-supervised separation system with no constraint.The results reveal that adding temporal continuity constraint can improve the separation performance compared with the baseline and separation system with only sparseness constraint.Keywor ds: non-negative matrix factorization • speech and music separation•sparse coding•temporal continuity•semi-supervised learning 1 Intr oduction Speech and music separation belongs to one specific problem of audio source separation, the applications of which include: speech enhancement when talking to mobile phone in loud music background, online video transcript interfered by music due to the amateurism of the uploader, in-car speech recognition with the background music from CD player or FM radio, etc.Although the research on audio source separation has achieved good results in these years and some of the methods also show feasibility on separation of speech and music, the problem of 1