Co-channel speaker separation using constrained nonlinear optimization

Daniel S. Benincasa, Michael I. Savic · 2002

This paper describes a technique to separate the speech of two speakers recorded over a single channel. The main focus of this research is to separate overlapping voiced speech signals using constrained nonlinear optimization. Based on the assumption that voiced speech can be modeled as a slowly-varying vocal tract filter with a quasi-periodic train of impulses, the speech waveform is represented as a sum of sine waves with time-varying amplitude, frequency and phase. In this work the unknown parameters of our speech model are the amplitude, frequency and phase of the harmonics of both speech signals. Using constrained nonlinear optimization, we determine, on a frame by frame basis, the best possible parameters that provides the least mean square error (LMSE) between the original co-channel speech signal and the sum of the reconstructed speech signals,.

Read the paper · More papers on PaperTik