VOCAL-TRACT MODELING FOR SPEAKER INDEPENDENT SINGLE CHANNEL SOURCE SEPARATION

Michael M. Stark, Franz Pernkopf, Van Tuan Pham, Gernot Kubin · 2010

In this paper, we investigate two statistical models for the source-filter based single channel speech separation task. We incorporate source-driven aspects by pitch estimation in the model-driven method which models the vocal-tract part as a priori knowledge. This approach results in a speaker independent (SI) source separation method. For modeling the vocal tract filters Gaussian mixture models (GMM) and non-negative matrix factorization are considered. For both methods, the final fusion of the source and filter parameters results in a reformulation of the models that finally are used for separation. Furthermore, for the GMM method we propose a new gain compensation and pitch adjustment method. Performance is evaluated and compared to the speaker dependent (SD) factorial Hidden Markov Model [1]. Although the SD method delivers the best quality our SI methods show promising results and possess a lower complexity in terms of used parameters. Index Terms — single channel speech separation, source-filter modeling, Gaussian mixture model, non-negative matrix factorization 1.

Read the paper · More papers on PaperTik