Single-Channel Multitalker Speech Recognition

Steven J. Rennie, John R. Hershey, Peder A. Olsen · IEEE Signal Processing Magazine · 2010

We have described some of the problems with modeling mixed acoustic signals in the log spectral domain using graphical models, as well as some current approaches to handling these problems for multitalker speech separation and recognition. We have also reviewed methods for inference on FHMMs (factorial hidden Markov model) and methods for handling the nonlinear interaction function in the log spectral domain. These methods are capable of separating and recognizing speech better than human listeners on the SSC task.

Read the paper · More papers on PaperTik