A minimum mean squared error estimator for single channel speaker separation
Aarthi M. Reddy, Bhiksha Raj · 2004
The problem of separating out the signals for multiple speakers from a single mixed recording has received considerable atten- tio ni n recent times. Most current techniques are based on the principle of masking :i n order the separate out the signal for any speaker, frequency components that are not believed to be- long to that speaker are suppressed. The signals for the speaker is reconstructed fro mt hepartial spectral information that re- mains. In this paper we present a different kind of technique - one that attempts to estimate all spectral components for the desired speaker. Separated signals are derived from the com- plete spectral descriptions so obtained. Experiments show that this method results in superior reconstruction to masking based methods. form representations of the various speakers by hidden Markov models (HMMs). The parameters of the HMM for any speaker are learnt from training data recorded from the speaker. In addi- tion, Roweis assumes that the log energy in any frequency band of the mixed signal at any time can be attributed to only one of the speakers. This log-max assumption is justified by two observations. First, when two or more speakers speak simulta- neously, at any time, any given frequency band is usually domi- nated by a single speaker. Second, in any given frequency band the disparity in the energy levels of the dominant speaker and the other speakers is such that the logarithm of the sum of the energies of the individual speakers can be well approximated by the logarithm of the energy of the dominant speaker. In or- der to reconstruct the signal for any speaker, Roweis estimates the mask for that speaker, i.e. the identity of the time-frequency locations where the speaker dominates. The entire signal is re- constructed entirely from the masked spectrum for the speaker, i.e. fro mt he spectral components identified by the mask. The results achieved with this method are remarkably good. Hershey et. al. (6) augment audio recordings with visual features, such as lip and facial movement, in order to enhance the separation. Additionally, the ys eparate the signal into mul- tiple frequency bands, which are then processed independently. As in Roweis' algorithm, the signals for the individual speakers are reconstructed from masked spectra.