Speech enhancement based on log spectral envelope model and harmonicity-derived spectral mask, and its coupling with feature compensation

Takuya Yoshioka, Tomohiro Nakatani · 2011

The use of a speech spectral envelope model defined in the log spectrum-type domain is a common approach to feature enhancement for noise robust speech recognition. However, from the noise reduction viewpoint, this approach ignores non-peak components of a spectrum and thus suffers from the poor SNR improvement during voiced periods. This paper proposes a speech enhancement method that exploits a log spectral envelope model and a harmonic structure. The key to the method is its use of a harmonic structure to define the prior distribution of a spectral mask, which is used for both accurate noise estimation and attenuation. In addition, we combine log mel-frequency feature enhancement with the above method to take advantage of low dimensionality. The whole proposed method outperforms a state-of-the-art speech enhancement method in four different noise environments.

Read the paper · More papers on PaperTik