Importance of Energy and Spectral Features in Gaussian Source Model for Speech Dereverberation
Tomohiro Nakatani, Biing-Hwang Fred Juang, Takuya Yoshioka, Keisuke Kinoshita, Masato Miyoshi · 2007
This paper introduces speech dereverberation based on a time-varying Gaussian source model (GSM) and investigates its behavior to provide a better perspective on solving the dereverberation problem. GSM is a generalization of the autocorrelation codebook (ACC) that has recently been shown to enable us to achieve high quality speech dereverberation with only a few seconds' observation. Based on GSM, the speech dereverberation is formulated as a likelihood maximization problem with multi-channel linear prediction, where the reverberant speech signal is transformed into one that is probabilistically more like clean speech. For investigation purposes, the autocorrelation matrix of the GSM is first decomposed into energy, vocal tract filter, and excitation signal features by adopting an autoregressive GSM (ARGSM), and then analyzed based on experiments. They reveal that the energy feature in the models plays a major role in reducing the reverberation components. It is also shown that the other spectral features in the models further contribute to the recovery of the short-time characteristics of the dereverberated signals.