Time-varying residual noise feature model estimation for multi-microphone speech recognition
Takuya Yoshioka, Emmanuel Y. J. Ternon, Tomohiro Nakatani · 2012
This paper proposes a method for compensating for the effect of noise remaining in a signal generated by a multi-microphone signal enhancer in the feature domain as a post-processing. The proposed method assumes that the multi-microphone signal enhancer generates estimates of both the target and original environmental noise signals. To obtain a time-varying residual noise feature model that responds to noise changes quickly and is consistent with a clean feature model, the proposed method leverages both the multiple signal estimates provided by the signal enhancer and the clean feature model. Specifically, the proposed method first roughly estimates residual noise features on a frame-by-frame basis by comparing the target and noise signal estimates. Then, these rough estimates are refined by using the clean feature model to yield a time-varying residual noise feature model. Experimental results show the effectiveness of the proposed method and its wide applicability.