Time-frequency methods for enhancing speech
Owen Patrick Kenny, Douglas J. Nelson · Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE · 1997
Speech signals have the property that they are broad-band white conveying information at a very low rate. The resulting signal has a time-frequency representation which is redundant and slowly varying in both time and frequency. In this paper, a new method for separating speech from noise and interference is presented. This new method uses image enhancement techniques applied to time- frequency representations of the corrupted speech signal. The image enhancement techniques are based on the assumption that speech and/or the noise and interference may be locally represented as a mixture of two-dimensional Gaussian distributions. The signal surface is expanded using a Hermite polynomial expansion and the signal surface is separated from the noise surface by a principal- component process. a Wiener gain surface is calculated from the enhanced image, and the enhanced signal is reconstructed from the Wiener gain surface using a time varying filter constructed from a basis of prolate-spheroidal filters.