Exploiting the harmonic structure for speech enhancement
Eunjoon Cho, Julius O. Smith, Bernard Carlos Widrow · 2012
We provide a single channel speech enhancement method leveraging the harmonic structure of voiced speech. A sinusoidal model, based on the pitch of the speaker, is used to filter noisy speech and remove any noise components that lie between the harmonics. To remove noise that lie on each harmonic frequency, we use a noise estimation procedure that exploits spectral sparsity of voiced speech. By measuring the power spectrum at frequencies that correspond to the zero crossings of the windowing function, we can estimate the noise levels even in frames that have voiced speech. We also provide a constrained linear least squares formulation to reduce “musical noise” which arises from difficulty in estimating speech and noise power spectral densities. We show that our method yields high perceptual performance over existing methods, and can easily adapt to conditions in which the noise characteristics are constantly changing.