A comparison of three algorithms for estimating aspiration noise in dysphonic voices
Mina Goor, Rahul Shrivastav, John G. Harris · The Journal of the Acoustical Society of America · 2004
Dysphonic voices are frequently characterized by increased aspiration noise. Several algorithms have been proposed to quantify the noise present in such voices. Yet, an independent analysis of these algorithms has not been reported. These algorithms differ in a number of aspects such as the theory underlying these measurements, the procedures used for estimating the noise, and the nature of their output. Three different algorithms for estimating noise in dysphonic voices were implemented in MATLAB and their output for synthetic and natural voice samples compared. These algorithms include: (1) the pitch-predictive signal-to-noise ratio reported by Milenkovic [Workshop on Acoustic Voice Analysis: Proceedings (1994)], which analyzes signals in the time domain and provides a time waveform of the aspiration noise: (2) the harmonics-to-noise ratio reported by deKrom [J. Speech Hear Res., 36, 254–266 (1993)], which performs a cepstral analysis to provide the spectrum of the aspiration noise; and (3) the glottal-noise excitation reported by Michaelis, Frohlich, and Strube [J. Acoust. Soc. Am. 103, 1628–1639 (1998)], which measures the correlation of the Hilbert transform of different frequency bands. The result of the study will help identify the algorithm(s) most suitable in the prediction of the listener judgments of voice quality. [Research supported by NIH/R21DC006690-01.]