A Time-Frequency Modulation Model of Speech Quality
James Mitchell Kates, Kathryn Hoberg Arehart · 2007
A new speech-quality metric, based on time-frequency modulation, is introduced in this paper. The metric uses a cochlear model, with the signal envelope in each frequency band converted to dB above threshold. Envelopes sampled across the frequency bands give short-time spectra that are approximated using a set of mel cepstrum coefficients. The correlation between the cepstral coefficient sequences for the clean and degraded signals is used to compute the quality metric. The metric accurately models quality judgments made by normal-hearing and hearing-impaired listeners for speech degraded by additive noise, nonlinear distortion, and dynamic-range compression.