Estimation model for the speech-quality dimension “Noisiness.”
Lu Huo, Marcel Wältermann, Ulrich Heute, Sebastian Möller · The Journal of the Acoustical Society of America · 2008
State-of-the-art assessment method of speech-transmission quality (e.g., PESQ or TOSQA) predict the mean-opinion score (MOS) quite accurately, but cannot provide diagnostic information, which is, however, highly desirable for system developers. In our research project, we aim at the development of an attribute-based speech-quality measure, which provides estimates of different attributes of speech samples and then maps them to one integral-quality estimate. Three dominant, mutually orthogonal perceptual dimensions were firstly identified by auditory experiments and multidimensional analysis (MDA) for narrow-band speech transmission: “directness/ frequency content,” “continuity,” and “noisiness.” The present paper focuses on the further decomposition and measurement of the global dimension “Noisiness.” Therefore, an auditory test including samples degraded by different kinds of noises has been conducted. The subsequent MDA indicates that at least two sub-dimensions (SD), “Speech Contamination” and (perceived) “Additive-Noise Level,” are further describing the global dimension “Noisiness.” The first SD characterizes the degree the noise distorts the speech signal as such, whereas the second SD reflects the degree the additive circuit or background noise itself annoys the listener. The instrumental estimation methods for both SDs and the mapping to the integral-quality ratings are presented in this paper.