Measuring Speech Quality of System Input while Observing only System Output
Stephen Voran · 2021
We present a set of relatively small-scale proof-of-concept experiments where we construct no-reference (NR) speech quality estimators that give reliable values of system-under-test (SUT) input speech quality in spite of the fact that NR estimators can only access SUT output speech. We then explain why this success is not as counter-intuitive as it might initially seem. Next we demonstrate that this advance can be used to adjust NR relative speech quality values to obtain the much more desirable and useful NR absolute speech quality values. The experiments start with over seven hours of studio-quality speech. A processor adds filtering, reverberation, and noise to simulate the somewhat lower quality speech that often must be used to test systems. Four different established full-reference speech quality estimators provide ground-truth values for these experiments.