Analysis of call-quality prediction performance for speech-only and audio-visual telephony
Benjamin M. Weiss, Sebastian Möller, Benjamin Belmudez, Błażej Lewcio · 2014
In this paper, we will compare several approaches for predicting the quality of entire speech-only or audio-visual telephone calls from subjective judgments or predictions of the quality of individual segments of these calls. The comparison will be done using several databases which have been collected in a test paradigm to simulate conversational structures. The subjective judgments for individual conversation segments obtained in these tests, as well as instrumental quality estimations for these segments, are used as the basis for predicting episode-final quality ratings. Although different modeling approaches reach a similar performance, an optimum model is proposed which leads to comparably high prediction accuracy for speech-only and audio-visual telephony. Both the test paradigms as well as the call-quality prediction models are subject to standardization activities of ITU-T Study Group 12, for which this analysis is considered an input.