Same but different? — Using speech signal features for comparing conversational VoIP quality studies
Sebastian K. Egger, Raimund Schatz, Katrin Schoenenberg, Alexander Raake, Gernot Kubin · 2012
In this paper we demonstrate how speech signal features can be used to detect and explain differences in human to human conversation tests. To this end, we compare the results of two conversational VoIP quality experiments designed to quantify the impact of network delay on perceived speech quality. Both studies followed the same procedures and used the same scenarios, but were conducted in two different labs. Our comparison shows that the two studies, despite having been executed correctly using the same test design, still can produce surprisingly different results regarding the users quality perception on a MOS scale. In this respect, speech signal features extracted from conversation recordings help identifying divergent participant behavior as plausible cause for such differences. Our in-depth analysis reveals how novel parameters developed by us like Intended and Unintended Interruption Rate (IIR, UIR) and the corrected Speaker Alternation Rate SARcorrcan be used to successfully determine the extent to which the results of different conversational speech quality studies are directly comparable and thus eligible for pooling, or not.