On Inter-Rater Reliability for Crowdsourced QoE

Tobias Hoßfeld, Michael Seufert, Babak Naderi · 2021

Crowdsourcing offers a faster, cheaper, and more scalable approach than the traditional laboratory quality assessment tests. However, participants perform the test in their own working environment, using their own hardware and without direct supervision of a test moderator, leading to different types of biases on the ratings. In this paper, we compare several reliability metrics that are commonly applied to the subjective ratings in terms of their sensitivity to identify typical issues of crowdsourced media quality tests. Following the subject bias theory, we simulate the ratings of different user groups with different bias and various magnitudes of uncertainty, while also considering the presence of unreliable raters. We apply traditional reliability metrics on the ratings and compare their sensitivity in identifying the severity of the raters' biases and uncertainties. Our results show that the average Spearman's rank correlation coefficient between raters can serve as a strong indicator for issues with the crowdsourcing study. This means that scoring too low for this metric should encourage researchers to revisit their study design in order to eventually improve the reliability of results from crowdsourcing-based quality studies.

Read the paper · More papers on PaperTik