POLQA+: Adding voice quality assessment to ITU-T’sPerceptual Objective Listening Quality Assessment(POLQA)

T.M. Schoegje · Utrecht University Repository (Utrecht University) · 2017

Most often, speech quality is modelled as the audio quality of a system-enabled transfer between talker and listener. However, with a reference signal being absent, solely audio quality cannot explain subjective quality assessments' results. Here, we propose to extend the current speech quality model and include audio as well as voice, and pronunciation quality aspects. To achieve this, a large feature space of speech characteristics is explored and the composite model is trained and tested, then validated, using two distinct data sets. This yields POLQA+: ITU-T's speech quality standard (Recommendation P.863), extended with solely the talker's medium pitch during vowels. POLQA+ explains 71.18% of data speech quality on unseen signals, with a RMSE* of only 0.011. Additionally, we show that adding features would easily improve this result; but, could not yet be expected to generalize to unseen contexts. This issue should be studied further to unleash POLQA+'s true impact and present a fine-mazed speech quality model.

Read the paper · More papers on PaperTik