Performance Analysis of Deep Learning Based Speech Quality Model with Mixture of Features
Rahul Kumar Jaiswal · 2022
Speech is one of the convenient mediums of communication among humans. However, the quality of speech deteriorates due to the surrounding noise. To fulfil the expected level of quality of experience (QoE) of the end-user while exploiting distinct applications, such as Microsoft Skype, Apple FaceTime to name a few, it is important to measure the speech quality and monitor it in the real-time. To this end, this paper investigates a series of deep neural network (DNN)-based objective no-reference speech quality models (SQMs) in accurately measuring speech quality. Three speech features, namely, line spectral frequencies (LSF), mel-frequency cepstral coefficients (MFCC), and multi-resolution auditory model (MRAM) are extracted from the speech signal after processing it through a voice activity detector (VAD). A series of DNN-based SQMs is, then, developed by incorporating either a single or a mixture of speech features. The standard no-reference speech quality prediction model (P.563) is employed as a baseline model. Results demonstrate that the DNN-based SQM trained with MRAM feature performs better in accurately measuring speech quality as compared to the baseline model and other DNN-based SQMs trained with different speech features or their mixtures.