PMSQA: Pre-Training Based Multi-Scale Non-Intrusive Speech Quality Assessment

Yi Jiang, Sixing Liu, Shaohan Liu, Qun Yang · 2024

Speech quality assessment plays a crucial role in many voice communication applications. However, the accuracy of the model-based non-intrusive speech assessment method is still not satisfactory. Therefore, this paper proposes a Pre-training based Multi-scale Speech Quality Assessment model (PMSQA) to improve the accuracy of neural network-based non-intrusive speech assessment. We begin by enhancing the feature extraction process of the speech quality assessment model, which aims to improve the model’s accuracy in predicting Mean Opinion Scores (MOS). The main contributions are as follows: 1) We employ coordinate attention to enhance the feature extractor, thereby improving the model’s ability to capture the inherent properties of the Mel spectrum. 2) We propose a novel multi-scale feature generation method and design a corresponding multi-scale position encoding mechanism to enable the model to extract richer information. 3) We propose a contrastive learning pre-training strategy that allows the model to learn perceptually relevant features. Experimental results demonstrate that our proposed model outperforms previous work on three widely used datasets.

Read the paper · More papers on PaperTik