RAMP+: Retrieval-Augmented MOS Prediction With Prior Knowledge Integration
Hui Wang, Shiwan Zhao, Xiguang Zheng, Jiaming Zhou, Xuechen Wang, Yong Qin · IEEE Transactions on Audio Speech and Language Processing · 2025
Neural network-based automatic Mean Opinion Score (MOS) prediction plays a pivotal role in assessing the perceptual quality of synthetic speech. Historically, this field has struggled with limited data resources, a problem not fully resolved by existing methodologies. On the one hand, self-supervised learning (SSL) models could enhance the representational power of the feature extractor by extensive pre-training on large datasets, but fail to address the data scarcity problem affecting the downstream decoder. On the other hand, methods such as data augmentation and the incorporation of individual judges' assessments could improve data utilization, but hard to ensure a satisfactory balance between the scale and quality of the additional dataset. To address these challenges of decoder inefficiency and suboptimal data utilization, we introduce a retrieval-augmented MOS prediction method with prior knowledge integration. This method employs a new approach to using SSL features by retrieving similar instances in the feature space to obtain scores, thereby boosting decoder performance. We also optimize dataset utilization by exploiting quality-related information from the dataset as prior knowledge. Furthermore, our framework includes a dynamic retrieval scope adaptation and a fusion network, enhancing both robustness and accuracy while minimizing the need for manual parameter tuning. Experimental results demonstrate that our retrieval approach, enhanced by prior knowledge, significantly outperforms existing state-of-the-art methods in the synthetic speech MOS prediction task across various scenarios.