SW-SQA: A Speech Quality Assessment Method Based on Extended Siamese Neural Network
Yuxi Lai, Yitong Liu, Jing Wang, Yun Cheng Shen, Hongwen Yang · 2025
Assessing speech quality, particularly the Mean Opinion Score (MOS) directly related to human perception, is crucial for communication service providers. Neural networks can effectively address the time and financial costs associated with traditional methods for obtaining MOS. However, neural network methods still face challenges, due to the noise introduced by subjective evaluations and the difficulty of establishing a relationship between speech and MOS. Observing the differences in spectrograms and waveforms among samples with varying MOS values, a novel non-intrusive SQA method is proposed, based on an Extended Siamese Neural Network and termed SW-SQA, ensuring that the distance between encoder-generated features for two samples is positively correlated with their MOS differences, effectively capturing the differences in input data among samples with varying MOS values. Furthermore, a weighted loss function is designed to direct the network's attention to samples with lower subjective noise, thereby mitigating the impact of subjective noise. Experimental results show that SW-SQA reduces RMSE by 12.4 % compared to the state-of-the-art and improves PLCC and SRCC by 2.9 % and 2.5 %, respectively.