ADTMOS – Synthesized Speech Quality Assessment Based on Audio Distortion Tokens

Qiao Liang, Ying Shen, T. Chen, Lin Zhang, Shengjie Zhao · IEEE Transactions on Audio Speech and Language Processing · 2025

In the fields of voice conversion (VC) and text-to-speech (TTS), recent years have witnessed a growing interest in developing synthesized speech quality assessment (SQA) systems. For such systems, it is essential to reliably and accurately evaluate the quality of synthesized speech produced by VC and TTS systems, as this remains a crucial issue requiring further exploration. Among various evaluation standards of speech quality, the mean opinion score (MOS) is the most commonly used SQA metric. The rapid advancement of deep learning (DL) techniques has propelled the emergence of DL-based MOS-based SQA algorithms. Unfortunately, none of these methods incorporate listeners' perceptions of audio distortions, which are considered one of the key factors affecting listeners' MOS ratings. To fill such a research gap to some extent, we propose a novel speech quality assessment framework, namely ADTMOS (Audio Distortion Token-Guided Deep MOS Predictor). ADTMOS consists of three parts: a public encoding layer which encodes the audio embeddings, an audio distortion token extractor which extracts ADT scores related to the subjective perceptions of audio distortions, and a frame-wise MOS score generator which is responsible for computing frame-level MOS scores. Experimental results demonstrate that compared to the LDNet baseline, ADTMOS achieves a 0.83% improvement on the VCC2018-CSMSC dataset and a 4.58% increase on the BVCC dataset in the system-level Spearman's rank correlation coefficient (SRCC). Furthermore, two innovative data augmentation techniques have been developed for the SQA task, aiming to mitigate the challenges of data scarcity and uneven sample distribution commonly encountered in SQA datasets.

Read the paper · More papers on PaperTik