Prediction of Perceived Synthesized Speech Quality with Wav2Vec2 Features on Small Dataset

Ivan Halim Parmonangan, Jennifer Santoso · 2022

The perceived quality of the synthesized audio is one of many factors that may determine the success of a speech synthesizer system in the market. Assessing the perceived speech audio quality is usually done subjectively and objectively. Since the subjective approach is costly and time-consuming, a faster and more efficient objective approach is developed. In an ob-jective approach, only the audio signal is used; The assessment is done either by comparing the difference between the original and synthesized speech or by creating a model to predict the quality. In a text-to-speech system, there is no audio input in the implementation phase. Therefore, evaluating its quality should be done by model prediction. However, human-labeled data is scarce, which leads to generalization issues with machine learning methods. In this work, we propose a transfer-learning method for predicting perceived synthesized speech quality, where the features are extracted from self-supervised pretrained Wav2Vec2. We experimented and evaluated the difference in the prediction performance of the basic Mel spectrogram features with the Wav2Vec2's feature-encoder, which learns the features from the input signal and contextual-encoder features, which learns the encoded features passed from the feature-encoder and using its embedding as the label for the self-supervised training purpose. Our result shows that the use of feature-encoder features yields better performance compared to the one using contextual-encoder features.

Read the paper · More papers on PaperTik