A Preliminary Study on Wav2Vec 2.0 Embeddings for Text-to-Speech

Yohan Lim, Nam-Hyeong Kim, Seung Yun, Sanghun Kim, Seung‐Ik Lee · 2021 International Conference on Information and Communication Technology Convergence (ICTC) · 2021

Wav2Vec 2.0 (W2V), a self-supervised speech representation trained with massive unlabeled speech data, showed promising results on Automatic Speech Recognition (ASR). In spite of several evidences showing that W2V can generate unique acoustic features, it has been rarely utilized in Text-to-Speech (TTS) task. In this paper, we adapt W2V embed dings to TTS as feature vectors. Our TTS model consists of two components: Text2Vec, which converts a character-level text sequence into W2V embeddings, and GAN-based vocoder, which decodes a W2V embedding sequence to waveform signals. From the experiments, we observe that W2V embeddings have considerable potential as acoustic features for TTS.

Read the paper · More papers on PaperTik