Vec-Tok Speech: Speech Vectorization and Tokenization for Neural Speech Generation

Xinfa Zhu, Yuanjun Lv, Yi Lei, Tao Li, Wendi He, Hongbin Zhou, Heng Lu, Lei Xie · IEEE Transactions on Audio Speech and Language Processing · 2025

Language models (LMs) have recently flourished in natural language processing and computer vision, generating high-quality texts and images in various tasks. While current speech LMs have made significant progress, there are still challenges to overcome in terms of achieving optimal speech quality and broad task generalization. This paper presents Vec-Tok Speech, an extensible framework that resembles multiple speech generation tasks, generating expressive and high-fidelity speech. Specifically, we propose a novel speech codec based onspeech vectorsandsemantic tokens. Speech vectors contain acoustic details contributing to high-fidelity speech reconstruction, while semantic tokens focus on the linguistic content of speech, facilitating language modeling. Based on the proposed speech codec, Vec-Tok Speech leverages an LM to undertake the core of speech generation. Moreover, Byte-Pair Encoding (BPE) is introduced to reduce the token length and bit rate for lower exposure bias and longer context, improving the performance of LMs. Vec-Tok Speech can be used for intra- and cross-lingual zero-shot voice conversion (VC), zero-shot speaking style transfer text-to-speech (TTS), speech-to-speech translation (S2ST), speech denoising, and speaker de-identification and anonymization. Experiments show that Vec-Tok Speech, built on 50,000 hours of speech, performs better than other SOTA models.

Read the paper · More papers on PaperTik