Continual Gated Adapter for Bilingual Codec Text-to-Speech
Li-Jen Yang, Jen‐Tzung Chien · 2024
Text-to-speech (TTS) system has achieved desirable performance in monolingual setting with high-resourced data. However, it is challenging to carry out the cross-lingual low-resourced voice cloning by utilizing the pre-trained multilingual model given by a monolingual speaker. This paper presents an adaptive cross-lingual voice cloning under low-resourced scenario which can be flexibly performed with speech prompt driven by a neural codec TTS. In particular, a few-shot speech synthesizer is leveraged with an adapter tuning for speaker adaptation on a bilingual codec backbone model. The rehearsal-based continual learning is implemented to avoid forgetting of source language data due to fine-tuning. Importantly, a novel learnable gated adapter is introduced to enable conditional computation to handle different input features with specialized adjustment. Experimental results show a significant improvement in naturalness and speaker similarity by using the continual gated adapter.