Semi-supervised ASR based on Iterative Joint Training with Discrete Speech Synthesis
Keiya Takagi, Tomoyosi Akiba, Hajime Tsukada · 2022 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) · 2022
This paper proposes Iterative Joint Training (IJT) with discrete speech synthesis for semi-supervised ASR. Thanks to the recent advances in Text-to-Speech synthesis, synthesized speech has become near human-level. However, a large amount of paired speech and text is indispensable to training a high quality TTS model. That prevents low-resource ASR from using TTS for data augmentation. To overcome that problem, we propose to use discrete speech representations, which may be easier to synthesize than continuous representations. The proposed method trains two models from the initial paired data; a speech recognition model whose input is discrete speech representation and a speech synthesis model that generates discrete speech from texts. Both models are repeatedly updated by pseudo-pair data generated by unpaired data using the previous models. Experimental results showed the effectiveness of the proposed method in low-resource settings. It successfully improved recognition performance by using discrete speech representations instead of conventional acoustic features in IJT experiments with a single-speaker speech corpus. Furthermore, the method improved the performance of multi-speaker speech recognition that used only single speaker pair data, unpaired multi-speaker speech, and unpaired text data from the conventional IJT approach using the conventional acoustic features.