SpeechUT: Bridging Speech and Text with Hidden-Unit for Encoder-Decoder Based Speech-Text Pre-training

Ziqiang Zhang, Long Zhou, Junyi Ao, Shujie Liu, Li-Rong Dai, Jinyu Li, Furu Wei · 2022

The rapid development of single-modal pretraining has prompted researchers to pay more attention to cross-modal pre-training methods.In this paper, we propose a unified-modal speech-unit-text pre-training model, SpeechUT, to connect the representations of a speech encoder and a text decoder with a shared unit encoder.Leveraging hidden-unit as an interface to align speech and text, we can decompose the speech-to-text model into a speech-to-unit model and a unit-to-text model, which can be jointly pre-trained with unpaired speech and text data respectively.Our proposed SpeechUT is fine-tuned and evaluated on automatic speech recognition (ASR) and speech translation (ST) tasks.Experimental results show that SpeechUT gets substantial improvements over strong baselines, and achieves state-of-the-art performance on both the Lib-riSpeech ASR and MuST-C ST tasks.To better understand the proposed SpeechUT, detailed analyses are conducted.The code and pretrained models are available at https://aka. ms/SpeechUT.

Read the paper · More papers on PaperTik