Improving End-to-End Speech-to-Text Translation With Document-Level Context
Xinyu Tian, Haoran Wei, Zhengxian Gong, Junhui Li, Jun Xie · IEEE Transactions on Audio Speech and Language Processing · 2025
In recent years, end-to-end speech-to-text translation (E2E-ST) has emerged as a promising approach. Existing ST models mostly focus on learning from sentence-level speech while neglecting the valuable contextual information carried by document-level speech context. To leverage document-level context, in this paper we proposeContext-AwareSpeech-to-textTranslation (CAST), a context-aware ST model which uses document-level context to enhance the encoding of the current speech sentence under a multi-task training framework. To better leverage the bimodal document-level context during training, on the one hand, we adopt a mixup strategy that mixes up the speech and text representations at sentence-level. On the other hand, we train the model to selectively utilize contextual information by employing a selective strategy. The experimental results from the MuST-C benchmark indicate that CAST significantly enhances the sentence-level baseline, yielding an average BLEU score of 30.4 and a COMET score of 78.7 across the eight translation directions.