Boosting Context-Aware Speech Translation With Large Language Models
Yue Zhou, Yuxuan Yuan, Chengwei Zhang, Xiaodong Shi · IEEE Signal Processing Letters · 2025
With the rise of large language models (LLMs), numerous studies have incorporated LLMs into the speech domain, yielding substantial improvements in sentence-level speech-to-text translation (ST) performance. However, when faced with complex context-aware speech translation tasks, the performance of LLMs often declines, sometimes even underperforming compared to existing context-aware ST models. This paper explores how to enhance the performance of LLMs in context-aware speech translation. Specifically, we optimize the LLM's context-aware speech translation capabilities through instruction tuning, while introducing multiple related tasks, such as automatic speech recognition and machine translation, to improve translation quality in complex contexts. Additionally, we propose a Task Distribution Regularization method to promote consistency across tasks and strengthen the model's understanding of context. We also design a multi-task hybrid learning strategy to ensure efficient fine-tuning of the LLM across tasks. Experimental results demonstrate that our approach achieves excellent performance on the MuST-C and CoVoST2 benchmarks, significantly improving context-aware ST performance.