NTTSU at WMT2024 General Translation Task
Minato Kondo, Ryo Fukuda, Xiaotian Wang, Katsuki Chousa, Masato Nishimura, Kosei Buma, Takatomo Kano, Takehito Utsuro · 2024
The NTTSU team's submission leverages several large language models developed through a training procedure that includes continual pre-training and supervised fine-tuning.For paragraph-level translation, we generated synthetic paragraph-aligned data and used these data for training.In the task of translating Japanese to Chinese, we focused on speech domain translation.Specifically, we built Whisper models for Japanese automatic speech recognition (ASR).Since the dataset used for Whisper training contains many noisy data pairs, we combined the Whisper outputs using ROVER (Fiscus, 1997) to refine the transcriptions.Furthermore, we employed forward translation from audio as data augmentation, using both ASR models and a base translation model.To select the best translation from multiple hypotheses of the models, we applied Minimum Bayes Risk decoding after Quality Estimation (Fernandes et al., 2022), incorporating scores such as COMET-QE, COMET, and cosine similarity by LaBSE.We explored three different reranking strategies to handle two types of candidates from sentence-and paragraph-level translation and employed a fusion method that integrates all three.