SKIM at WMT 2023 General Translation Task
Keito Kudo, Takumi Ito, Makoto Morishita, Jun Suzuki · 2023
The SKIM team's submission used a standard procedure to build ensemble Transformer models, including base-model training, backtranslation of base models for data augmentation, and retraining of several final models using back-translated training data.Each final model had its own architecture and configuration, including up to 10.5B parameters, and substituted self-and cross-sublayers in the decoder with a cross+self-attention sublayer (Peitz et al., 2019).We selected the best candidate from a large candidate pool, namely 70 translations generated from 13 distinct models for each sentence, using an MBR reranking method using COMET and COMET-QE (Fernandes et al., 2022).We also applied data augmentation and selection techniques to the training data of the Transformer models.