DiDi Labs’ End-to-end System for the IWSLT 2020 Offline Speech TranslationTask
Arkady Arkhangorodsky, Yiqi Huang, Amittai E. Axelrod · 2020
We describe the DiDi Labs system submitted for the IWSLT 2020 Offline Speech Translation Task (Ansari et al., 2020).We trained an end-to-end system that translates audio from English TED talks to German text, without producing intermediate English text.Our base system used the S-Transformer architecture (Di Gangi et al., 2019b), trained using the MuST-C dataset (Di Gangi et al., 2019a).We extended the system via decoder pre-training, pre-trained speech features, and text translation, but these extensions did not yield improved results.