Efficient and High-Quality Neural Machine Translation with OpenNMT
Guillaume Klein, Dakun Zhang, Clément Chouteau, Josep Crego, Jean Sénellart · 2020
This paper describes the OpenNMT submissions to the WNGT 2020 efficiency shared task.We explore training and acceleration of Transformer models with various sizes that are trained in a teacher-student setup.We also present a custom and optimized C++ inference engine that enables fast CPU and GPU decoding with few dependencies.By combining additional optimizations and parallelization techniques, we create small, efficient, and highquality neural machine translation models.