Addressing unknown word problem for neural machine translation using distributee representations of words as input features
Tomoki Nishimura, Tomoyosi Akiba · 2017
In recent years, the machine translation system based on neural network, called Neural Machie Translation, have attracted much attention, in which the entire translation steps are implemented in a single large neural network. In this framework, dealing with a large vocabulary size on its input (source) and output (target) often make the training computationally intractable. Therefore, the most frequent words in training data are retained to form a small vocabulary (shortlist), and the other, not frequent, words are all mapped to a single shared token. That causes so-called the unknown word problem. In this work, we propose three, rather simple, methods to overcome the unknown word problem based on distributed representation of words. We compared the translation performances of baseline and the proposed methods through experimental evaluation. Though the proposed methods did not improved the baseline in terms of BLEU, we found several evidences that the proposed methods successfully select appropriate target words even if their source words are out-of-vocabulary.